Regional atmospheric pollutant distribution prediction method and system
By combining numerical simulation and machine learning methods, a neural network group for pollutant distribution prediction is established and trained, which solves the accuracy and efficiency problems of numerical simulation methods in the existing technology and the problem of insufficient data of machine learning methods, and achieves higher accuracy and efficiency of pollutant distribution prediction.
Patent Information
- Application Number
- CN202510122343.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The numerical simulation method in the prior art has problems such as difficulty in establishing a model, poor calculation accuracy, and low simulation efficiency, and the machine learning method lacks a sufficient number of measurement points to provide a large amount of accurate data for learning training.
Combining numerical simulation and machine learning methods, by meshing and numerical simulation of the prediction area, the data set of the whole domain variables in the timing are calculated, and a neural network group for pollutant distribution prediction is established and trained based on this.
Improves the accuracy and efficiency of pollutant distribution prediction, provides higher resolution and more accurate data sets, reduces the cost of data acquisition, and improves the interpretability of machine learning.
Smart Images

Figure CN120030898A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental protection technology, and in particular to a method and system for predicting regional atmospheric pollutant distribution. Background Art
[0002] The state has strict requirements for pollutant emissions from different factories and systems. Local regulatory authorities mainly conduct real-time grid-connected supervision of the pollutant measurement points of production enterprises. In addition, for regulatory purposes, regulatory agencies or corporate environmental protection departments can also install measurement points at other locations to further regulate factory production behavior and supervise the pollutant emissions of production enterprises. However, the distribution of these measurement points is discrete and sparse, which is not enough to characterize the distribution of pollutants in the entire area; on the other hand, especially for industrial areas where production enterprises are relatively concentrated, it is also necessary to accurately define the scope of factory responsibility. At the same time, the impact of weather, climate, etc. on the diffusion and evolution of pollutants should be considered to more accurately lock the direct or indirect source of responsibility of pollutants and issue control or warning instructions. The basis for pollutant concentration prediction is wind environment analysis or meteorological analysis. Existing technologies mainly include the following two categories: numerical simulation methods of different scales, but these methods often have problems such as difficulty in model establishment, poor calculation accuracy, and low simulation efficiency. Machine learning methods are used based on sufficient databases, but they require a large amount of accurate data as support for learning and training, and the number of existing measurement point settings is not enough to provide sufficient and accurate data.
[0003] The "method for predicting atmospheric pollutant concentration based on WRF-Chem" disclosed in Chinese patent documents has a publication number of CN118013769B and a publication date of 2024-06-14, which includes the following steps: S1, WRF-Chem parameterization scheme selection: obtain meteorological data, topographic data and pollution source list of the target area, and obtain the simulated value of air pollutant concentration at each grid point in the area through WRF-Chem model simulation; S2, WRF-Chem simulation result correction: use the observation data of the observation station to correct the simulated value of air pollutant concentration at each grid point to obtain the correction coefficient; S3, obtain the future forecast data of air pollutants in the target area through the NCEP global numerical weather forecast model GFS data, and based on the correction coefficient obtained in step S2, realize the prediction of atmospheric pollution concentration at any position in the target area. However, this technology is a means of simply using numerical simulation, which is highly professional, complex and difficult to establish. The large scale of the simulation grid will make the error between the calculation result and the actual result large, but the pursuit of high precision will make the solution process extremely lengthy and the simulation efficiency low. Summary of the invention
[0004] The present invention aims to overcome the problems in the prior art of difficult model establishment, poor calculation accuracy and low simulation efficiency of numerical simulation methods, and the lack of a sufficient number of measurement points to provide a large amount of accurate data for learning and training using machine learning methods, and provides a method and system for predicting the distribution of regional atmospheric pollutants.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A method for predicting regional atmospheric pollutant distribution, comprising: Establish a computational domain for the prediction area and perform grid division; Using the data of pollutant measuring points and meteorological measuring points and the intensity of pollutant emission sources as known conditions, non-steady-state numerical simulation is performed to calculate the data set of global variables in time series; Based on the time series data set of global variables, a neural network group for pollutant distribution prediction is established and trained; Based on the data collected from pollutant measuring points and meteorological measuring points, the pollutant distribution in the area to be predicted is predicted through a neural network group.
[0006] The present invention first analyzes the data association relationship based on the mechanism, and uses the analysis results to guide the establishment of a neural network group for pollutant prediction. In the specific field of pollutant diffusion and evolution in the region, the prior theoretical knowledge in fluid mechanics and chemical reactions is used for the structural design of the neural network group, thereby achieving reliability that is better than that of a general neural network in a database of equal quality, and improving the interpretability of machine learning. Through a large number of calculations performed through numerical simulation in the first step, the grid scale can be much smaller than the grid scale of other commonly used numerical calculations, providing sufficient and accurate data sets for the second step of machine learning prediction. The data volume and spatial resolution are significantly higher than the existing discrete measurement point method, and the cost of obtaining data is also much lower than that of point collection using fixed measurements, drone patrols, etc.
[0007] Preferably, the step of establishing a computational domain for the area to be predicted and performing grid division comprises: Determine the range of the area to be predicted according to the object of interest, obtain the three-dimensional spatial range of the object of interest and its contents, simplify them, and use the simplified spatial area as the calculation domain; The computational domain is divided into a number of grids, each of which is empty or includes at least one of a pollutant measurement point, a meteorological measurement point, and a pollutant emission source.
[0008] Preferably, the neural network group includes: a distribution subnetwork that converts numerical simulation results into variable distribution fields; a pollution source subnetwork that uses the output of the distribution subnetwork and the numerical simulation results as inputs, outputs the impact results of pollution sources to the prediction subnetwork; a convection subnetwork that outputs the impact results of pollutant convection to the prediction subnetwork; and a diffusion subnetwork that outputs the impact results of pollutant diffusion to the prediction subnetwork. The prediction subnetwork outputs the final pollutant distribution prediction results.
[0009] Preferably, the distribution subnetwork uses the environmental data of the meteorological measuring point at a certain moment in the numerical simulation results combined with the concentration data of the pollutant measuring point and the intensity data of the pollutant emission source as input; Taking the distribution of environmental parameters and pollutant concentrations in each grid at the same time as output, a neural network sub-model is established and trained.
[0010] Preferably, in the pollution source sub-network, the output result of the distribution sub-network at the current moment is combined with the input of the distribution sub-network at the next moment as input; Taking the pollutant concentration increment contributed by the pollution sources in each grid as the output, a neural network sub-model is established and trained.
[0011] Preferably, in the convection subnetwork, the output result of the distribution subnetwork at the current moment is used as input together with the environmental data of the meteorological measuring point and the concentration data of the pollutant measuring point at the next moment; The pollutant concentration increment in each grid due to convection with adjacent grids is used as output to establish and train a neural network sub-model.
[0012] Preferably, in the diffusion subnetwork, the output result of the distribution subnetwork at the current moment is taken as input together with the environmental data of the meteorological measuring point and the concentration data of the pollutant measuring point at the next moment; The pollutant concentration increment in each grid due to mutual diffusion with adjacent grids is used as output, and a neural network sub-model is established and trained.
[0013] Preferably, in the prediction subnetwork, the output result of the pollution source subnetwork is used as input in combination with the output results of the diffusion subnetwork and the convection subnetwork; Taking the total amount of pollutant concentration changes in each grid as output, a neural network sub-model is established and trained.
[0014] Preferably, a pollutant concentration threshold is set, and the pollutant concentration predicted in each grid is compared with the pollutant concentration threshold. If the pollutant concentration predicted in the grid is greater than the pollutant concentration threshold, a pollutant emission warning is issued for the grid that exceeds the pollutant concentration threshold.
[0015] A prediction system for regional atmospheric pollutant distribution, comprising: The data acquisition module collects meteorological and pollutant measurement point data and obtains the intensity of pollutant emission sources; the regional preprocessing module establishes the calculation domain and divides the grid for the predicted area; The numerical simulation module uses the data of pollutant measurement points and meteorological measurement points and the intensity of pollutant emission sources as known conditions to perform non-steady-state numerical simulation and calculate the data set of global variables in time series; The neural network module establishes and trains a neural network group for pollutant distribution prediction based on the data set of global variables in time series.
[0016] The present invention has the following beneficial effects: it analyzes the data association relationship based on the mechanism and guides the establishment of a neural network group for pollutant prediction. In the specific field of pollutant diffusion and evolution in the region, the prior theoretical knowledge in fluid mechanics and chemical reactions is used for the structural design of the neural network group, and the reliability is better than that of the general neural network in the database of the same quality, and the interpretability problem of machine learning is improved; through the first step of numerical simulation to perform a large number of calculations, the grid scale can be much smaller than the grid scale of other commonly used numerical calculations, providing sufficient and accurate data sets for the second step of machine learning prediction, and the data volume and spatial resolution are significantly higher than the existing discrete measurement point method, and the cost of obtaining data is also much lower than that of point collection using fixed measurement, drone patrol, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of a method for predicting regional atmospheric pollutant distribution in the present invention.
[0018] Figure 2 It is a schematic diagram of meshing the computational domain in an embodiment of the present invention.
[0019] In the figure: 1. Grid; 2. Pollutant emission source; 3. Meteorological measurement point; 4. Pollutant measurement point. DETAILED DESCRIPTION
[0020] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.
[0021] The basis for pollutant concentration prediction is wind environment analysis or meteorological analysis. Existing technologies mainly include the following two categories: numerical simulation methods of different scales, mainly including macro-scale meteorological simulation, meso-scale meteorological simulation, micro-scale wind environment simulation, computational fluid dynamics (CFD) simulation, multi-scale mixed simulation, etc. However, these methods often have problems such as difficulty in model establishment, poor calculation accuracy, and low simulation efficiency. Using machine learning methods based on sufficient databases, historical meteorological and pollutant distribution data can be analyzed for correlation, and output values can be predicted under given input conditions, but there are problems with insufficient training data and calculation efficiency.
[0022] In order to overcome the problems and shortcomings of using only one of the above two methods, the present invention combines numerical simulation with machine learning methods, uses the results of numerical simulation to provide data for machine learning, and provides Figure 1 A method for predicting the distribution of regional atmospheric pollutants is shown, comprising: Establish a computational domain for the prediction area and perform grid division; Using the data of pollutant measuring points and meteorological measuring points and the intensity of pollutant emission sources as known conditions, non-steady-state numerical simulation is performed to calculate the data set of global variables in time series; Based on the time series data set of global variables, a neural network group for pollutant distribution prediction is established and trained; Based on the data collected from pollutant measuring points and meteorological measuring points, the pollutant distribution in the area to be predicted is predicted through a neural network group.
[0023] It should be noted that, in the present invention, the data correlation relationship based on the mechanism is first analyzed, and the analysis results are used to guide the establishment of a neural network group for pollutant prediction. In the specific field of pollutant diffusion and evolution in the region, the prior theoretical knowledge in fluid mechanics and chemical reactions is used for the structural design of the neural network group, so as to achieve reliability better than that of general neural networks in databases of equal quality, and improve the interpretability of machine learning. Through a large number of calculations through numerical simulation in the first step, the grid scale can be much smaller than the grid scale of other commonly used numerical calculations, providing sufficient and accurate data sets for the second step of machine learning prediction. The data volume and spatial resolution are significantly higher than the existing discrete measurement point method, and the cost of data acquisition is also much lower than that of point collection using fixed measurements, drone patrols, etc.
[0024] As a specific embodiment, the neural network group includes: a distribution subnetwork for converting numerical simulation results into a variable distribution field; The output of the distribution subnetwork and the numerical simulation results are used as input: the pollution source subnetwork outputs the impact results of the pollution source to the prediction subnetwork; the convection subnetwork outputs the impact results of the pollutant convection to the prediction subnetwork; the diffusion subnetwork outputs the impact results of the pollutant diffusion to the prediction subnetwork; The prediction subnetwork outputs the final pollutant distribution prediction results.
[0025] It should be noted that the neural network group in the present invention is composed of a total of five sub-neural networks. For the sub-neural networks, the most basic fully connected neural network can be used to achieve the purpose of the present invention. Similarly, other existing neural network structures can be selected according to actual needs. In the neural network group, the distribution sub-network is the top-level neural network, which simulates the process of obtaining the distribution state of global variables in each grid during the numerical simulation process; the pollution source sub-network, convection sub-network and diffusion sub-network are the neural networks of the middle layer, and the influence of different influencing factors on the distribution of pollutants is obtained according to the distribution state of global variables in each grid; the prediction sub-network is the bottom-level neural network, and the final pollutant distribution prediction is comprehensively performed based on the output results of the middle-level neural network. All sub-neural networks are trained with the data set of numerical simulation results, and finally form a neural network group that can predict pollutant concentrations.
[0026] It is worth mentioning that the application object of the present invention is the flow reaction problem with extremely strong nonlinearity. According to the control equation that describes its detailed mechanism, it is divided into three different influencing factors: convection, diffusion and source, and the sub-neural networks are trained separately to form a neural network group to calculate and solve the flow field distribution and evolution.
[0027] Specifically, for y=y with a strong mechanistic description 1 +y 2 +y 3 =f(x), the traditional CFD-machine learning coupling method is to solve y directly from x; while the method of the present invention uses multiple sub-neural networks to achieve x->y' 1 , x->y' 2 , x->y' 3 The approximate solution is obtained, and then further modified to obtain the final output y. Therefore, compared with the existing technology, the present invention has the following two advantages: First, the sub-neural networks obtained by this splitting method have clear physical meanings, which improves the defects of the traditional CFD-machine learning coupled black box method in terms of interpretability; second, because the solution is carried out one by one on the computational domain grid, y i =y 1,i +y 2,i +y 3,i(i is the grid number), it is easy to know that in the whole domain, the y value range that the traditional network can predict is: y min ≥y 1,min +y 2,min +y 3,min ,y max ≤y 1,max +y 2,max +y 3,max The y value range that the split network can predict is: y min =y 1,min +y 2,min +y 3,min ,y max =y 1,max +y 2,max +y 3,max Only a simple modification of the prediction subnetwork is required (such as adding physical constraints: y' i =y' 1,i +y' 2,i +y' 3,i ) can be achieved. Note that this correction only increases the error caused by the strong physical constraint in the prediction subnetwork; obviously, this simple correction is extremely difficult to implement in traditional networks. In summary, from the perspective of variable value range, this splitting method has the ability to cover a wider range of parameters and has a wider range of applicability than existing technologies.
[0028] The prior knowledge used in the present invention mainly comes from the basic equations and mechanisms in the disciplines of fluid mechanics, heat and mass transfer, chemical reactions, etc. The pollutant concentration increment in a grid is decomposed into the contribution of the source, convection and diffusion, and a neural network group consisting of multiple sub-neural networks is established accordingly. Therefore, the purpose of the invention can also be achieved by constructing neural networks with different structures or other machine learning algorithms in other ways; however, the algorithms constructed in this way may not have clear physical meanings, that is, they are not easily interpretable and are difficult to iterate or maintain.
[0029] As a specific example, Figure 2 As shown in the figure, the calculation domain is established for the prediction area, and the grid division includes: Determine the range of the area to be predicted according to the object of interest, obtain the three-dimensional spatial range of the object of interest and its contents, simplify them, and use the simplified spatial area as the calculation domain; The calculation domain is divided into a number of grids 1 , each of which is empty or includes at least one of a pollutant measurement point 4 , a meteorological measurement point 3 , and a pollutant emission source 2 .
[0030] It should be noted that, according to the terrain and building surveying information of the Geographic Information System (GIS system), the calculation domain of the prediction area is determined and the calculation grid is divided, and the grid information including the node position is stored. The calculation domain and grid division include the following steps: The subject implementing this method (government departments, regulatory agencies, etc.) can determine the scope of the area to be predicted based on the specific object of concern. For example, the Environmental Protection Bureau of District B in City A can take the entire jurisdiction as the object of concern; the head of the Environmental Protection Department of Factory C can take the entire factory area as the object of concern.
[0031] The three-dimensional spatial scope and contents of the object of interest are obtained from GIS, and a certain simplification is performed. For example, the outer boundary of the computational domain is the spatial boundary of the object of interest, large buildings are retained inside the computational domain (the outer wall of large buildings becomes the solid wall boundary in the computational domain), and small buildings are ignored (it is assumed that the place is ventilated on all sides and there are no obstacles blocking the air flow), etc. The simplified spatial area is the computational domain.
[0032] The method of dividing the computational domain is consistent with the grid division method in CFD calculations. The computational domain is divided into grids of similar geometric sizes, that is, the size of the divided grid is within the allowable size fluctuation range. If the computing power allows, generally speaking, the smaller the grid size, the more accurate the result. In actual operation, the CFL criterion is generally used to preliminarily determine the grid size, and then the grid is improved by increasing or decreasing the grid size, locally encrypting the boundary layer grid, and other operations to ensure that the "grid independence" requirements are met, that is: if the grid is continued to be encrypted, the accuracy will not be significantly improved. To expand to three-dimensional space, you only need to divide the space into multiple small blocks (cubes are fine, deformation is allowed).
[0033] The role of nodes is to transform continuous space into discrete points, and use discrete points to approximate continuous distribution in space. By drawing a grid, you can intuitively see the location of the nodes. Different CFD algorithms have different ways of determining grids and nodes. Commonly used algorithms include: finite volume method (FVM), finite difference method (FDM), finite element method (FEM), meshless method (MeshlessMethods), etc. Taking the most classic FVM SIMPLE algorithm as an example, two sets of nodes are generally used: one set places the node at the center of the grid to store scalars such as density; the other set places the node at the center of the grid surface to store vectors such as velocity. There are also those that directly use grid vertices as the nodes themselves. Each method has different requirements for nodes. For the same calculation object, there can also be different methods for determining grids and nodes. It only needs to be verified that it meets the "grid independence" requirement.
[0034] As a specific embodiment, the data of pollutant measuring points and meteorological measuring points and the intensity of pollutant emission sources are used as known conditions to perform non-steady-state numerical simulation and calculate a data set of global variables in time series.
[0035] Specifically, according to the wind speed {V m}、Pressure {P m}、Temperature{T m}、Concentration of pollutant at measuring point {C p}、 Pollutant emission source intensity {S q}, perform unsteady computational fluid dynamics (CFD) simulation to obtain the time series data set of global variables in the computational domain, including the wind speed V of any grid i i , any component concentration C i,j , the convection term Conv in each control equation i , Diffusion term Diff i , pollution source term S i Evolution over time. The source term includes pollutant emission source term and chemical reaction source term. i is the grid number and j is the component number. The subscript m corresponds to the meteorological measurement point number; the subscript p corresponds to the pollutant measurement point number; and the subscript q corresponds to the pollutant emission source number.
[0036] In the numerical simulation part, the monitoring area is divided according to the discrete and sparsely distributed pollutant measurement points, meteorological measurement points and pollutant emission sources inside and outside the predicted area. The measurement points can be composed of grid-connected pollutant emission sources of enterprise groups such as production enterprises, and other meteorological or pollutant measurement points arranged by environmental supervision departments or production enterprises can also be added.
[0037] For the measurement points at the pollutant emission sources, at least the type, flow rate, temperature and pressure of the pollutants emitted should be included (i.e., sufficient to make the emission intensity of the pollution source clear). For meteorological measurement points, at least environmental parameters such as temperature, pressure and wind speed should be included; for pollutant measurement points, the main data collected is the pollutant concentration data.
[0038] Regarding the selection of grid size / resolution: For grid division in general problems, the grid scale needs to be determined by balancing the accuracy requirements and computing power. Because the numerical simulation calculation results of the present invention will be included in the pollutant distribution database, which has the purpose of long-term use, more computing resources are invested in the numerical simulation part, and high-resolution grid division is performed to build a reliable database for subsequent neural network group training.
[0039] The pollutant emission source information from the production enterprises is included in the source term of the control equation, and the control equations and models of computational fluid dynamics, mass transfer (the first two are necessary), heat transfer, and chemical reaction (the latter two are optional) are solved by non-steady-state numerical simulation to obtain the distribution form and change law of variables such as air temperature, wind speed, and pollutant concentration in the monitoring area. For example, for the prediction of haze (particulate pollutants), Ansys Fluent commercial software is used. According to the aforementioned calculation domain, grid division and boundary conditions, the Euler model is selected for fluid flow solution, and the DPM (Discrete Phase Model) model is selected for particle flow solution. The particle event model is added through UDF (User Defined Function) (this step is optional, and you can add, for example: collision-coalescence / fragmentation model driven by kernel function), and the distribution of particle pollutants in each calculation domain can be obtained.
[0040] As a specific embodiment, in the part where the neural network group performs prediction, the high-resolution distribution results obtained in the numerical simulation step are used as the data source for machine learning to establish a pollutant distribution database. New pollutant distribution data can be added to the database at any time through numerical simulation calculations using newly acquired data (as long as any point in the meteorological measurement point data or the pollutant measurement point data changes, it can be called new). The data in the database is preprocessed and divided into training sets and test sets. Three common division methods can be used: retention method, cross-validation method, and self-help method. Determine the structure of the neural network group, use the meteorological measurement point data and pollutant measurement point data in the database as input, and use the pollutant concentration distribution of each grid in the region as output. Through training and testing, the neural network group that is ultimately used for pollutant concentration prediction is obtained.
[0041] The distribution subnetwork takes the environmental data of the meteorological measuring point at a certain moment in the numerical simulation results, the concentration data of the pollutant measuring point, and the intensity data of the pollutant emission source as input; Taking the distribution of environmental parameters and pollutant concentrations in each grid at the same time as output, a neural network sub-model is established and trained.
[0042] Specifically, the input of the distribution subnetwork is: the wind speed {V m,t}、Pressure {P m,t}、Temperature{T m,t}、Concentration of pollutant at measuring point {C p,t}、 Pollutant emission source intensity {S q,t The output of the distribution subnetwork is: the wind speed distribution field {V i,t}、Pressure distribution field {P i,t}、Temperature distribution field {T i,t}、Pollution concentration distribution field{C i,t The distribution subnetwork is trained using numerical simulation data sets, so that the trained distribution subnetwork can obtain the meteorological data distribution and pollutant concentration distribution in each grid in the area to be predicted at that moment based on the meteorological measurement point data, pollutant measurement point data and pollutant emission source intensity input at a certain moment.
[0043] In the pollution source sub-network, the output result of the distribution sub-network at the current moment is combined with the input of the distribution sub-network at the next moment as input; Taking the pollutant concentration increment contributed by the pollution sources in each grid as the output, a neural network sub-model is established and trained.
[0044] Specifically, the input of the pollution source subnetwork is: the wind speed distribution field {V i,t}、Pressure distribution field {P i,t}、Temperature distribution field {T i,t}、Pollution concentration distribution field{C i,t}, and the wind speed {V m,t+dt}、Pressure {P m,t+dt}、Temperature{T m,t+dt}、Concentration of pollutant at measuring point {C p,t+dt}、 Pollutant emission source intensity {S q,t+dt}. (When the value at time t+dt is not available, any reasonable user-defined value can be used, such as empirical estimates, simulation values, extrapolated values obtained based on trends, or directly using the measured value at time t instead).
[0045] The output of the pollutant subnetwork is: pollutant emission source S i Contributed pollutant concentration increment C a,i , the physical meaning is the change in pollutant concentration caused by pollutant emission sources and chemical reaction source items in grid i within dt. The pollution source subnetwork is trained using the output results of the distribution subnetwork at the same time and the data of meteorological measurement points, pollutant measurement points and pollutant emission sources at the next time to obtain a pollution source subnetwork that meets the requirements of the neural network group.
[0046] In the convection subnetwork, the output result of the distribution subnetwork at the current moment is combined with the environmental data of the meteorological measurement point and the concentration data of the pollutant measurement point at the next moment as input; The pollutant concentration increment in each grid due to convection with adjacent grids is used as output to establish and train a neural network sub-model.
[0047] Specifically, the input of the convection subnetwork is: the wind speed distribution field {V i,t}、Pressure distribution field {P i,t}、Temperature distribution field {T i,t}、Pollution concentration distribution field{C i,t}, and the wind speed {V m,t+dt}、Pressure {P m,t+dt}、Temperature{T m,t+dt}、Concentration of pollutant at measuring point {C p,t+dt}. (When the value at time t+dt is not available, any reasonable user-defined value can be used, such as empirical estimates, simulation values, extrapolated values obtained based on trends, or directly using the measured value at time t instead).
[0048] The output of the convection subnetwork is: Conv i and Conv in Contributed pollutant concentration increment C b,i , the physical meaning is the change of pollutant concentration caused by convection in grid i within dt. The subscript in means i-neighbor, which represents all grids adjacent to grid i. The output results of the distribution subnetwork at the same time and the results of the numerical simulation at the next time are used to train the convection subnetwork to obtain a convection subnetwork that meets the requirements of the neural network group.
[0049] In the diffusion subnetwork, the output result of the distribution subnetwork at the current moment is combined with the environmental data of the meteorological measurement point and the concentration data of the pollutant measurement point at the next moment as input; The pollutant concentration increment in each grid due to mutual diffusion with adjacent grids is used as output, and a neural network sub-model is established and trained.
[0050] Specifically, the input of the diffusion subnetwork is: the wind speed distribution field {V i,t}、Pressure distribution field {P i,t}、Temperature distribution field {T i,t}、Pollution concentration distribution field{C i,t}, and the wind speed {V m,t+dt}、Pressure {P m,t+dt}、Temperature{T m,t+dt}、Concentration of pollutant at measuring point {C p,t+dt}. (When the value at time t+dt is not available, any reasonable user-defined value can be used, such as empirical estimates, simulation values, extrapolated values obtained based on trends, or directly using the measured value at time t instead).
[0051] The output of the diffusion subnetwork is: Diff i and Diff in Contributed pollutant concentration increment C c,i, the physical meaning is the change in pollutant concentration caused by diffusion in grid i within dt. The subscript in means i-neighbor, representing all grids adjacent to grid i. The diffusion subnetwork is trained using the output results of the distribution subnetwork at the same time and the results of the numerical simulation at the next time to obtain a diffusion subnetwork that meets the requirements of the neural network group.
[0052] In the prediction subnetwork, the output of the pollution source subnetwork is used as input together with the output of the diffusion subnetwork and the convection subnetwork. Taking the total amount of pollutant concentration changes in each grid as output, a neural network sub-model is established and trained.
[0053] Specifically, the input of the prediction subnetwork is: the pollutant emission source S at any time t i Contributed pollutant concentration increment C a,i 、Conv i and Conv in Contributed pollutant concentration increment C b,i and Diff i and Diff in Contributed pollutant concentration increment C c,i The output of the prediction subnetwork is: C total,i , the physical meaning is the total change of pollutant concentration of grid i within dt. The pollutant concentration distribution at any time t combined with the total change of pollutant concentration of each grid within dt can get the pollutant concentration prediction result at the next time t+dt. Other outputs of the prediction subnetwork also include the distribution of environmental parameters in each grid; the neural network group obtained by combining the five subnetworks can predict the pollutant concentration of each grid.
[0054] After training the neural network group composed of five sub-networks, the wind speed {V m}、Pressure {P m}、Temperature{T m}、Concentration of pollutant at measuring point {C m}Input into the neural network group for result prediction.
[0055] Optionally, when predicting the pollutant concentration distribution, it is only necessary to use the distribution subnetwork once at time t = 0 to obtain the distribution at time t = 0 (i.e., obtain a relatively accurate initial distribution field) as the starting input of other subnetworks. 0 The wind speed at the meteorological point measured {V m,t0}、Pressure {P m,t0}、Temperature{T m,t0}、Concentration of pollutant at measuring point {C p,t0}、 Pollutant emission source intensity {S q,t0} as the input of the distribution sub-network. The output of the distribution sub-network is: 0 The wind speed distribution field {V i,t0}、Pressure distribution field {P i,t0}、Temperature distribution field {T i,t0}、Pollution concentration distribution field{C i,t0}. Then the subsequent sub-network is used for t 1 =t 0 +dt time result prediction. 1 =t 0 The result prediction at time +dt can be used as new input data to be input into the middle layer sub-network and the bottom layer sub-network for t 2 =t 1 +dt time result prediction; then only the middle layer and the bottom layer sub-network can be used to continuously complete t n =t n-1 +dt time prediction advance.
[0056] As an optional embodiment, a pollutant concentration threshold is set, and the pollutant concentration predicted in each grid is compared with the pollutant concentration threshold. If the pollutant concentration predicted in the grid is greater than the pollutant concentration threshold, a pollutant emission warning is issued for the grid that exceeds the pollutant concentration threshold.
[0057] Specifically, the current pollutant limit requirements can be used as the standard (which can be national standards, local standards, or any reasonable values artificially limited by the competent authorities); in the simulation and prediction of pollutant distribution, the above standards are used to determine whether there are grids in the region with pollutant concentrations exceeding the limit in the current or future period, and pollutant emission warnings are issued for grids that exceed the limit.
[0058] In addition to the method for predicting the distribution of regional atmospheric pollutants, the present invention also provides a system for predicting the distribution of regional atmospheric pollutants, including: The data acquisition module collects meteorological and pollutant measurement point data and obtains the intensity of pollutant emission sources; the regional preprocessing module establishes the calculation domain and divides the grid for the predicted area; The numerical simulation module uses the data of pollutant measurement points and meteorological measurement points and the intensity of pollutant emission sources as known conditions to perform non-steady-state numerical simulation and calculate the data set of global variables in time series; The neural network module establishes and trains a neural network group for pollutant distribution prediction based on the global variable time series data set; Database, which saves the simulation results obtained by the numerical simulation module.
[0059] It should be noted that the existing machine learning prediction methods mainly analyze and study historical monitoring data, explore the connections between data, find pollution patterns, and establish prediction models. The meteorological data used for analysis by this method comes from the public or paid databases of domestic and foreign meteorological agencies (the grid resolution is low, and the scale is generally at the kilometer level or even higher), and the pollutant distribution data comes from limited measurement points. In other words, the existing conditions make it difficult to provide sufficient and detailed data for accurate statistical predictions. This pure black box method lacks interpretability and cannot reveal the natural science laws in the process. It is difficult to extrapolate. That is, it is difficult to predict new situations beyond the scope of historical data. In summary, pure statistical forecasting methods cannot meet the theoretical analysis requirements in cutting-edge research and the requirements for responding to changes in practical applications.
[0060] Statistical prediction methods require sufficient and accurate data as input, but the objective and actual extremely sparse distribution of pollutant concentration measurement points makes it difficult to meet the requirements for accurate prediction of any point in the entire area. The present invention can perform a large number of calculations through the numerical simulation CFD algorithm in the first step to provide sufficient and accurate data sets for the second step of machine learning prediction. Its data volume and spatial resolution are significantly higher than those of highly discrete measurement methods, and the cost of data acquisition is much lower than that of point collection using fixed measurements, drone patrols, etc.
[0061] The existing numerical simulation method is based on emission source data, and calculates the evolution of pollutant distribution from the theoretical perspectives of chemistry, atmospheric physics, fluid mechanics, etc., and analyzes the changing laws of air quality. The problems with this method are: large scale, coarse grid, and difficult modeling. Specifically: the existing model solution grid generally has a low resolution (mostly on the scale of kilometers), and often ignores geographical factors such as terrain and buildings; in addition, the existing pollutant concentration measurement points are sparsely distributed, and the atmospheric movement itself is very complex, resulting in large errors between the simulation results and the actual results (the above factors are not purely parallel, but affect each other). If you want to reduce the grid scale, increase the measurement points, and implement detailed modeling and solution, the calculation efficiency will be significantly reduced. In summary, pure numerical simulation methods cannot meet the timeliness and cost requirements in practical applications.
[0062] In the present invention, a large amount of computing power is invested in the early stage, and a database is constructed through sophisticated mathematical model solutions for later machine learning training and prediction. In other words, because the results obtained by numerical simulation are placed in the database for long-term use, after amortizing the initial large amount of computing power on a long-term scale, a reasonable (even high) calculation cost is acceptable to the present invention. Therefore, the grid scale of the numerical simulation method in the present invention can be much smaller than the grid scale (kilometer level) of other commonly used numerical simulations, and various necessary mechanisms (such as: particle agglomeration, secondary pollutant reaction) can be added to complete the model and comprehensively improve the solution accuracy.
[0063] The above embodiments are further elaborations and illustrations of the present invention for ease of understanding, and are not limitations of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for predicting the distribution of regional atmospheric pollutants, characterized in that: include: Establish a computational domain for the prediction area and perform grid division; Using the data of pollutant measuring points and meteorological measuring points and the intensity of pollutant emission sources as known conditions, non-steady-state numerical simulation is performed to calculate the data set of global variables in time series; Based on the time series data set of global variables, a neural network group for pollutant distribution prediction is established and trained; Based on the data collected from pollutant measuring points and meteorological measuring points, the pollutant distribution in the area to be predicted is predicted through a neural network group.
2. A method for predicting the distribution of regional air pollutants according to claim 1, characterized in that: The establishment of a computational domain for the area to be predicted and the grid division include: Determine the range of the area to be predicted according to the object of interest, obtain the three-dimensional spatial range of the object of interest and its contents, simplify them, and use the simplified spatial area as the calculation domain; The computational domain is divided into a number of grids, each of which is empty or includes at least one of a pollutant measurement point, a meteorological measurement point, and a pollutant emission source.
3. A method for predicting the distribution of regional atmospheric pollutants according to claim 1 or 2, characterized in that: The neural network group includes: a distribution subnetwork that converts numerical simulation results into a variable distribution field; The output of the distribution subnetwork and the numerical simulation results are used as input: the pollution source subnetwork outputs the impact results of the pollution source to the prediction subnetwork; the convection subnetwork outputs the impact results of the pollutant convection to the prediction subnetwork; the diffusion subnetwork outputs the impact results of the pollutant diffusion to the prediction subnetwork; The prediction subnetwork outputs the final pollutant distribution prediction results.
4. A method for predicting the distribution of regional air pollutants according to claim 3, characterized in that: The distribution subnetwork uses the environmental data of the meteorological measuring point at a certain moment in the numerical simulation results, the concentration data of the pollutant measuring point and the intensity data of the pollutant emission source as input; Taking the distribution of environmental parameters and pollutant concentrations in each grid at the same time as output, a neural network sub-model is established and trained.
5. A method for predicting the distribution of regional air pollutants according to claim 4, characterized in that: In the pollution source subnetwork, the output result of the distribution subnetwork at the current moment is combined with the input of the distribution subnetwork at the next moment as input; the pollutant concentration increment contributed by the pollution source in each grid is used as output, and a neural network submodel is established and trained.
6. A method for predicting regional atmospheric pollutant distribution according to claim 4, characterized in that: In the convection subnetwork, the output result of the distribution subnetwork at the current moment is combined with the environmental data of the meteorological measurement point and the concentration data of the pollutant measurement point at the next moment as input; The pollutant concentration increment in each grid due to convection with adjacent grids is used as output to establish and train a neural network sub-model.
7. A method for predicting regional atmospheric pollutant distribution according to claim 4, characterized in that: In the diffusion subnetwork, the output result of the distribution subnetwork at the current moment is combined with the environmental data of the meteorological measurement point and the concentration data of the pollutant measurement point at the next moment as input; The pollutant concentration increment in each grid due to mutual diffusion with adjacent grids is used as output, and a neural network sub-model is established and trained.
8. A method for predicting the distribution of regional air pollutants according to any one of claims 4 to 7, characterized in that: In the prediction subnetwork, the output result of the pollution source subnetwork is used as input in combination with the output results of the diffusion subnetwork and the convection subnetwork; Taking the total amount of pollutant concentration changes in each grid as output, a neural network sub-model is established and trained.
9. A method for predicting regional atmospheric pollutant distribution according to claim 8, characterized in that: A pollutant concentration threshold is set, and the predicted pollutant concentration in each grid is compared with the pollutant concentration threshold. If the predicted pollutant concentration in the grid is greater than the pollutant concentration threshold, a pollutant emission warning is issued for the grid that exceeds the pollutant concentration threshold.
10. A prediction system for regional atmospheric pollutant distribution, applicable to the prediction method according to any one of claims 1 to 9, characterized in that: include: The data collection module collects meteorological measurement point data and pollutant measurement point data, and obtains the intensity of pollutant emission sources; The regional preprocessing module is used to establish the calculation domain and divide the grid for the prediction area; The numerical simulation module uses the data of pollutant measurement points and meteorological measurement points and the intensity of pollutant emission sources as known conditions to perform non-steady-state numerical simulation and calculate the data set of global variables in time series; The neural network module establishes and trains a neural network group for pollutant distribution prediction based on the data set of global variables in time series.
Citation Information
Patent Citations
A method for predicting atmospheric pollutant concentrations based on WRF-Chem
CN118013769B
Pollutant point source diffusion model construction method based on VBEM method
CN114329996A
Atmospheric pollution traceability prediction method and system based on continuous online observation data
CN114662344A
Air mass concentration field analog simulator based on deep learning
CN116070500A
Urban air quality prediction method in combination with pollution diffusion index
CN117171546A
Cited By
Diesel engine pollutant emission early warning method
CN120911276A
Karst region double-path pollution transmission simulation method
CN122414002A