Dynamic Distribution Regulation Method of Urban Pollution Sources Based on Reinforcement Learning

Through improved Informer network and game theory multi-agent reinforcement learning, the problems of global optimal control and dynamic adaptation in pollution source regulation are solved, and high-precision pollutant concentration prediction and intelligent pollution control are achieved.

CN119904083BActive Publication Date: 2025-07-04安徽配隆天环保科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510399752.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing pollution source control methods are difficult to achieve global optimal control, lack dynamic adaptability, cannot effectively deal with complex terrain and non-steady meteorological conditions, and the coordinated governance between multiple pollution sources has not been fully considered.

Method used

The improved Informer network is used to predict pollutant concentrations in time and space, and combined with game theory, multi-agents reinforce learning, optimize pollution control strategies, optimize pollution source control strategies through Nash equilibrium, and optimize the reward mechanism for pollution historical data, and dynamically adjust pollution source control strategies.

Benefits of technology

The accuracy and spatial resolution of pollutant concentration prediction are improved, the adaptability and global optimization of pollution source regulation are achieved, and the governance plan can be dynamically adjusted under different environmental conditions, which improves the economic and environmental benefits of pollution control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904083B_ABST
    Figure CN119904083B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dynamically distributing and regulating urban pollution sources based on reinforcement learning, which includes the following steps: S1, collecting source data of air pollution, preprocessing it, and constructing a pollution data set; S2, using an improved Informer network to perform time series prediction on the pollution data set; S3, calculating the influence degree of different pollution sources on air quality, and constructing a pollution source regulation task data set; S4, establishing a game theory multi-agent reinforcement learning environment, using Nash equilibrium to optimize the regulation strategy of pollution source agents, optimizing the reward mechanism based on pollution historical data, and dynamically adjusting the pollution source regulation strategy; S5, calculating the pollution source adjustment plan and optimizing the pollution regulation parameters; S6, executing the pollution regulation strategy and performing iterative optimization. The present invention combines an improved Informer network and game theory multi-agent reinforcement learning to achieve accurate prediction of pollutant concentration and intelligent regulation of pollution sources, and has the advantages of high precision, strong self-adaptability, and global optimization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental pollution control, and particularly to a method for dynamically distributing and regulating urban pollution sources based on reinforcement learning. Background Art

[0002] The sources of air pollution are complex and diverse, including industrial emissions, vehicle exhaust, construction dust, and residential emissions. There are significant differences in the emission characteristics and diffusion patterns of different pollution sources. Therefore, how to effectively monitor, predict, and intelligently regulate the emissions of urban pollution sources has become one of the core issues in the current environmental governance field.

[0003] Existing pollution source regulation methods mainly rely on atmospheric physical diffusion models, such as the Gaussian plume model, Lagrangian particle model, and Eulerian grid model. These methods rely on mathematical modeling of pollutant diffusion laws. Although they can simulate the propagation process of pollutants in the atmosphere to a certain extent, their applicability is limited by the assumptions of the models. For example, the Gaussian plume model assumes that pollutant diffusion is a uniform and continuous process, making it difficult to handle complex terrains, non-steady meteorological conditions, and the interactive effects of multi-source pollution. In addition, traditional air pollution prediction methods rely on statistical analysis and machine learning models, such as autoregressive moving average models, support vector machines, and random forests. Although these methods have a certain degree of accuracy in short-term prediction, they still have significant limitations in long-term prediction, modeling of non-linear pollution evolution processes, and co-prediction of multiple pollutants.

[0004] In recent years, deep learning techniques have demonstrated superior performance in time series prediction tasks. Among them, long short-term memory networks and Transformer architectures have been widely applied to pollutant concentration prediction. However, the standard Transformer has a huge computational cost when dealing with long time series data. The Informer network reduces the computational complexity through a sparse self-attention mechanism and improves the prediction ability of long sequence data. It has been applied to fields such as traffic flow prediction and meteorological prediction, but there is still room for improvement in its application to air pollution prediction, such as optimizing model training for sparse data and improving the accuracy of pollutant spatial interpolation.

[0005] In terms of pollution source regulation, traditional pollutant treatment measures mainly rely on fixed regulation strategies, such as setting industry emission standards, designating low-emission areas, and implementing odd-even license plate restrictions. Although these methods can reduce pollutant emissions to a certain extent, they lack dynamic adaptability and cannot flexibly adjust treatment plans according to real-time pollution conditions and meteorological conditions. In addition, the optimal control of a single pollution source often cannot take into account the coordinated treatment of multiple pollution sources. For example, the interactive effects between industrial pollution sources and vehicle emissions are not fully considered. Therefore, existing methods are difficult to achieve the global optimal control of pollution sources.

[0006] In recent years, multi-agent reinforcement learning has been widely applied in distributed decision-making optimization tasks. For example, in the fields of intelligent traffic control, UAV cooperation, and smart grid scheduling, multi-agent systems optimize the overall benefit through collaborative learning and improve decision-making efficiency. However, in the field of pollution source regulation, research based on multi-agent reinforcement learning is still in its initial stage and mainly faces the following challenges: (1) The non-linear dynamic characteristics of pollutant diffusion lead to a complex state space for agent decision-making; (2) The interaction between pollution sources makes the optimization of the reward function non-convex, making it difficult to solve the optimal solution; (3) How to design a reasonable reward mechanism so that agents can take into account the overall pollution control goal while optimizing their own benefits. Most existing studies adopt single reinforcement learning methods, such as deep Q-network, proximal policy optimization, etc., without fully considering the game characteristics of pollution control, resulting in the strategies learned by agents may converge to local optimal solutions and cannot achieve the optimal strategy for global pollution regulation. Summary of the Invention

[0007] An object of the present invention is to propose a dynamic distribution regulation method for urban pollution sources based on reinforcement learning. The improved Informer network of the present invention is used for spatio-temporal prediction of pollutant concentration, and the pollution control strategy is optimized based on game theory multi-agent reinforcement learning. Through the improved Informer network, long-time series data is processed to improve the accuracy of pollutant concentration prediction, and the diffusion trend of pollutants is calculated by combining spatial interpolation methods. A game theory multi-agent reinforcement learning environment is established, the regulation strategy of pollution source agents is optimized using Nash equilibrium, and the reward mechanism is optimized by combining pollution historical data, so that each agent continuously learns and optimizes in a dynamic environment, and finally converges to the global optimal strategy, improving the executability and effectiveness of the pollution source regulation plan.

[0008] A dynamic distribution regulation method for urban pollution sources based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect the source data of air pollution, and perform data cleaning, format conversion, and standardization processing on the source data to construct a pollution data set;

[0010] S2. Use the improved Informer network to perform time series prediction on the pollution data set, calculate the future concentration change trend of different pollutants, and calculate the diffusion of pollutants in different regions by combining spatial interpolation methods to generate pollutant spatio-temporal distribution prediction data;

[0011] S3. Based on the pollutant spatio-temporal distribution prediction data, construct a pollution source impact factor matrix, calculate the impact degree of different pollution sources on air quality, and construct a pollution source regulation task data set by combining the pollution source emission and pollution diffusion characteristics;

[0012] S4. Based on the pollution source control task data set, establish a game theory multi-agent reinforcement learning environment, set the state space, action space and reward function of the pollution source agent, optimize the control strategy of the pollution source agent using Nash equilibrium, optimize the reward mechanism based on historical pollution data, and dynamically adjust the pollution source control strategy in combination with the pollution control resource constraint conditions to generate a globally optimal control strategy;

[0013] S5. Combine the globally optimal control strategy, calculate the pollution source adjustment plan, and optimize the pollution control parameters;

[0014] S6. Execute the pollution control strategy and iteratively optimize the pollution control strategy.

[0015] Optionally, the source data includes air pollutant concentration data and meteorological data.

[0016] Optionally, the specific steps of S2 include:

[0017] S21. Construct a pollutant time series data matrix, process the time series data, and calculate the change rate and acceleration of the pollutant concentration:

[0018] ;

[0019] ;

[0020] Among them, represents the pollutant time series data matrix, represents time position the pollutant concentration at, represents the total number of time steps, represents the geographical location index, represents time position the change rate of the pollutant concentration at, represents time position the acceleration of the pollutant concentration at, represents time position the pollutant concentration at, represents time position the pollutant concentration at, represents the time interval between adjacent time steps, represents time position the change rate of the pollutant concentration at;

[0021] S22. Set the input sequence length to , and define the input matrix:

[0022] ;

[0023] Among them, represents the input matrix, describing the input pollutant time series data, represents the input vector, represents time position the pollutant concentration at the position, represents time position the change rate of the pollutant concentration at the position, represents time position the acceleration of the pollutant concentration at the position, represents time the meteorological data matrix at the position;

[0024] S23. Use the improved Informer network for pollutant time series prediction, perform a linear transformation on the input matrix and map it to a high-dimensional embedding vector. The improved Informer network includes a data embedding layer, a sparse self-attention mechanism, and a decoder:

[0025] ;

[0026] Among them, represents the high-dimensional embedding vector, represents the input matrix, represents the embedding layer weight matrix, represents the embedding layer bias term;

[0027] S24. Use the sparse self-attention mechanism to calculate the high-dimensional embedding vector, obtain query, key, and value vectors, and calculate the attention scores:

[0028] ;

[0029] Among them, represents the attention score, represents the query vector, represents the key vector, represents the dimension of the key vector, represents the indicator function, which is only retained when the attention score is greater than the threshold and represents the threshold;

[0030] S25. Use the sampling strategy to only retain the attention scores with the greatest influence and calculate the attention-weighted high-dimensional embedding vector:

[0031] ;

[0032] Among them, represents the high-dimensional embedding vector after attention weighting, represents the filtered attention scores, only retaining the attention scores corresponding to the first important time steps, represents the residual connection coefficient, represents the value vector, represents the sampling strategy, only retaining the attention scores with the greatest influence;

[0033] S26. The decoder based on the improved Informer network decodes the high-dimensional embedding vector after attention weighting through reverse calculation to generate the pollutant concentration prediction. The decoder includes a multi-head attention mechanism, a feed-forward network, a residual connection, and a normalization layer. The feed-forward network is a multi-layer perceptron:

[0034] ;

[0035] Among them, represents the predicted pollutant concentration sequence, represents time position where the predicted value of the pollutant concentration is located, represents the prediction step;

[0036] S27. Combine the predicted pollutant concentration and the spatial interpolation method to calculate the spatial diffusion trend of the pollutant, and define the pollutant spatial diffusion equation:

[0037] ;

[0038] Among them, represents the temporal change of the pollutant concentration, represents the convective diffusion term of the pollutant under the influence of the wind speed , represents the diffusion coefficient, represents the wind speed, represents the Laplace diffusion term of the pollutant concentration, represents the pollutant concentration;

[0039] S28. Based on the predicted pollutant concentration and the pollutant spatial diffusion equation, calculate the concentration distribution of the future pollutant in different regions to obtain the predicted value of the pollutant spatio-temporal distribution, and generate the pollutant spatio-temporal distribution prediction data:

[0040] ;

[0041] Among them, represents the predicted value of the pollutant concentration at position . represents the time position where the pollutant concentration is represents the time position where the rate of change of the pollutant concentration is represents the time step, represents the spatial diffusion gradient of the pollutant, represents the spatial step, represents the diffusion coefficient, represents the diffusion term, represents the Laplace diffusion term of the pollutant concentration.

[0042] Optionally, the S4 specifically includes:

[0043] S41. Based on the pollution source control task dataset, establish a reinforcement learning environment and construct a Markov decision process:

[0044] ;

[0045] Among them, represents the Markov decision process of the pollution source control task, represents the set of pollution source agents, represents the state space, including pollutant concentration, pollution change rate, and environmental carrying capacity, represents the pollution source agent whose action space is the pollution adjustment strategy, represents the state transition probability, describing the impact of different adjustment strategies taken by the pollution source agent on the pollutant concentration, represents the reward function, measuring the effect of pollution control, represents the discount factor, controlling the weight of future rewards;

[0046] S42. Set the state space and action space of the pollution source agent. The state space includes pollutant concentration, pollutant time change rate, pollutant acceleration, environmental carrying capacity, and pollution source control weight. The action space selects to adjust the emission amount;

[0047] S43. Construct the reward function of the pollution source agent. The reward function considers pollutant concentration, pollutant time change rate, and pollution source control weight:

[0048] ;

[0049] Among them, represents the revenue function of the pollution source agent, , and respectively represent the influence weights of pollutant concentration, pollutant time change rate, and pollution source regulation weight, represents the pollutant concentration, represents the pollutant time change rate, represents the pollution source regulation weight;

[0050] S44. Generate a pollution source regulation strategy based on reinforcement learning, optimize the pollution source agent regulation strategy using the policy gradient method, and calculate the regulation policy gradient:

[0051] ;

[0052] Among them, represents the regulation policy gradient of the policy objective function of the pollution source agent, represents the policy objective function of the pollution source agent, represents the mathematical expectation, represents the revenue function of the pollution source agent, represents the policy function, represents the parameters of the policy function, represents the gradient calculation, represents the action space of the pollution source agent, represents the state space of the pollution source agent, represents the time step;

[0053] S45. Optimize the pollution source agent regulation strategy using Nash equilibrium, and calculate the optimal strategy of each pollution source agent when the strategies of other agents are fixed:

[0054] ;

[0055] Among them, represents the pollution source agent 's optimal strategy, represents the policy function, represents the optimal strategies of all other agents, represents the revenue function of the pollution source agent, represents the discount factor;

[0056] S46. Optimize the reward mechanism based on pollution historical data, and construct a time-weighted reward function:

[0057] ;

[0058] Among them, represents the time-weighted reward function, represents the historical reward impact factor, represents the time decay factor, which controls the weight of historical pollution data, represents the variable for the accumulation operation, represents the natural exponential function;

[0059] S47. Combine the pollution control resource constraints, dynamically adjust the pollution source control strategy, and optimize the pollution source control strategy using the gradient descent method:

[0060] ;

[0061] Among them, represents the adjusted pollution source control strategy, represents the pollution source control strategy, represents the learning rate, which controls the step size of the control strategy adjustment, represents the gradient of the time-weighted reward function with respect to the control strategy;

[0062] S48. Generate the global optimal pollution source control strategy based on the intelligent agent strategy after reinforcement learning training:

[0063] ;

[0064] Among them, represents the set of global optimal pollution source control strategies for all pollution source intelligent agents, represents the number of pollution source intelligent agents.

[0065] Optionally, the S5 specifically includes:

[0066] S51. Generate a pollution source adjustment vector based on the optimal strategies of each pollution source intelligent agent in the global optimal control strategy. The pollution source adjustment vector includes the emission adjustment amplitude, the start and end times of the control, and the regional execution weight;

[0067] S52. Construct an optimization objective function for the pollution source control parameters. The objective function aims to minimize the pollutant concentration prediction error, reduce the pollution diffusion intensity, and improve the resource utilization efficiency;

[0068] S53. Use the Lagrange multiplier method to calculate the optimal solution of the pollution source control parameters;

[0069] S54. Generate a pollution source adjustment plan based on the calculated optimal solution, and determine the emission limit value, the execution time period, and the regional control priority for each pollution source.

[0070] The beneficial effects of the present invention are:

[0071] First, the improved Informer network proposed by the present invention can effectively overcome the limitations of traditional time series prediction methods. By using the sparse self-attention mechanism, it reduces the computational complexity and enables stronger modeling capabilities when dealing with long-time series pollutant concentration prediction tasks. In addition, by constructing a pollutant time series data matrix and calculating the change rate and acceleration of pollutant concentration, the present invention can more comprehensively consider the time dynamic characteristics of pollutants when predicting future pollutant concentration trends. Combining with the spatial interpolation method, the present invention further makes up for the deficiencies of existing models in pollutant spatial diffusion modeling, making the prediction of pollutant diffusion trends in different regions more accurate and providing reliable data support for subsequent pollution source regulation.

[0072] Second, in terms of pollution source regulation, the present invention adopts the game theory multi-agent reinforcement learning method, which has stronger self-adaptability and global optimization capabilities compared with traditional rule-based pollution control strategies. By constructing the state space, action space and reward function of the pollution source agent, the present invention enables the agent to autonomously learn in a complex pollution environment and uses the Nash equilibrium to optimize the pollution source regulation strategy, enabling each agent to optimize its own benefits without harming the overall environmental interests. Different from the existing single-agent reinforcement learning method, the present invention is based on the game theory framework, enabling different pollution source agents to cooperate and optimize the regulation strategy, avoiding the local optimum problem and improving the overall effect of pollution source regulation. In addition, the present invention combines historical pollution data to optimize the reward mechanism, enabling the agent to better adapt to different pollution evolution patterns and improving the stability and generalization ability of the pollution regulation strategy.

[0073] Finally, the present invention also introduces pollution control resource constraints during the pollution source regulation process, making the finally generated pollution regulation strategy theoretically optimal. At the same time, the regulation strategy of the present invention can dynamically adapt to environmental changes. For example, it strengthens pollution source control under adverse meteorological conditions (such as high humidity and low wind speed) and appropriately relaxes restrictions when the air self-purification ability is strong, thereby maximizing the economic and environmental benefits of pollution control.

[0074] In summary, the present invention improves the accuracy of pollutant concentration prediction by introducing an improved Informer network and optimizes the pollution source regulation strategy using game theory multi-agent reinforcement learning, overcoming the problems of insufficient pollution prediction accuracy and weak self-adaptability of pollution regulation in the prior art. The present invention improves the spatio-temporal resolution of pollutant concentration prediction and, through continuous optimization of the reinforcement learning strategy, enables the pollution source regulation plan to dynamically adapt to environmental changes, realizing the intelligentization, autonomy and global optimization of pollution control. Description of the Drawings

[0075] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:

[0076] Figure 1 is a flowchart of a method for dynamically regulating the distribution of urban pollution sources based on reinforcement learning proposed by the present invention;

[0077] Figure 2 is a schematic diagram of an improved Informer network structure for a method for dynamically regulating the distribution of urban pollution sources based on reinforcement learning proposed by the present invention. Detailed implementation manners

[0078] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0079] Refer to Figure 1 and Figure 2 , a method for dynamically regulating the distribution of urban pollution sources based on reinforcement learning, includes the following steps:

[0080] S1. Collect source data related to air pollution, and perform data cleaning, format conversion, and standardization processing on the source data to construct a pollution data set;

[0081] S2. Use the improved Informer network to perform time series prediction on the pollution data set, calculate the future concentration change trends of different pollutants, and combine the spatial interpolation method to calculate the diffusion of pollutants in different regions to generate pollutant spatio-temporal distribution prediction data;

[0082] S3. Based on the pollutant spatio-temporal distribution prediction data, construct a pollution source impact factor matrix, calculate the impact degree of different pollution sources on air quality, and combine the pollution source emission amount and pollution diffusion characteristics to construct a pollution source regulation task data set;

[0083] S4. Based on the pollution source regulation task data set, establish a game theory multi-agent reinforcement learning environment, set the state space, action space, and reward function of the pollution source agent, use the Nash equilibrium to optimize the regulation strategy of the pollution source agent, optimize the reward mechanism based on the pollution history data, and combine the pollution control resource constraint conditions to dynamically adjust the pollution source regulation strategy to generate a globally optimal regulation strategy;

[0084] S5. Combine the globally optimal regulation strategy, calculate the pollution source adjustment plan, and optimize the pollution control parameters;

[0085] S6. Execute the pollution control strategy and iteratively optimize the pollution control plan.

[0086] In this embodiment, the source data includes air pollutant concentration data and meteorological data.

[0087] In this embodiment, S2 specifically includes:

[0088] S21. Construct a pollutant time series data matrix, process the time series data, and calculate the change rate and acceleration of the pollutant concentration:

[0089] ;

[0090] ;

[0091] Among them, represents the pollutant time series data matrix, represents time position of the pollutant concentration at, represents the total number of time steps, represents the geographical location index, represents time position of the change rate of the pollutant concentration at, represents time position of the acceleration of the pollutant concentration at, represents time position of the pollutant concentration at, represents time position of the pollutant concentration at, represents the time interval between adjacent time steps, represents time position of the change rate of the pollutant concentration at;

[0092] S22. Set the input sequence length to , and define the input matrix:

[0093] ;

[0094] Among them, represents the input matrix, describing the input pollutant time series data, represents the input vector, represents time position of the pollutant concentration at, represents time position of the change rate of the pollutant concentration at, Indicates time Position The acceleration of the pollutant concentration at the Indicates time The meteorological data matrix at the

[0095] S23. Use the improved Informer network for pollutant time series prediction. Perform a linear transformation on the input matrix and map it to a high-dimensional embedding vector. The improved Informer network includes a data embedding layer, a sparse self-attention mechanism, and a decoder:

[0096] ;

[0097] Among them, Represents the high-dimensional embedding vector, Represents the input matrix, Represents the embedding layer weight matrix, Represents the embedding layer bias term;

[0098] S24. Use the sparse self-attention mechanism to calculate the high-dimensional embedding vector to obtain query, key, and value vectors, and calculate the attention scores:

[0099] ;

[0100] Among them, Represents the attention score, Represents the query vector, Represents the key vector, Represents the dimension of the key vector, Represents the indicator function, which is only retained when the attention score is greater than the threshold and Represents the threshold;

[0101] S25. Use the sampling strategy to only retain the attention scores with the greatest influence and calculate the attention-weighted high-dimensional embedding vector:

[0102] ;

[0103] Among them, Represents the attention-weighted high-dimensional embedding vector, Represents the filtered attention scores, only retaining the attention scores corresponding to the first important time steps, Represents the residual connection coefficient, Represents the value vector, Represents the sampling strategy, only retaining the attention scores with the greatest influence;

[0104] S26. The decoder based on the improved Informer network decodes the high-dimensional embedded vector after attention weighting through reverse calculation to generate the pollutant concentration prediction. The decoder includes a multi-head attention mechanism, a feed-forward network, a residual connection, and a normalization layer. The feed-forward network is a multi-layer perceptron:

[0105] ;

[0106] where represents the predicted pollutant concentration sequence, represents time position the predicted value of the pollutant concentration at, represents the prediction step;

[0107] S27. Combine the predicted pollutant concentration prediction and the spatial interpolation method to calculate the spatial diffusion trend of the pollutant, and define the pollutant spatial diffusion equation:

[0108] ;

[0109] where represents the time change of the pollutant concentration, represents the wind speed the convective diffusion term of the pollutant affected by, represents the diffusion coefficient, represents the wind speed, represents the Laplace diffusion term of the pollutant concentration, represents the pollutant concentration;

[0110] S28. Based on the predicted pollutant concentration and the pollutant spatial diffusion equation, calculate the concentration distribution of the future pollutant in different regions to obtain the predicted value of the pollutant spatio-temporal distribution, and generate the pollutant spatio-temporal distribution prediction data:

[0111] ;

[0112] where represents time position the predicted value of the pollutant concentration at, represents time position the pollutant concentration at, represents time position the change rate of the pollutant concentration at, represents the time step, Represents the spatial diffusion gradient of pollutants, Represents the spatial step size, Represents the diffusion coefficient, Represents the diffusion term, Represents the Laplace diffusion term of the pollutant concentration.

[0113] In this embodiment, S4 specifically includes:

[0114] S41. Based on the pollution source regulation task dataset, establish a reinforcement learning environment and construct a Markov decision process:

[0115] ;

[0116] Among them, Represents the Markov decision process of the pollution source regulation task, Represents the set of pollution source agents, Represents the state space, including pollutant concentration, pollution change rate, and environmental carrying capacity, Represents the pollution source agent 's action space, which is the pollution adjustment strategy, Represents the state transition probability, describing the impact of different adjustment strategies taken by the pollution source agent on the pollutant concentration, Represents the reward function, measuring the effect of pollution control, Represents the discount factor, controlling the weight of future rewards;

[0117] S42. Set the state space and action space of the pollution source agent. The state space includes pollutant concentration, pollutant time change rate, pollutant acceleration, environmental carrying capacity, and pollution source regulation weight. The action space selects to adjust the emission amount;

[0118] S43. Construct the reward function of the pollution source agent. The reward function considers pollutant concentration, pollutant time change rate, and pollution source regulation weight:

[0119] ;

[0120] Among them, Represents the reward function of the pollution source agent, , and respectively represent the influence weights of pollutant concentration, pollutant time change rate, and pollution source regulation weight, Represents the pollutant concentration, Represents the pollutant time change rate, Represents the pollution source regulation weight;

[0121] S44. Generate pollution source regulation strategies based on reinforcement learning, optimize the regulation strategies of pollution source agents using the policy gradient method, and calculate the regulation policy gradient:

[0122] ;

[0123] Among them, represents the regulation policy gradient of the policy objective function of the pollution source agent, represents the policy objective function of the pollution source agent, represents the mathematical expectation, represents the reward function of the pollution source agent, represents the policy function, represents the parameters of the policy function, represents the gradient calculation, represents the action space of the pollution source agent, represents the state space of the pollution source agent, represents the time step;

[0124] S45. Optimize the regulation strategies of pollution source agents using Nash equilibrium. Calculate the optimal strategy for each pollution source agent under the condition that the strategies of other agents are fixed:

[0125] ;

[0126] Among them, represents the pollution source agent 's optimal strategy, represents the policy function, represents the optimal strategies of all other agents, represents the reward function of the pollution source agent, represents the discount factor;

[0127] S46. Optimize the reward mechanism based on historical pollution data and construct a time-weighted reward function:

[0128] ;

[0129] Among them, represents the time-weighted reward function, represents the historical reward influence factor, represents the time decay factor, which controls the weight of historical pollution data, represents the variable for the accumulation operation, represents the natural exponential function;

[0130] S47. Combine the resource constraints of pollution control, dynamically adjust the pollution source regulation strategies, and optimize the pollution source regulation strategies using the gradient descent method:

[0131] ;

[0132] Among them, represents the regulation strategy of the adjusted pollution source, represents the regulation strategy of the pollution source, represents the learning rate, which controls the step size of the regulation strategy adjustment, represents the gradient of the time-weighted reward function with respect to the regulation strategy;

[0133] S48. Generate the global optimal regulation strategy for the pollution source based on the intelligent agent strategy after reinforcement learning training:

[0134] ;

[0135] Among them, represents the set of global optimal regulation strategies for all pollution source intelligent agents, represents the number of pollution source intelligent agents.

[0136] In this embodiment, S5 specifically includes:

[0137] S51. Generate a pollution source adjustment vector based on the optimal strategies of each pollution source intelligent agent in the global optimal regulation strategy. The pollution source adjustment vector includes the emission adjustment amplitude, the start and end times of regulation, and the regional execution weight;

[0138] S52. Construct an optimization objective function for the pollution source regulation parameters. The objective function aims to minimize the pollutant concentration prediction error, reduce the pollution diffusion intensity, and improve the resource utilization efficiency;

[0139] S53. Use the Lagrange multiplier method to calculate the optimal solution of the pollution source regulation parameters;

[0140] S54. Generate a pollution source adjustment plan according to the calculated optimal solution, and determine the emission limit value, the execution time period, and the regional control priority of each pollution source.

[0141] Example 1:

[0142] To verify the feasibility of the present invention in implementation, the present invention is applied to a typical urban pollution source regulation scenario, and a city with a relatively high degree of industrialization is selected.

[0143] In the air pollution monitoring stations in this city, a plurality of sensors are arranged to collect pollutant concentration, meteorological data, and pollution source emission data. Among them, PM2.5, PM10, , , CO, The concentration data of major pollutants such as

[0144] are collected by ground monitoring stations and updated hourly; meteorological data including wind speed, wind direction, temperature, humidity, etc. are provided by local meteorological stations; pollution source emission data are obtained through channels such as industrial enterprise online monitoring systems and traffic flow monitoring systems. After data collection, the data are first cleaned, format-converted, and standardized to construct a pollution dataset, which is then input into an improved Informer network for spatio-temporal prediction of pollutant concentrations.

[0145] In practical applications, the present invention uses an improved Informer network to predict the concentrations of major pollutants in the city, combines spatial interpolation methods to calculate the diffusion of pollutants in different regions. In terms of pollution source regulation, based on the pollution source regulation task dataset, a game theory multi-agent reinforcement learning environment is established, and the state space, action space, and reward function of the pollution source agent are set. In the experiment, the agents respectively represent different pollution sources in the city, including industrial enterprises, motor vehicles, construction dust, etc. The reinforcement learning algorithm is used to optimize the pollution source emission strategy. Through interaction and learning with the environment, the agents gradually reduce pollutant emissions and optimize the overall air quality.

[0146] Table 1 Experimental comparison data table

[0147]

[0148] It can be seen from the above experimental data comparison table that the present invention shows significant advantages in both pollutant prediction and pollution source regulation. First, in terms of pollutant concentration prediction, after using the improved Informer network, the prediction errors (MSE) of PM2.5, and are all significantly reduced. Among them, the prediction error of PM2.5 has decreased by 36.7% compared with the traditional method, the prediction error has decreased by 34.9%, the prediction error has decreased by 34.2%. This result shows that compared with the traditional time series prediction method, the Informer network of the present invention improves the modeling ability of long-time series data and optimizes the prediction accuracy of pollutant concentrations through the sparse self-attention mechanism and sampling strategy.

[0149] In terms of pollution source regulation, the present invention uses game theory multi-agent reinforcement learning to optimize pollution emissions, resulting in a significant increase in the decline rate of pollutant concentrations. Compared with the traditional pollution control method, the decline rate of PM2.5 has increased from 9.8% to 22.7%, with an increase of up to 131.6%; The decline rate increased from 7.5% to 17.4%, with an increase of 132.0%. The decline rate increased from 5.9% to 13.9%, with an increase of 135.6%. This indicates that in the process of pollutant treatment, the reinforcement learning method can effectively optimize the emission strategy of pollution sources, making pollution treatment more accurate, intelligent, and capable of adapting to different meteorological conditions and pollution evolution trends.

[0150] In terms of the treatment effect of heavy pollution weather, when using traditional pollution control methods, the decline rate of PM2.5 was 9.8%, while the intelligent regulation strategy of the present invention made the decline rate of PM2.5 reach 19.3%, an increase of 96.9%. This shows that in extreme pollution weather, traditional methods may be difficult to achieve effective pollution control, while the present invention optimizes the pollution source emission strategy through reinforcement learning, making the pollution treatment plan have stronger dynamic adaptability and global optimality, and can further reduce pollutant concentrations and improve air quality during critical periods.

[0151] Generally speaking, the present invention accurately predicts pollutant concentrations through the improved Informer network and combines game theory multi-agent reinforcement learning to optimize the pollution source emission strategy, effectively overcoming the problems of low accuracy of traditional pollution prediction and poor adaptive ability of pollution regulation, and achieving more accurate pollution prediction and more efficient pollution treatment strategies.

[0152] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. A method for dynamically regulating the distribution of urban pollution sources based on reinforcement learning, characterized in that, It includes the following steps: S1. Collect the source data of air pollution, perform data cleaning, format conversion, and standardization processing on the source data, and construct a pollution dataset; S2. Use the improved Informer network to perform time series prediction on the pollution dataset, calculate the future concentration change trends of different pollutants, and combine the spatial interpolation method to calculate the diffusion of pollutants in different regions, and generate pollutant spatio-temporal distribution prediction data; S3. Based on the pollutant spatio-temporal distribution prediction data, construct a pollution source impact factor matrix, calculate the impact degree of different pollution sources on air quality, and combine the pollution source emissions and pollution diffusion characteristics to construct a pollution source regulation task dataset; S4. Based on the pollution source regulation task dataset, establish a game theory multi-agent reinforcement learning environment, set the state space, action space, and reward function of the pollution source agent, use the Nash equilibrium to optimize the regulation strategy of the pollution source agent, optimize the reward mechanism based on the pollution historical data, and combine the pollution control resource constraint conditions to dynamically adjust the pollution source regulation strategy to generate a globally optimal regulation strategy; S5. Combine the globally optimal regulation strategy, calculate the pollution source adjustment plan, and optimize the pollution control parameters; S6. Execute the pollution control strategy and iteratively optimize the pollution control strategy; The improved Informer network in S2 adopts a sparse self-attention mechanism and a Top-u sampling strategy; The specific content of S5 includes: S51. Based on the optimal strategies of each pollution source agent in the globally optimal regulation strategy, generate a pollution source adjustment vector, and the pollution source adjustment vector includes the emission adjustment amplitude, the start and end time of regulation, and the regional execution weight; S52. Construct an optimization objective function for pollution source regulation parameters, and the objective function takes minimizing the pollutant concentration prediction error, reducing the pollution diffusion intensity, and improving the resource utilization efficiency as the optimization objectives; S53. Use the Lagrange multiplier method to calculate the optimal solution of the pollution source regulation parameters; S54. According to the calculated optimal solution, generate a pollution source adjustment plan, and determine the emission limit value, the execution time period, and the regional control priority of each pollution source.

2. The dynamic distribution regulation method of urban pollution sources based on reinforcement learning according to claim 1, wherein, The source data includes air pollutant concentration data and meteorological data.

3. The dynamic distribution regulation method of urban pollution sources based on reinforcement learning according to claim 1, characterized in that The specific content of S2 includes: S21. Construct a pollutant time series data matrix, process the time series data, and calculate the change rate and acceleration of the pollutant concentration: ; ; Among them, represents the pollutant time series data matrix, represents time position the pollutant concentration at the position, represents the total number of time steps, represents the geographical location index, represents time position the pollutant concentration change rate at the position, represents time position the acceleration of the pollutant concentration at the position, represents time position the pollutant concentration at the position, represents time position the pollutant concentration at the position, represents the time interval between adjacent time steps, represents time position the pollutant concentration change rate at the position; S22. Set the input sequence length to , and define the input matrix: ; Among them, represents the input matrix, describing the input pollutant time series data, represents the input vector, represents time position the pollutant concentration at the position, represents time position the rate of change of the pollutant concentration at the position, represents time position the acceleration of the pollutant concentration at the position, represents time the meteorological data matrix at the position; S23. An improved Informer network is adopted for time series prediction of pollutants, and a linear transformation is performed on the input matrix to map it to a high-dimensional embedding vector. The improved Informer network includes a data embedding layer, a sparse self-attention mechanism, and a decoder: ; Among them, represents the high-dimensional embedding vector, represents the input matrix, represents the weight matrix of the embedding layer, represents the bias term of the embedding layer; S24. Use the sparse self-attention mechanism to calculate the high-dimensional embedding vectors to obtain query, key, and value vectors, and calculate the attention scores: ; Among them, represents the attention score, represents the query vector, represents the key vector, represents the dimension of the key vector, represents the indicator function, which is retained only when the attention score is greater than the threshold ; represents the threshold; S25. Adopt the Top-u sampling strategy to only retain the attention scores with the greatest influence and calculate the attention-weighted high-dimensional embedding vectors: ; Among them, represents the high-dimensional embedding vector after attention weighting, represents the filtered attention scores, only retaining the attention scores corresponding to the first u important time steps, represents the residual connection coefficient, represents the value vector, represents the sampling strategy, only retaining the attention scores with the greatest influence; S26. Based on the decoder of the improved Informer network, decode the high-dimensional embedded vector after attention weighting through reverse calculation to generate the pollutant concentration prediction. The decoder includes a multi-head attention mechanism, a feed-forward network, a residual connection, and a normalization layer. The feed-forward network is a multi-layer perceptron: ; Among them, represents the predicted pollutant concentration sequence, represents time position the predicted value of the pollutant concentration at the position, represents the prediction step; S27. Prediction in combination with predicted pollutant concentration and the spatial interpolation method to calculate the spatial diffusion trend of pollutants, and define the spatial diffusion equation of pollutants: ; Among them, represents the time variation of the pollutant concentration, represents the wind speed and the convective diffusion term of the pollutant under the influence, represents the diffusion coefficient, represents the wind speed, represents the Laplace diffusion term of the pollutant concentration, represents the pollutant concentration; S28. Based on the predicted pollutant concentration and the pollutant spatial diffusion equation, calculate the concentration distribution of future pollutants in different regions, obtain the predicted values of pollutant spatio-temporal distribution, and generate pollutant spatio-temporal distribution prediction data: ; Among them, represents the predicted value of the pollutant concentration at the position at time represents the pollutant concentration at the position at time represents the rate of change of the pollutant concentration at the position at time represents the time step, represents the spatial diffusion gradient of the pollutant, represents the spatial step, represents the diffusion coefficient, represents the diffusion term, represents the Laplace diffusion term of the pollutant concentration.

4. A method for dynamically distributing and regulating urban pollution sources based on reinforcement learning according to claim 1, characterized in that, The specific content of S4 includes: S41. Based on the pollution source regulation task dataset, establish a reinforcement learning environment and construct a Markov decision process; ; Among them, represents the Markov decision process of the pollution source regulation task, represents the set of pollution source agents, represents the state space, including pollutant concentration, pollution change rate, and environmental carrying capacity, represents the pollution source agent 's action space, which is the pollution adjustment strategy, represents the state transition probability, describing the impact of different adjustment strategies adopted by the pollution source agent on the pollutant concentration, represents the reward function, measuring the effect of pollution control, represents the discount factor, controlling the weight of future rewards; S42. Set the state space and action space of the pollution source agent, and the state space includes the pollutant concentration, the pollutant time change rate, the pollutant acceleration, the environmental carrying capacity, and the pollution source regulation weight, and the action space selects to adjust the emission amount; S43. Construct a reward function for the pollution source agent, and the reward function considers the pollutant concentration, the pollutant time change rate, and the pollution source regulation weight: ; Among them, represents the revenue function of the pollution source agent, , and respectively represent the influence weights of pollutant concentration, pollutant time change rate, and pollution source regulation weight, represents the pollutant concentration, represents the pollutant time change rate, represents the pollution source regulation weight; S44. Generate pollution source regulation strategies based on reinforcement learning, optimize the regulation strategies of pollution source agents using policy gradient methods, and calculate the regulation policy gradients: ; Among them, represents the regulatory policy gradient of the policy objective function of the pollution source agent, represents the policy objective function of the pollution source agent, represents the mathematical expectation, represents the revenue function of the pollution source agent, represents the policy function, represents the parameters of the policy function, represents the gradient calculation, represents the action space of the pollution source agent, represents the state space of the pollution source agent, represents the time step; S45. Optimize the regulation strategies of pollution source agents using Nash equilibrium. Calculate the optimal strategies for each pollution source agent while keeping the strategies of other agents fixed: ; Among them, represents the optimal strategy of the pollution source agent ; represents the policy function represents the optimal strategies of all other agents represents the revenue function of the pollution source agent represents the discount factor; S46. Optimize the reward mechanism based on historical pollution data and construct a time-weighted reward function: ; Among them, represents the time-weighted reward function, represents the historical reward impact factor, represents the time decay factor, which controls the weight of historical pollution data, represents the variable for the accumulation operation, represents the natural exponential function; S47. Combine the constraints of pollution control resources, dynamically adjust the pollution source regulation strategies, and optimize the pollution source regulation strategies using the gradient descent method: ; Among them, represents the regulation strategy of the adjusted pollution source, represents the regulation strategy of the pollution source, represents the learning rate, which controls the step size of the regulation strategy adjustment, represents the gradient of the time-weighted reward function with respect to the regulation strategy; S48. Generate the globally optimal pollution source regulation strategy based on the agent strategies after reinforcement learning training: ; Among them, represents the global optimal regulation strategy set of all pollution source agents, represents the number of pollution source agents.

Citation Information

Patent Citations

  • Coal mining intelligent scheduling method based on artificial intelligence platform

    CN119090243A