Major pollutant discharge right quota dynamic optimization deep reinforcement learning system and method
Through deep reinforcement learning technology, an intelligent dynamic optimization system for pollutant emission rights quota has been built, which solves the shortcomings of traditional methods in responding to environmental changes and balancing environmental and economic needs, and achieves more efficient and flexible environmental management and resource utilization.
Patent Information
- Application Number
- CN202510172495.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional pollution discharge rights quota management method lacks flexibility, cannot respond to environmental changes in a timely manner, is difficult to balance environmental protection and economic development needs, and lacks the comprehensive understanding and precise modeling ability of complex environmental systems, is difficult to deal with complex scenarios with multiple pollutants and multiple subjects, and lacks the ability to predict future trends and adaptive adjustment mechanisms.
The dynamic optimization system for the main pollutant emission rights quota based on deep reinforcement learning is adopted, including the database management module, environmental status assessment module, multi-agent reinforcement learning module and strategy evaluation module. Through deep neural network and multi-agent reinforcement learning algorithm, the intelligent and dynamic management of pollution rights quota is achieved.
Real-time response to environmental changes, dynamically adjust the pollutant emission rights quota, improve the flexibility and effectiveness of environmental management, can balance environmental protection and economic development needs, accurately model complex environmental systems, handle multiple pollutants and multi-subject scenarios, and have prediction and adaptability.
Smart Images

Figure CN119990668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental science and technology, and in particular to a deep reinforcement learning system and method for dynamic optimization of emission rights quotas for major pollutants. Background Art
[0002] The traditional method of pollution emission quota management mainly relies on fixed quota allocation and simple market trading mechanism, which has exposed many problems in practice.
[0003] First, the fixed quota allocation method lacks flexibility and cannot respond to dynamic changes in environmental conditions in a timely manner. The environmental system is a complex dynamic system that is affected by many factors, such as seasonal changes, climate conditions, economic development, etc. Fixed quota allocation cannot reflect these dynamic changes, which may lead to overly conservative allocation in some periods, which inhibits economic development, and overly loose allocation in other periods, which endangers environmental safety.
[0004] Secondly, a simple market trading mechanism cannot fully consider the balance between environmental carrying capacity and economic development needs. Relying solely on market regulation may lead to excessive concentration of pollution discharge rights in certain regions or industries, causing excessive local environmental pressure. At the same time, the market mechanism may also be manipulated by a small number of dominant enterprises, affecting the fairness and efficiency of resource allocation.
[0005] Furthermore, the existing emission quota management system often lacks comprehensive understanding and accurate modeling capabilities of complex environmental systems, which results in the inability to fully consider the interactions between various environmental factors and the long-term impact of emission behavior on the environment when formulating quota strategies.
[0006] In addition, traditional methods are unable to cope with complex scenarios involving multiple pollutants and multiple entities. In actual situations, there are often complex situations where multiple pollutants are discharged at the same time and multiple pollutant discharge entities affect each other. Simple linear models or static optimization methods are difficult to effectively handle this complexity.
[0007] Finally, existing technologies lack the ability to predict future trends and adaptive adjustment mechanisms. The formulation and implementation of environmental policies need to consider long-term effects, but traditional methods are mostly limited to responding to current conditions and it is difficult to foresee the long-term impacts that policy adjustments may have.
[0008] In view of the above problems, there is an urgent need for an intelligent system that can dynamically optimize the quota of emission rights for major pollutants. The system should be able to respond to environmental changes in real time, balance environmental protection and economic development needs, accurately model complex environmental systems, handle multi-pollutant and multi-agent scenarios, and have predictive and adaptive capabilities. Summary of the invention
[0009] The present invention aims at the above technical problems and proposes a system and method for dynamic optimization of emission quotas of major pollutants based on deep reinforcement learning. The invention realizes intelligent and dynamic management of emission quotas by introducing advanced artificial intelligence technology, especially deep reinforcement learning algorithm.
[0010] The present invention proposes a deep reinforcement learning system for dynamic optimization of the quota of emission rights for major pollutants, including:
[0011] Database management module for:
[0012] Store water quality information and sewage discharge information;
[0013] Update and maintain emission quota data;
[0014] The environmental status assessment module is in communication with the database management module and is used to:
[0015] Based on the water quality information of the water body, assessing water quality concentration;
[0016] Calculate the load of various major pollutants in the water body according to the sewage discharge information;
[0017] A multi-agent reinforcement learning module is communicatively connected to the environment state assessment module and is used to:
[0018] Receiving environmental status information and pollutant load information sent by the environmental status assessment module;
[0019] Based on the environmental state information and the sewage load information, executing a multi-agent reinforcement learning algorithm;
[0020] Generate an optimized emission quota allocation plan;
[0021] A strategy evaluation module is communicatively connected with the multi-agent reinforcement learning module and the database management module, and is used to:
[0022] Receiving the emission rights quota allocation plan generated by the multi-agent reinforcement learning module;
[0023] Evaluate the overall effect of the emission quota allocation plan;
[0024] Feeding back the evaluation results to the multi-agent reinforcement learning module for strategy adjustment;
[0025] The optimal quota allocation scheme is stored in the database management module.
[0026] Preferably, the environmental status assessment module comprises:
[0027] Water quality concentration assessment submodule, used to:
[0028] Receiving water quality information of water bodies sent by the database management module;
[0029] Based on the water quality information of the water body, calculating the concentration of each pollutant in the water body;
[0030] Generate water quality concentration assessment reports;
[0031] The sewage load assessment submodule is used to:
[0032] Receiving the sewage discharge information sent by the database management module;
[0033] Based on the sewage discharge information, calculating the load of each major pollutant in the water body;
[0034] Generate sewage load assessment reports.
[0035] Preferably, the multi-agent reinforcement learning module includes a plurality of single-agent sub-modules, each of which includes:
[0036] Action space submodule, used to:
[0037] Divide the action space according to the total quota of emission rights currently owned by the agent;
[0038] Generates a set of optional quota adjustment actions;
[0039] The policy network submodule is used to:
[0040] Receive environmental status information;
[0041] Based on deep neural networks, evaluate the policy value of each action in the action space;
[0042] The value network submodule is used to:
[0043] Receive environmental status information;
[0044] Based on deep neural networks, the value of each action is evaluated;
[0045] Dynamic programming submodule for:
[0046] Combine the outputs of the policy network and the value network to estimate the long-term return of each action;
[0047] Select the best action;
[0048] State transfer submodule, used for:
[0049] Execute the selected action;
[0050] Update the environment status;
[0051] Reward network submodule, used to:
[0052] Calculate the immediate reward after performing the action;
[0053] Adjust the policy network and value network based on the reward information.
[0054] Preferably, the action space divided in the action space submodule is:
[0055] {a1, a2, a3, a4}, where
[0056] a1 means the current Agent does not change the quota.
[0057] a2 means the current Agent reduces the quota.
[0058] a3 means the current Agent maintains the quota,
[0059] a4 indicates that the current Agent increases the quota.
[0060] Preferably, the policy value function of the policy network submodule is:
[0061] π(a t |s t ;θ)=P(a t |s t ; θ),
[0062] Among them, s t Indicates the environmental state, a t represents the action, and θ represents the neural network parameters.
[0063] Preferably, the value function of the value network submodule is:
[0064] V(s t )=E[R t |s t ],
[0065] Among them, R t Indicates that from state s t The cumulative discount reward starts.
[0066] Preferably, the strategy evaluation module comprises:
[0067] The return evaluation submodule is used to:
[0068] Receiving the emission rights quota allocation plan generated by the multi-agent reinforcement learning module;
[0069] Calculate the environmental and economic benefits of the plan;
[0070] Generate comprehensive return assessment report;
[0071] Solution optimization submodule, used to:
[0072] Based on the comprehensive return evaluation report, adjust and optimize the emission rights quota allocation plan;
[0073] Generate the final optimal pollution discharge policy.
[0074] As a preference, it also includes:
[0075] The quota trading module is in communication with the database management module and the multi-agent reinforcement learning module and is used to:
[0076] Receiving a quota allocation plan generated by the multi-agent reinforcement learning module;
[0077] Carry out transaction matching of emission rights quota;
[0078] The transaction results are updated to the database management module.
[0079] Preferably, the database management module includes:
[0080] Water quality information database, used to store water quality parameters such as concentration of various pollutants, pH value, dissolved oxygen, etc.
[0081] Pollutant discharge information database, used to store information such as the amount of pollutants discharged, the type of pollutants discharged, and the time of pollutants discharged by each pollutant-discharging enterprise;
[0082] The quota information database is used to store the initial quota, current quota, historical transaction records and other information of each pollutant-discharging enterprise.
[0083] The deep reinforcement learning method for dynamic optimization of the quota of emission rights for major pollutants includes the following steps:
[0084] (1) Store and update water quality information and sewage discharge information through the database management module;
[0085] (2) The environmental status assessment module assesses the water quality concentration based on the water quality information of the water body, and calculates the load of various major pollutants in the water body according to the sewage discharge information;
[0086] (3) The multi-agent reinforcement learning module receives environmental status information and sewage load information and performs the following sub-steps:
[0087] a. Divide the action space for each agent;
[0088] b. Use deep neural networks to build policy networks and value networks;
[0089] c. Based on the current state, select an action through the policy network;
[0090] d. Perform the selected action and observe environmental feedback and rewards;
[0091] e. Update the value network and strategy network;
[0092] f. Repeat steps ce until convergence or the preset number of iterations is reached;
[0093] (4) The strategy evaluation module evaluates the global effect of the emission quota allocation scheme generated by the multi-agent reinforcement learning module;
[0094] (5) Based on the evaluation results, the strategy evaluation module provides feedback to the multi-agent reinforcement learning module for adjusting the learning strategy;
[0095] (6) storing the optimal quota allocation plan in the database management module;
[0096] (7) Repeat steps (1)-(6) regularly to adapt to environmental changes and policy adjustments.
[0097] The beneficial effects of the present invention are embodied in multiple aspects:
[0098] From a macro perspective, the present invention provides a powerful decision-making support tool for environmental management departments. Through real-time data analysis and intelligent algorithms, the system can timely adjust the emission quota allocation strategy according to the dynamic changes in environmental conditions. This dynamic optimization capability significantly improves the flexibility and effectiveness of environmental management, making policy formulation more scientific and reasonable.
[0099] At the system architecture level, the present invention adopts a modular design, including a database management module, an environmental status assessment module, a multi-agent reinforcement learning module, and a strategy assessment module. The collaborative work between these modules realizes the intelligent management of the entire process from data collection, environmental assessment to strategy generation. The close collaboration between modules not only improves the overall efficiency of the system, but also enhances the scalability and adaptability of the system.
[0100] In terms of algorithmic innovation, the deep reinforcement learning technology introduced in this invention provides a new approach to solving decision-making problems in complex environments. The multi-agent reinforcement learning module can simultaneously consider the interests and behaviors of multiple pollutant discharge entities and find the optimal solution through competition and cooperation. This method effectively solves the difficulties of traditional methods in dealing with complex multi-agent scenarios.
[0101] In terms of environmental modeling, the present invention achieves accurate modeling of complex environmental systems through an environmental status assessment module. This module not only considers direct indicators such as water quality concentration, but also calculates derived indicators such as pollutant load, providing comprehensive and accurate environmental status information for subsequent decision optimization.
[0102] In terms of the balance between economic benefits and environmental protection, the strategy evaluation module of the present invention takes into account both environmental benefits and economic benefits through a multi-objective evaluation method. This comprehensive evaluation mechanism ensures that the generated quota allocation plan can protect the environment while minimizing the negative impact on economic development.
[0103] In terms of the adaptability and learning ability of the system, the deep reinforcement learning algorithm of the present invention has the characteristics of continuous learning and self-optimization. As time goes by and data accumulates, the decision-making ability of the system will continue to improve, gradually adapting to the management needs of emission quotas in different regions and types.
[0104] In practical application, the quota trading module of the present invention provides technical support for the market-based trading of emission rights. Through intelligent transaction matching algorithms, the system can maximize the overall transaction efficiency while ensuring the interests of both parties to the transaction, and promote the rational flow and efficient use of emission rights resources.
[0105] In general, this invention effectively solves the problems existing in the traditional emission quota management through deep reinforcement learning technology, modular system design and multi-objective optimization method. It not only improves the scientificity and effectiveness of environmental management, but also provides a new technical means for balancing environmental protection and economic development, which has important theoretical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 It is the main flow chart of the overall system of the present invention;
[0107] Figure 2 is an internal structure diagram of the data acquisition module of the present invention;
[0108] Figure 3 is a work flow chart of the environmental status assessment module of the present invention;
[0109] Figure 4 is a structural diagram of the multi-agent reinforcement learning module of the present invention;
[0110] Figure 5 It is a decision flow chart of a single Agent submodule of the present invention;
[0111] Figure 6 This is a flow chart of the strategy evaluation and quota trading of the present invention. DETAILED DESCRIPTION
[0112] Please refer to Figure 1-6 The present invention provides a deep reinforcement learning system and method for dynamic optimization of emission quotas of major pollutants. The system includes a database management module, an environmental status assessment module, a multi-agent reinforcement learning module, and a strategy assessment module. These modules work together to achieve intelligent dynamic optimization of emission quotas.
[0113] Specifically, the database management module 1 is used to store and manage water quality information and pollution discharge information. This module is not just a simple data storage unit, but an intelligent system that can update and maintain pollution discharge quota data in real time. For example, when new water quality monitoring data or pollution discharge data is input, the database management module 1 can automatically update the relevant records and maintain the consistency and integrity of the data. Preferably, the module can also analyze historical data, identify water quality change trends or pollution discharge behavior patterns, and provide data support for subsequent decision-making.
[0114] The environmental status assessment module 2 establishes a communication connection with the database management module 1. Its main function is to assess the water quality concentration based on the water quality information of the water body and calculate the load of various major pollutants in the water body based on the sewage discharge information. In one embodiment of the present invention, the environmental status assessment module 2 can use a complex water quality model to assess the health of the water body. For example, the comprehensive pollution index can be calculated using the following formula:
[0115] The calculation formula of the comprehensive pollution index is:
[0116]
[0117] Among them, P is the comprehensive pollution index, C i is the measured concentration of the i-th pollutant, S i is the standard concentration of the i-th pollutant, and n is the number of pollutant types. When P>1, it means that the water body is polluted, and the larger the P value, the more serious the pollution. The multi-agent reinforcement learning module 3 is the core of this system, and it establishes a communication connection with the environmental status assessment module 2. The module receives the environmental status information and pollution load information sent by the environmental status assessment module 2, and then executes the multi-agent reinforcement learning algorithm to generate an optimized pollution discharge quota allocation plan. In a preferred embodiment of the present invention, the multi-agent reinforcement learning module 3 adopts a deep Q network (DQN) algorithm. The core of the algorithm is to approximate the Q function through a neural network: Q(s,a;θ)≈Q * (s,a) where s represents the state, a represents the action, and θ represents the parameters of the neural network. The goal of the algorithm is to minimize the following loss function:
[0119] Among them, r represents the immediate reward, γ represents the discount factor, and θ - Represents the parameters of the target network.
[0120] The strategy evaluation module 4 establishes a communication connection with the multi-agent reinforcement learning module 3 and the database management module 1. The main task of this module is to evaluate the global effect of the emission quota allocation scheme generated by the multi-agent reinforcement learning module 3. In one embodiment of the present invention, the strategy evaluation module 4 can adopt a multi-objective evaluation method, taking into account both environmental benefits and economic benefits. For example, the following comprehensive evaluation function can be used:
[0121] E=w1E env +w2E eco ,
[0122] Among them, E is the comprehensive evaluation value, E env is the environmental benefit index, E eco is the economic benefit index, w1 and w2 are weight coefficients, and w1+w2=1. By adjusting the weight coefficients, the needs of environmental protection and economic development can be flexibly balanced.
[0123] The system of the present invention realizes the dynamic optimization of the quota of emission rights for major pollutants through the collaborative work of the above modules. Compared with traditional methods, this system has the following advantages: first, it can respond to environmental changes in real time, adjust the emission rights quota in a timely manner, and improve the flexibility and effectiveness of environmental management; second, through the deep reinforcement learning algorithm, the system can continuously learn and optimize the decision-making strategy, and gradually improve the rationality of quota allocation; finally, the multi-agent framework design enables the system to handle complex multi-agent game problems and better simulate the emission rights trading scenario in the real world.
[0124] In practical applications, this system can adjust parameters and optimize modules according to specific circumstances. For example, in water quality assessment, appropriate evaluation indicators can be selected according to the characteristics of different water bodies; in reinforcement learning algorithms, the structure and learning rate of the neural network can be adjusted according to the complexity of the problem; in strategy evaluation, the weights of each indicator can be adjusted according to local policies and development needs. This flexibility enables this system to adapt to the quota management needs of different regions and types of pollution rights.
[0125] In summary, the deep reinforcement learning system and method for dynamic optimization of the quota of emission rights for major pollutants provided by the present invention provide a powerful decision-making support tool for environmental management departments, which helps to achieve the coordination and unity of environmental protection and economic development.
[0126] In a preferred embodiment of the present invention, the multi-agent reinforcement learning module 3 includes multiple single-agent sub-modules. Each single-agent sub-module represents a pollutant-discharging enterprise, and optimizes the allocation of pollutant emission quotas through mutual cooperation and competition. This design fully considers the complex relationship between pollutant-discharging enterprises in the real world, so that the system can better simulate the actual situation.
[0127] Specifically, each single agent submodule includes an action space submodule 31, a policy network submodule 32, a value network submodule 33, a dynamic programming submodule 34, a state transfer submodule 35, and a reward network submodule 36. These submodules work together to realize the intelligent decision-making process of a single agent.
[0128] The action space submodule 31 is responsible for dividing the action space and generating a set of optional quota adjustment actions according to the total quota of emission rights currently owned by the Agent. In one embodiment of the present invention, the action space is divided into four basic actions: keep the existing quota, reduce the quota, increase the quota and remain unchanged. This design simplifies the decision-making process and retains sufficient flexibility.
[0129] The policy network submodule 32 and the value network submodule 33 are the core components of this system, and together they form the basis of the Actor-Critic architecture. The policy network is responsible for selecting the best action based on the current state of the environment, while the value network evaluates the value of each state. These two networks are implemented through deep neural networks, which can handle high-dimensional state spaces and improve the learning and generalization capabilities of the system.
[0130] The dynamic programming submodule 34 combines the outputs of the policy network and the value network to estimate the long-term reward of each action and select the optimal action. This process can be achieved through algorithms such as Monte Carlo Tree Search (MCTS), further improving the quality and efficiency of decision making.
[0131] The state transfer submodule 35 is responsible for executing the selected action and updating the environmental state. In practical applications, this process involves interaction with other agents and feedback from the environment. For example, when an agent decides to increase its emission quota, it may affect the decisions of other agents and also change the overall environmental state.
[0132] The reward network submodule 36 calculates the immediate reward after executing the action, and adjusts the strategy network and the value network according to the reward information. In a preferred embodiment of the present invention, the reward function not only takes into account the economic benefits, but also includes the factor of environmental protection. For example, the following reward function can be used:
[0133] R=w1·P-w2·E-w3·|QQ t |,
[0134] Among them, R is the total reward, P is the economic benefit, E is the environmental pollution index, Q is t is the current quota, Q tis the target quota, w1, w2 and w3 are weight coefficients. By adjusting these weights, the needs of economic development and environmental protection can be balanced. The action space divided by the system of the present invention in the action space submodule 31 is {a1, a2, a3, a4}. Among them,
[0135] a1 means "the current Agent does not change the quota",
[0136] a2 means "the current agent reduces the quota",
[0137] a3 means "the current Agent maintains the quota",
[0138] a4 means "the current agent increases the quota".
[0139] This design simplifies the decision-making process while retaining sufficient flexibility to deal with different situations. In the policy network submodule 32, the policy value function takes the following form:
[0140] π(a t |s t ;θ)=P(a t |s t ; θ),
[0141] Among them, s t represents the environmental state at time t, a t represents the action at time t, and θ represents the neural network parameters. This function represents the action at a given state s t and parameter θ, select action a t In the value network submodule 33, the value function is defined as:
[0142] V(s t )=E[R t |s t ],
[0143] Among them, R t Indicates that from state s t The cumulative discounted reward at the beginning. This function evaluates the total benefit that may be obtained in the future under a certain state.
[0144] Through this design, the system of the present invention can effectively learn the optimal emission quota allocation strategy. Each agent not only considers its own interests, but also needs to weigh the impact of its behavior on the overall environment, thereby achieving a balance between individual interests and collective interests.
[0145] In practical applications, the performance of the system may be affected by many factors, such as the dimension of the state space, the structure of the neural network, the setting of the learning rate, etc. Therefore, in the specific implementation, it is necessary to adjust and optimize the parameters according to the actual situation. For example, the optimal neural network structure and hyperparameters can be selected through methods such as cross-validation.
[0146] In general, the multi-agent reinforcement learning module 3 provided by the present invention realizes intelligent decision-making in complex environments through its carefully designed sub-module structure. This method can not only adapt to dynamically changing environments, but also find a balance point between multiple goals, providing strong technical support for the dynamic optimization of the quota of emission rights for major pollutants.
[0147] The system of the present invention further includes a quota trading module 5, which establishes a communication connection with the database management module 1 and the multi-agent reinforcement learning module 3. The introduction of the quota trading module 5 enables the system to not only optimize the initial allocation of emission rights quotas, but also support dynamic trading of quotas, further improving the efficiency of resource allocation.
[0148] Specifically, the quota trading module 5 first receives the quota allocation plan generated by the multi-agent reinforcement learning module 3. Based on this initial plan, the quota trading module 5 performs transaction matching of emission rights quotas. In a preferred embodiment of the present invention, the transaction matching process adopts a bilateral matching algorithm, which can maximize the overall transaction efficiency while ensuring the interests of both parties to the transaction.
[0149] The core idea of the transaction matching algorithm is to find complementary buying and selling needs and match them under certain constraints. For example, the following objective function can be used to optimize the matching results:
[0150] max∑ i,j x ij (p i -c j ),
[0151] Among them, x ij represents the transaction volume between buyer i and seller j, p i represents the highest acceptable price of buyer i, c j represents the minimum accepted price of seller j. This optimization problem needs to meet a series of constraints, such as:
[0152] The total transaction volume does not exceed the tradable quota;
[0153] Σ i∈B ∑ j∈S x ij ≤Q total ,
[0154] Among them, xij represents the transaction volume between buyer i and seller j; B is the set of all buyers; S is the set of all sellers; Q total is the total tradable quota of the system.
[0155] The transaction volume of a single enterprise does not exceed its permitted range, etc.
[0156] For Buyeri:
[0157]
[0158] in is the maximum buying quota of buyer i.
[0159] For Seller j:
[0160]
[0161] in is the maximum selling quota of seller j.
[0162] The transaction price must meet the buyer's maximum acceptable price and the seller's minimum acceptable price.
[0163] P min,j ≤p ij ≤P max,i ,
[0164] Among them, p ij is the transaction price between buyer i and seller j; P min,j is the lowest price accepted by seller j; P max,i is the maximum price accepted by buyer i.
[0165] After the transaction, the enterprise's quota balance cannot be lower than its minimum reserved quota, nor can it exceed its maximum permitted quota.
[0166] For Buyeri:
[0167]
[0168] in: is the initial quota of buyer i;
[0169] is the minimum reserved quota of buyer i; is the maximum allowed quota for buyer i.
[0170] For Seller j:
[0171]
[0172] Volume and price must be non-negative values.
[0173]
[0174] The above constraints are combined into the optimization objective function to form a complete optimization model. For example, the optimization objective can be to minimize transaction costs or maximize overall benefits:
[0175] MaximizeZ=Σ i∈B Σ j∈S u ij x ij ,
[0176] where u ij is the transaction utility between buyer i and seller j.
[0177] These constraints will ensure that the solution complies with both market rules and the capabilities and limitations of each participant.
[0178] After completing the transaction matching, the quota trading module 5 updates the transaction results to the database management module 1. This process not only includes updating the quota data of each enterprise, but also includes recording the transaction history, calculating the transaction price index, etc. This information provides an important basis for subsequent decision-making and also provides valuable market dynamic information for regulatory authorities.
[0179] The database management module 1 of the present invention includes three main sub-databases: a water quality information database 11, a sewage information database 12 and a quota information database 13. This hierarchical design not only improves the efficiency of data management, but also enhances the scalability and flexibility of the system.
[0180] The water quality information database 11 is used to store water quality parameters such as the concentration of various pollutants in the water body, pH value, dissolved oxygen, etc. In one embodiment of the present invention, the water quality information can be stored in time series to facilitate trend analysis and prediction. For example, a time series database such as InfluxDB can be used to store these high-frequency time series data to improve query and analysis efficiency.
[0181] The pollutant discharge information database 12 is used to store information such as the discharge amount, type of discharge, and discharge time of each pollutant discharge enterprise. The design of this database needs to take into account the characteristics of different types of pollutants. For example, different data structures may be required to represent continuous and intermittent emissions. In addition, the pollutant discharge information database 12 can also include relevant data such as the production information and technical level of the enterprise to provide comprehensive information support for subsequent policy formulation.
[0182] The quota information database 13 is used to store information such as the initial quota, current quota, and historical transaction records of each pollutant-discharging enterprise. In a preferred embodiment of the present invention, the quota information database 13 not only records static quota data, but also contains dynamic transaction information. For example, detailed information such as the time, transaction parties, transaction volume, and transaction price of each transaction can be recorded. These data can be used to analyze market dynamics, evaluate policy effects, and even predict future quota demand.
[0183] The present invention also provides a deep reinforcement learning method for dynamic optimization of the quota of emission rights of major pollutants corresponding to the above system. The method comprises the following steps:
[0184] First, water quality information and sewage discharge information of water bodies are stored and updated through the database management module 1. This step provides basic data support for the subsequent optimization process. In practical applications, water quality data can be collected in real time through IoT devices, and sewage discharge data can be obtained through the enterprise online monitoring system to ensure the timeliness and accuracy of the data.
[0185] Secondly, the environmental status assessment module 2 assesses the water quality concentration based on the water quality information of the water body, and calculates the load of various major pollutants in the water body according to the sewage discharge information. This step converts the raw data into environmental status information that can be used for decision-making. In one embodiment of the present invention, a complex water quality model can be used to assess the health status of the water body, such as the WASP (Water Quality Analysis Simulation Program) model.
[0186] Next, the multi-agent reinforcement learning module 3 receives the environmental status information and the sewage load information and executes the deep reinforcement learning algorithm. This process includes the following sub-steps:
[0187] a. Divide the action space for each agent;
[0188] b. Use deep neural networks to build policy networks and value networks;
[0189] c. Based on the current state, select an action through the policy network;
[0190] d. Perform the selected action and observe environmental feedback and rewards;
[0191] e. Update the value network and strategy network;
[0192] f. Repeat steps ce until convergence or the preset number of iterations is reached.
[0193] In this process, advanced reinforcement learning algorithms such as PPO (Proximal Policy Optimization) or SAC (Soft Actor-Critic) can be used to improve learning efficiency and stability.
[0194] Then, the strategy evaluation module 4 evaluates the global effect of the emission quota allocation scheme generated by the multi-agent reinforcement learning module 3. The evaluation indicators may include multiple dimensions such as the degree of water quality improvement, economic impact, and social fairness. Based on the evaluation results, the strategy evaluation module 4 provides feedback to the multi-agent reinforcement learning module 3 for adjusting the learning strategy.
[0195] Finally, the optimal quota allocation plan is stored in the database management module 1, and the above steps are repeated regularly to adapt to environmental changes and policy adjustments.
[0196] Through this method, the system of the present invention can achieve dynamic optimization of emission quotas and find a balance between protecting the environment and promoting economic development. Compared with the traditional fixed quota method, this method has stronger adaptability and flexibility, and can adjust the quota allocation strategy in time according to actual conditions, thereby achieving more efficient use of resources and better protection of the environment.
[0197] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A deep reinforcement learning system for dynamic optimization of emission quotas of major pollutants, characterized by: include: Database management module for: Store water quality information and sewage discharge information; Update and maintain emission quota data; The environmental status assessment module is in communication with the database management module and is used to: Based on the water quality information of the water body, assessing water quality concentration; Calculating the load of various major pollutants in the water body based on the sewage discharge information; A multi-agent reinforcement learning module is communicatively connected to the environment state assessment module and is used to: Receiving environmental status information and pollutant load information sent by the environmental status assessment module; Based on the environmental state information and the sewage load information, executing a multi-agent reinforcement learning algorithm; Generate an optimized emission quota allocation plan; A strategy evaluation module is communicatively connected with the multi-agent reinforcement learning module and the database management module, and is used to: Receiving the emission rights quota allocation plan generated by the multi-agent reinforcement learning module; Evaluate the overall effect of the emission quota allocation plan; Feedback the evaluation results to the multi-agent reinforcement learning module for strategy adjustment; The optimal quota allocation scheme is stored in the database management module.
2. The system according to claim 1, characterized in that The environmental status assessment module comprises: Water quality concentration assessment submodule, used to: Receiving water quality information of water bodies sent by the database management module; Based on the water quality information of the water body, calculating the concentration of each pollutant in the water body; Generate water quality concentration assessment reports; The sewage load assessment submodule is used to: Receiving the sewage discharge information sent by the database management module; Based on the sewage discharge information, calculating the load of each major pollutant in the water body; Generate sewage load assessment reports.
3. The system according to claim 1, characterized in that The multi-agent reinforcement learning module includes multiple single-agent sub-modules, each of which includes: Action space submodule, used to: Divide the action space according to the total quota of emission rights currently owned by the agent; Generates a set of optional quota adjustment actions; The policy network submodule is used to: Receive environmental status information; Based on deep neural networks, evaluate the policy value of each action in the action space; The value network submodule is used to: Receive environmental status information; Based on deep neural networks, the value of each action is evaluated; Dynamic programming submodule for: Combine the outputs of the policy network and the value network to estimate the long-term return of each action; Select the best action; State transfer submodule, used for: Execute the selected action; Update the environment status; Reward network submodule, used to: Calculate the immediate reward after performing the action; Adjust the policy network and value network based on the reward information.
4. The system according to claim 3, characterized in that The action space divided in the action space submodule is: {a1, a2, a3, a4}, where a1 means the current Agent does not change the quota. a2 means the current Agent reduces the quota. a3 means the current Agent maintains the quota, a4 indicates that the current Agent increases the quota.
5. The system according to claim 3, characterized in that The policy value function of the policy network submodule is: π(a t |s t ;θ)=P(a t |s t ;i), Among them, s t Indicates the environmental state, a t represents the action, and θ represents the neural network parameters.
6. The system according to claim 3, characterized in that The value function of the value network submodule is: V(s t )=E[R t |s t ], Among them, R t Indicates that from state s t The cumulative discount reward starts.
7. The system according to claim 1, characterized in that The strategy evaluation module includes: The return evaluation submodule is used to: Receiving the emission rights quota allocation plan generated by the multi-agent reinforcement learning module; Calculate the environmental and economic benefits of the plan; Generate comprehensive return evaluation report; Solution optimization submodule, used to: Based on the comprehensive return evaluation report, adjust and optimize the emission rights quota allocation plan; Generate the final optimal pollution discharge policy.
8. The system according to claim 1, characterized in that Also includes: The quota trading module is in communication with the database management module and the multi-agent reinforcement learning module and is used to: Receiving a quota allocation plan generated by the multi-agent reinforcement learning module; Carry out transaction matching of emission rights quota; The transaction results are updated to the database management module.
9. The system according to claim 1, characterized in that The database management module includes: Water quality information database, used to store water quality parameters such as concentration of various pollutants, pH value, dissolved oxygen, etc. Pollutant discharge information database, used to store information such as the amount of pollutants discharged, the type of pollutants discharged, and the time of pollutants discharged by each pollutant-discharging enterprise; The quota information database is used to store information such as the initial quota, current quota, and historical transaction records of each pollutant-discharging enterprise.
10. A deep reinforcement learning method for dynamic optimization of the quota of emission rights of major pollutants, using the system described in any one of claims 1 to 9, characterized in that: The following steps are involved: (1) Store and update water quality information and sewage discharge information through the database management module; (2) The environmental status assessment module assesses the water quality concentration based on the water quality information of the water body, and calculates the load of various major pollutants in the water body according to the sewage discharge information; (3) The multi-agent reinforcement learning module receives environmental status information and sewage load information and performs the following sub-steps: a. Divide the action space for each agent; b. Use deep neural networks to build policy networks and value networks; c. Based on the current state, select an action through the policy network; d. Perform the selected action and observe environmental feedback and rewards; e. Update the value network and strategy network; f. Repeat steps ce until convergence or the preset number of iterations is reached; (4) The strategy evaluation module evaluates the global effect of the emission quota allocation scheme generated by the multi-agent reinforcement learning module; (5) Based on the evaluation results, the strategy evaluation module provides feedback to the multi-agent reinforcement learning module for adjusting the learning strategy; (6) storing the optimal quota allocation plan in the database management module; (7) Repeat steps (1)-(6) regularly to adapt to environmental changes and policy adjustments.
Citation Information
Cited By
Multi-agent system for water quality purification material development
CN120524711A