Industrial park water-carbon collaborative optimization configuration method based on reinforcement learning

By employing a reinforcement learning-based water-carbon co-optimization allocation method, the problem of co-optimization between water recycling and carbon emissions in industrial parks was solved, achieving near-zero low-carbon wastewater discharge and efficient water resource utilization, thus supporting the goal of carbon neutrality.

CN121599210APending Publication Date: 2026-03-03HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511748884.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies lack synergistic optimization of water recycling and carbon emission processes at the industrial park scale, making it difficult to achieve near-zero low-carbon wastewater discharge.

Method used

A reinforcement learning-based method for coordinated water-carbon allocation in industrial parks is adopted. Through water network superstructure simulation, supply and demand point identification, operating cost and carbon emission simulation, control variable setting, and reinforcement learning algorithm optimization, a multi-objective optimization model is constructed to optimize water resource allocation and carbon emission management.

Benefits of technology

It has achieved near-zero low-carbon emissions of wastewater from industrial parks, improved water resource utilization efficiency, reduced carbon emissions, adapted to different development scenarios, and supported the goal of carbon peaking and carbon neutrality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121599210A_ABST
    Figure CN121599210A_ABST
Patent Text Reader

Abstract

The invention relates to an industrial park water-carbon collaborative optimization configuration method based on reinforcement learning. The method comprises the following steps: simulating a water network superstructure; simulating operation cost and carbon emission; determining a target function, a constraint condition and a decision variable; designing an algorithm and a network structure, determining a reinforcement learning algorithm, and constructing the network structure; establishing a Markov chain; training parameter setting; optimization performance analysis: analyzing algorithm performance through interval and substitute distance indexes; and optimization result analysis: analyzing the optimization result, providing a water-saving and carbon-reducing weight setting proportion under different scenes, setting park optimization weights according to different emphasis points, and analyzing the optimization result. Parameters are adjusted according to the specific conditions of the industrial park, water-carbon collaborative optimization is carried out on the industrial park under different development situations, the adaptive capacity is high, and technical support is provided for achieving the purpose that carbon reaches the peak and carbon neutralization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a reinforcement learning-based method for the coordinated optimization of water and carbon allocation in industrial parks, belonging to the technical field of carbon emission management. Background Technology

[0002] To achieve dual carbon targets, ensure that enterprises can reduce carbon emissions more efficiently, complete industrial restructuring, and achieve high-quality development, industrial park wastewater should be discharged near zero under low-carbon constraints. Balancing the water treatment and reuse process in industrial parks with its carbon emission processes is the key to achieving near-zero wastewater discharge in industrial parks under low-carbon constraints.

[0003] Current research focuses primarily on improving the efficiency of wastewater treatment and reuse technologies in enterprises, neglecting the synergy between water recycling and carbon emission processes at the industrial park scale. While this achieves wastewater treatment and high-proportion reuse for individual enterprises, it lacks optimization of the water network at the industrial park scale, making it difficult to support near-zero low-carbon emissions of wastewater from industrial parks. Therefore, it is necessary to optimize the allocation of water-carbon synergy within industrial parks. Summary of the Invention

[0004] To address the aforementioned problems, this invention discloses a reinforcement learning-based method for the coordinated optimization of water and carbon resources in industrial parks, the specific technical solution of which is as follows:

[0005] A reinforcement learning-based method for the coordinated optimization of water and carbon allocation in industrial parks includes the following steps:

[0006] Step S1: Water network superstructure simulation: The unit water network superstructure includes the superstructure of water use units, water treatment units, and regeneration units;

[0007] Supply and demand point identification includes the water network supply and demand structure and matching probability matrix;

[0008] The park's water network superstructure includes the park's water balance and the park's water regeneration;

[0009] Step S2: Simulation of operating costs and carbon emissions, including:

[0010] Operating cost simulation: Operating costs generated by sludge discharge;

[0011] Carbon emission simulation: including direct and indirect emissions;

[0012] Step S3: Set control variables, determine the objective function, constraints, and decision variables, specifically:

[0013] Three objective functions: fresh water consumption, total carbon dioxide emissions, and operating costs.

[0014] Three constraints: water balance, supply and demand balance, and limited available water resources.

[0015] Four decision variables: treatment flow rate Qa during wastewater treatment, influent and effluent flow rates Qwtl during wastewater ecological treatment, and influent flow rate Qf and operating pressure Pf of the ultrafiltration stage during wastewater regeneration treatment.

[0016] Step S4: Algorithm and network structure design, determine the reinforcement learning algorithm, and construct the network structure;

[0017] Step S5: Markov chain construction;

[0018] Step S6: Setting training parameters;

[0019] Step S7: Optimize performance analysis: Analyze the algorithm performance using metrics, including: interval and generation distance;

[0020] Step S8: Optimization Result Analysis: Analyze the optimization results, propose water-saving and carbon-reduction weight setting ratios for different scenarios, set the park optimization weights according to different focuses, and analyze the optimization results.

[0021] Furthermore, in step S2, the operating cost is the total operating cost based on the operating cost (Cost) generated by the sludge discharge volume (SP). The specific evaluation criteria are shown in the formula:

[0022] Cost = SP

[0023]

[0024] In the formula, TSS is the concentration of suspended solids, X is the particulate matter content in the waste sludge, and Q is the concentration of suspended solids. w denoted as sludge flow rate, T as sludge discharge cycle, and t as start time.

[0025] Furthermore, the carbon emission simulation in step S2 adopts the emission factor method, and the specific calculation method is as follows:

[0026] (1) Direct discharge of CH4 from sewage pipe network

[0027]

[0028] In the formula, CH4 carbon emission intensity of wastewater pipe network, kg CO2-eq / m³ 3 ; denoted as CH4 emission factor in the sewage network, kg CH4 / kg COD. Based on the anaerobic reaction process of organic matter, the theoretical value is taken as 0.25 kg CH4 / kg COD. C represents the average concentration of organic matter in the sewage network, i.e., the initial COD value within the calculation boundary range. The average concentration of organic matter entering the municipal sewage network from the High-tech Zone is used as the initial COD value, kg COD / m³. 3 ηT t represents the anaerobic conversion rate of organic matter in the sewage network; t represents the average hydraulic retention time of sewage within the calculation boundary range, in days.

[0029] (2) Direct discharge of N2O from sewage pipe network

[0030]

[0031] In the formula, The carbon emission intensity of N2O from the wastewater pipe network, expressed as kg CO2-eq / m³ 3 ; TN0 represents the N2O emission factor from the wastewater pipe network, expressed as kg N2O-N / kg N, taken as 0.005 kg N2O-N / kg N; TN0 represents the initial average total nitrogen concentration within the calculation boundary range, expressed as kg N / m³. 3 The average total nitrogen concentration entering the municipal sewage pipe network in the High-tech Zone was used as the initial total nitrogen concentration; TN e To calculate the terminal average total nitrogen concentration within the boundary range, kgN / m 3 The total nitrogen concentration at the influent of the wastewater treatment plant is used as the final total nitrogen concentration.

[0032] (3) Direct discharge of CH4 during wastewater treatment process

[0033]

[0034] In the formula, B in The average influent BOD5 concentration of the wastewater treatment plant is expressed in mg BOD5 / L. The amount of CH4 gas recovered or removed is expressed in kg CH4 / m³. 3 ;

[0035] (4) Direct discharge of N2O during wastewater treatment process

[0036]

[0037] In the formula, TN in The average total nitrogen concentration in the influent of the wastewater treatment plant, in mg N / L; The amount of N2O gas recovered or removed is expressed in kgN2O / m³. 3 ;

[0038] (5) Direct CO2 emissions during sludge treatment and disposal process

[0039]

[0040] In the formula, CO2 emission intensity from fossil sources, kg CO2-eq / m³ 3 ;MFCF represents the percentage of CO2 emissions from fossil sources; M SSThe dry weight of the sludge to be treated is kg dry sludge / a; CF is the carbon content in the dry matter, %; OF is the oxidation factor, %; Q in To evaluate the total amount of domestic sewage treated within the year, m 3 / a;

[0041] (6) Direct emission of N2O during sludge treatment and disposal process

[0042]

[0043] (7) Indirect emissions from electricity consumption

[0044] C d =(E d ·EF d ) / Q

[0045] In the formula, C d Carbon emission intensity from electricity purchased for operation and maintenance, kg CO2-eq / m³ 3 E d Total electricity consumption for operation and maintenance during the year, kWh / a; EF d The regional electricity emission factor is kg CO2-eq / kWh; Q is the total water treatment volume in the park, m³ / h. 3 / a;

[0046] (8) Indirect emissions from material consumption

[0047]

[0048] In the formula, C cl Indirect carbon emission intensity from chemicals consumed in the operation of the water system, kg CO2-eq / m³ 3 M cl.i To evaluate the total consumption of the i-th agent within a year, kg / a; EF cl,i Let be the emission factor of the i-th agent, kg CO2-eq / kg.

[0049] Furthermore, the three objective functions in step S3—fresh water usage, total carbon dioxide emissions, and operating costs—are as follows:

[0050] (1) The amount of fresh water used, f1, is the smallest:

[0051]

[0052] In the formula, F i Let i be the fresh water consumption of each enterprise, t / d, i be the i-th enterprise, and n be the total number of enterprises in the park;

[0053] (2) Minimizes operating cost f2:

[0054]

[0055] In the formula, α represents the proportion of local water to fresh water, Rpi represents the amount of reclaimed water used by each enterprise from the reclaimed water network (t / d), and Rf represents the amount of reclaimed water used by each enterprise. 1,i The amount of reclaimed water used by each enterprise after treatment at its internal wastewater treatment plant, in t / d, Rf 2,i S represents the volume of treated wastewater used by each enterprise (t / d), S represents the volume of wastewater treated by the High-tech Zone Wastewater Treatment Plant (t / d), C represents the volume of water treated by the constructed wetland (t / d), and e represents the volume of wastewater treated by the constructed wetland (t / d). α e 1-α e Rpi e Rf1,i e Rf2,i e s e c The costs per ton of water treated are local water, water from other regions, reclaimed water, wastewater treated at in-house wastewater treatment plants, wastewater treatment plants, and constructed wetlands, respectively, in yuan / t.

[0056] (3) The total carbon dioxide emissions f3 are the smallest:

[0057]

[0058] In the formula, m represents the m-th calculation unit. Let tCO2-eq / a be the direct greenhouse gas emissions of carbon in the m-th calculation unit. Let tCO2-eq / a be the indirect carbon emissions of the m-th calculation unit. Let tCO2 be the carbon sink of the m-th calculation unit, tCO2-eq / a;

[0059] The three constraints are: water balance, supply and demand balance, and water resource availability limits, specifically:

[0060] (1) Horizontal measurement: The inflow of water to each water-contacting unit = the outflow of water + the loss of water, and the loss rate is denoted as β. i ,

[0061] (1-β i (F) i Rpi+Rf 1,i +Rf 2,i ) = E 1,i +E 2,i

[0062] In the formula, E 1,i The volume of water discharged from the enterprise's water use process to the in-situ treatment and reuse process, in t / d, E 2,i The amount of water discharged from the enterprise's water usage process to the wastewater treatment plant, in t / d, F i The amount of fresh water used after preparation, t / d, Rf 1,iThe amount of reclaimed water used by each enterprise after treatment at its internal wastewater treatment plant, in t / d, Rf 2,i Rpi represents the amount of wastewater treated by the wastewater treatment plant used by each enterprise, in t / d; Rpi represents the amount of reclaimed water used by each enterprise from the reclaimed water network, in t / d.

[0063] (2) Supply and demand balance: The water consumption of each enterprise is equal before and after the optimization.

[0064] F i +Rpi+Rf 1.i +Rf 2.i =F i,0

[0065] In the formula, F i,0 To optimize the fresh water consumption of each enterprise before configuration, t / d;

[0066] (3) Limitations on available water resources:

[0067]

[0068] In the formula, W0 is the maximum water supply of the water diversion project, t / d, W i The local water resource availability is expressed in t / d, and α represents the proportion of local water to fresh water.

[0069] i. Nonnegativity constraint:

[0070] F i ≥0

[0071] That is, the amount of fresh water used should not be less than 0;

[0072] ii. Water-carbon relationship index constraint:

[0073]

[0074] The carbon emission intensity per cubic meter of water in each water-related unit shall not exceed the maximum value;

[0075] In the formula, RW m Carbon emission intensity per cubic meter of water in the m-th water-contacting unit, kgCO2-eq / m 3 C m Let tCO2-eq / a be the carbon emissions of the m-th water-contact unit. RW represents the water treatment capacity of the m-th water-contacting unit, in ten thousand tons per year. max This represents the historical maximum value of carbon emissions and available water volume for the m-th water-related unit, i.e., the maximum carbon emission intensity per cubic meter of water, expressed as kgCO2-eq / m³. 3 ;

[0076] Four decision variables: treatment flow rate Qa during wastewater treatment, influent and effluent flow rates Qwtl during ecological wastewater treatment, and influent flow rate Qf and operating pressure Pf of the ultrafiltration stage during wastewater regeneration treatment. Specifically:

[0077] In order to obtain optimized setpoints for operational variables with different time scales simultaneously, decision variables reflecting the changes of operational variables over time were designed. The optimization time length was divided according to the time scale of each operational variable, and each operational variable corresponding to a unit time length became the decision variable. Finally, a set of decision variables composed of time series values ​​of multiple operational variables was obtained.

[0078] Furthermore, step S4, the algorithm and network structure design, specifically includes:

[0079] Choose the reinforcement learning algorithm—PPO algorithm;

[0080] The PPO algorithm's network structure design is based on the AC framework and consists of three neural networks: two Actor networks and one Critic network.

[0081] Furthermore, the PPO algorithm process is as follows: First, random action sampling is performed to obtain a set of training data through interaction with the environment, and the data is stored in the experience pool. Next, random sampling is performed in the experience pool, and the data is input into the Critic network to calculate the state value. Then, the advantage function GAE is used to calculate the advantage function value and the discounted cumulative reward. The discounted cumulative reward is used as the objective to update the Critic network and reduce the error between the Critic network's predicted value and the true value. The advantage function GAE is used as the weight to calculate the policy gradient and update the Actor network. After the Actor network has been updated for a certain period, the parameters of the Actor network are assigned to the old Actor network. The above steps are repeated until the optimal policy parameters corresponding to the objective function are found.

[0082] Furthermore, step S5 involves constructing a reinforcement learning Markov chain to improve algorithm performance, specifically as follows:

[0083] Multi-objective optimization problems can be viewed as non-zero-sum random games involving multiple agents, where multiple agents take actions by choosing the values ​​of decision variables and obtain their respective rewards from the park's water system according to the objective function. In reinforcement learning, agents rely on tuples {S, A, T, R} to interact with the environment and learn continuously through Markov decision-making (MDP) processes, thereby optimizing their strategies.

[0084] The system considers the optimal management of objectives and the optimization of wastewater treatment processes, forming a Markov game framework for a multi-objective optimization problem. In this multi-agent system, experience sharing among different agents can improve algorithm performance. Therefore, within the Markov game framework, each agent is allowed to observe the actions and rewards of others. Water variable information obtained by simulating the park's water system processes is used to construct the state space. The operating cost and carbon emissions in the objective function are calculated using this water variable information, forming the reward. After receiving feedback from the Markov game environment, the agents adjust control variables by interacting with the environment, thereby improving the optimization objective in the park's water system. First, a multi-agent decision-making model is built based on the tuple {S, A, T, R}.

[0085] (1) State S

[0086] The water-carbon system process involves multiple state variables x. Selecting variables with distinctive characteristics allows the agent to understand and become familiar with its current environment. Therefore, the agent's state at time t can be expressed as:

[0087]

[0088] In the formula, x n State variables representing the corresponding water-related processes in the park, including flow rate, water quality, carbon emissions, and cost, x in This represents the state variable of the Markov matrix in the i-th dimension, used for its state transition;

[0089] (2) Action A

[0090] Reinforcement learning agents learn by observing states, deriving optimal choices for the current environment, and optimizing their policies based on independently accumulated rewards. The agent's actions are defined as follows:

[0091]

[0092] In the formula, K L a5 represents the parameter simulating dissolved oxygen concentration, δ represents the parameter range, and Q... int ξ represents the parameter range for simulating nitrogen oxide concentration.

[0093] (3) Transition probability T

[0094] The state transition probability T is the probability that the agent chooses action A at time t and transitions from state St to the next state St+1. Evaluating the agent's performance through state transitions helps the agent converge to the optimal control policy more quickly. For all actions and states that satisfy the constraints in the above optimization problem, we have:

[0095] T(S t+1 ||S t A t)>0 and

[0096] (4) Reward R

[0097] The common goal of all agents in the model is to simultaneously minimize operating costs, fresh water consumption, and greenhouse gas emissions. Therefore, the reward function is set to point to the corresponding objective function, expressed as:

[0098] R1(S t A t )=R2(S t A t ) = r c -w g ·g(X l )-w h ·h(X m )-w j ·j(X n ).

[0099] In the formula, r c The penalty for an agent when water quality limits are exceeded; w g w h w j These represent the weights of operating costs, fresh water consumption, and total carbon dioxide emissions in the multi-objective optimization.

[0100] Furthermore, step S6 involves determining a set of training parameter settings through debugging:

[0101] The system parameters set during training include the experience pool capacity D, the number of samples M, the actor learning rate, the discount factor, the initial exploration rate, the final exploration rate, and the online neural network update frequency in the PPO algorithm. The initial exploration rate controls the balance between the agent's exploration and strategy utilization. As the number of training iterations increases, the exploration rate is multiplied by the discount factor and gradually decreases until the final exploration rate.

[0102] Furthermore, step S7 analyzes the algorithm performance using interval and generation distance metrics, specifically as follows:

[0103] (1) Spacing (SPA)

[0104] The principle of the Spacing Index (SPA) is based on the statistical distribution of nearest neighbor distances. By calculating the Euclidean standard deviation of each solution to other solutions in the Pareto front, SPA measures the uniformity of solution distribution on the Pareto front. The smaller the SPA value, the closer it is to 0, indicating a more uniform distribution of solutions on the Pareto front.

[0105]

[0106] Where PF represents the optimized Pareto solution space, and di represents the minimum distance from the i-th solution to other solutions in PF. This represents the average value of all solutions;

[0107] (2) General distance (GD)

[0108] Algebraic distance is a metric for evaluating the convergence of optimization results. It measures the average closeness between the Pareto approximation front found by the algorithm and the true Pareto front. It calculates the "distance" from each point in the approximation front to the true front and then takes the average. A smaller value indicates better convergence.

[0109]

[0110] Furthermore, step S8 specifically includes:

[0111] Weighting analysis quantifies the importance of each objective to determine the priority of decision optimization, including cost priority, water conservation priority, low carbon priority, and balanced development. The weights are evaluated using a weighted sum model.

[0112]

[0113] Where: M is the number of optimization objectives; w k It is the weight of the k-th objective; f norm,k,j Score is the value of the j-th solution on the k-th normalization objective. j It is the overall score of the solution.

[0114] The beneficial effects of this invention are:

[0115] Based on the analysis of water-carbon relationships in industrial parks and guided by the principle of "water-carbon synergistic configuration - wastewater source reduction", this invention develops a water-carbon synergistic optimization configuration technology for complex industrial parks to achieve near-zero low-carbon wastewater discharge in the region.

[0116] This invention is based on reinforcement learning technology in artificial intelligence, requiring less data and converging faster. The invention allows for parameter adjustments based on the specific conditions of industrial parks, enabling water and carbon co-optimization under different development scenarios. It exhibits high adaptability and provides technical support for achieving carbon peaking and carbon neutrality goals. Attached Figure Description

[0117] Figure 1 This is a diagram of the unit water network superstructure of the present invention.

[0118] Figure 2 This is a water network supply and demand structure diagram for industrial parks according to the present invention.

[0119] Figure 3 This is a schematic diagram of the regional water network superstructure of the present invention.

[0120] Figure 4 This is a superstructure diagram of the park water network of the present invention.

[0121] Figure 5 This is a diagram of the PPO algorithm architecture of the present invention.

[0122] Figure 6 This is a diagram of the Markov chain framework based on the PPO algorithm of this invention. Detailed Implementation

[0123] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0124] The core process of this invention is as follows:

[0125] Analyzing and quantifying the overall structure of the industrial park's water-carbon system is fundamental to constructing a water-carbon synergistic optimization model. It is also a necessary approach for the subsequent rational allocation of water resources and achieving near-zero wastewater discharge from industrial parks under low-carbon constraints. From a systemic and holistic perspective, this study simulates the composition of the park's water system based on a water network superstructure, including water supply, water use, drainage, water treatment, and water reuse—all aspects involving water quantity and quality. It also constructs simulations of the industrial park's operating costs and carbon emissions. Providing logical guarantees and a fundamental basis for subsequent algorithm construction and operation is a crucial step in building the industrial park's water-carbon synergistic optimization model, ensuring its stability and reliability.

[0126] Setting control variables is a prerequisite for constructing a water-carbon synergistic optimization model for industrial parks and a key step in clarifying the synergistic optimization objective. Control variables include the objective function, constraints, and decision variables. Based on simulation and in-depth research of the water-carbon system, reasonable control variables for the industrial park system are constructed, and the optimization logic is adjusted based on the corresponding variable settings. Furthermore, based on the characteristics of the water-carbon system in the industrial park, a reinforcement learning algorithm is selected, a reasonable network structure is built to ensure the convergence and stability of the optimization algorithm, the algorithm network architecture is drawn, a reinforcement learning Markov chain is constructed, and a Markov game framework is used to improve the algorithm's performance.

[0127] Through repeated debugging, effective algorithm training parameter settings were determined and applied to the reinforcement learning algorithm, ultimately yielding the results of water and carbon synergistic optimization in industrial parks. Based on this, the optimization performance of the algorithm was further analyzed, including convergence interval and convergence distance, reflecting the algorithm's convergence uniformity, scalability, and convergence. Observing that both the interval and distance are less than 0.1 indicates that the algorithm's optimization simulation results exhibit a uniform distribution, suggesting that the Pareto solution set found by the algorithm closely approximates the real global Pareto optimal frontier. This means that the optimized solution is very close in quality to the theoretically best possible solution. The convergence, distribution, and coverage exhibited by the performance indicators demonstrate the strong applicability of the multi-objective optimization model in the research of water and carbon synergistic problems in industrial parks. The optimal performance algorithm optimization results were selected, and water-saving and carbon-reduction weight settings were proposed for different scenarios to address the real-world problems of industrial park optimization. The optimization weights for industrial parks were set according to different emphases, and the optimization results were analyzed.

[0128] The specific execution process of the present invention is given below with reference to specific embodiments:

[0129] S1: Water network superstructure simulation. Variables highly correlated with carbon emissions from the water system are selected and included in the model. The unit water network superstructure mainly includes the superstructures of water-using units, water treatment units, and regeneration units. Supply and demand point identification mainly includes factors such as the water network supply and demand structure and the matching probability matrix. The park water network superstructure mainly includes factors such as park water balance and park water regeneration.

[0130] (1) Unit water network superstructure

[0131] The unit water network superstructure (also called the typical water-related unit superstructure) includes a superstructure of water-using units, water treatment units, and regeneration units. The water intake of water-using unit j includes fresh water (W_(F,j)), water from other water-using units (W_(p,j), W_(k,j)), and water from distributed and centralized regeneration units (W_(ij,2), W_(RM,j)). The water output includes water destined for the treatment unit (W_(WW,j,k)), water from other water-using units (W_(j,p)), water from the regeneration unit (W_(ji,1)), and water loss (W_(Lo,j)).

[0132] The water inlet of the water treatment unit is all the effluent from the water-using unit (W_(WW, j, k)), and the effluent includes water loss (W_(Lo, j)) and treated wastewater going to other units (W_(Out, WW, k)).

[0133] The water intake of the regeneration unit includes the water effluent from the water-using unit (W_(ji, 1)), and the water effluent includes water loss (W_(Lo, i)), drainage to the treatment unit (W_(WW, i)), and regenerated water to the water-using unit (W_(ij, 2)).

[0134] Specific unit water network superstructures such as Figure 1 As shown.

[0135] (2) Identification of supply and demand points

[0136] Supply and demand point identification is based on the superstructure of unit water networks. It analyzes the water quantity and quality characteristics of each stream at the inlet and outlet of water-related units to identify water-related units that can be recycled or utilize reclaimed water, providing a foundation for the construction of multi-scale water network superstructures. This invention exemplifies the identification of water supply and demand points for each type of water-related unit by simplifying and merging the water quality requirements and drainage characteristics of 86 water-related structural units from a water quality perspective.

[0137] When identifying the supply and demand points of water-related units, both water volume and water usage type are considered. Regarding water volume, establishing a wastewater reuse pathway is uneconomical when the water volume at the supply or demand point of a water-related unit is less than a certain value. Therefore, it is considered to merge the small flow at the supply point into a large flow, remove the demand point from the wastewater reuse connection, and consider supplying the demand point with tap water. Regarding water usage type, considering the high requirements for water quality and safety for domestic water use, reclaimed water is rarely used currently; therefore, the use of reclaimed water for domestic water use is not considered. Furthermore, the demand points of water-using units using pure water only correspond to the supply points of the pure water preparation unit, and there is no room for optimization; therefore, this type of demand point is not considered. Correspondingly, pure water produced by the pure water preparation unit is also not included in the scope.

[0138] Based on the above principles, the water supply and demand points of various types of water-related units are analyzed. Water supply points mainly exist in the circulating cooling unit, the pure water preparation unit, and the wastewater reclamation unit. The clean wastewater from the circulating cooling unit, the clean wastewater from the pure water preparation unit, and the reclaimed water produced by the wastewater reclamation unit can all be used for water supply. Water demand points mainly exist in the general water use unit, the circulating cooling unit, and the pure water preparation unit. The inlet water of the general water use unit, the makeup water of the circulating cooling unit, and the tap water or reclaimed water required by the pure water preparation unit all constitute water demand.

[0139] Based on the aforementioned water-related unit supply and demand points, a drainage network structure for the high-tech zone will be established according to the low-carbon development model, such as... Figure 2As shown, after being treated at the company's on-site wastewater treatment facilities or stations, the company's wastewater enters the High-tech Zone Wastewater Treatment Plant, which serves as the core of wastewater treatment for the entire High-tech Zone. The effluent from the wastewater treatment plant flows to a reclaimed water plant and an ecological treatment unit, respectively. The concentrated wastewater from the reclaimed water plant is returned to the wastewater treatment plant for further treatment. The treated reclaimed water that meets the standards is supplied to all companies through the reclaimed water supply network. The effluent from the ecological wetland is directly discharged into the receiving river.

[0140] The water supply and demand structure of industrial parks, such as Figure 2 As shown.

[0141] For the water supply and demand points in the water network of key enterprises in the industrial park, the matching possibility of supply and demand points within the park's water network is analyzed based on the water quality characteristics of each point. Factors considered include water quality, location, and water quantity at the supply and demand points.

[0142] Table 1. Supply and Demand Matching Probability Matrix

[0143]

[0144] First, consider the water quality matching degree between water supply points and water demand points. The probability of matching between supply and demand points can be determined according to their water quality categories, forming a supply-demand point matching probability matrix. The columns of the matrix represent each water demand point, the rows represent each water supply point, and the matrix elements represent the probability of matching. If the water quality of the supply and demand points meets the matching principle, the corresponding element has a value of 1; otherwise, the element has a value of 0. Finally, a matching probability matrix for all supply and demand points in the water network can be formed. Based on this matching probability matrix, the pipeline investment cost and water saving cost of different supply-demand matching methods are compared. The water demand point with the smallest difference between investment cost and water saving cost is the optimal choice for the water supply point. Finally, the supply-demand point matching probability matrix corresponding to this superstructure model is formed, as shown in Table 1.

[0145] (3) Superstructure of the park's water network

[0146] The park's water network superstructure is based on the water-related unit superstructure. Taking a holistic view of the park, it focuses on water supply, drainage, water treatment, and water discharge, considering the park's water system's connection with surrounding rivers and lakes, and expressing the possible connections between various units within the park. Based on commonly used layouts for industrial park water systems, a park water network superstructure is designed, emphasizing the park's water supply and water treatment units. Figure 3The superstructure represents the park's waterworks, wastewater treatment plant, reclaimed water plant, and constructed wetlands as water treatment units. Water sources at the park level include the park's waterworks, reclaimed water plant, and constructed wetlands. Both water supply and treatment units can supply water to water-using units, with water quality corresponding to the treatment effect of each unit. This superstructure reflects the drainage water quality characteristics of different water-related units within the park, providing a foundation for integrated direct and indirect wastewater reuse methods. It explores the potential for wastewater reuse within the park through differentiated water supply and on-demand water supply, while fully considering the needs of the surrounding aquatic ecological environment.

[0147] Based on the existing water balance in the park, and considering the park's water system needs and the characteristics of the influent and effluent water quality of each water-related unit, the connections between the various water-related units in the park are organized and reconstructed to form a superstructure model of the park's water network. Since most enterprises in the park have water usage processes that utilize reclaimed water, such as pure water preparation, the demand for reclaimed water is significant. In conjunction with the park's development plan to build a reclaimed water plant, the superstructure model of the park's water network identifies the reclaimed water plant as a crucial wastewater recycling unit. Because the effluent from the wastewater treatment plant has good water quality, meeting the Class A standard, and is characterized by a large, continuous, and relatively stable volume, the effluent from the wastewater treatment plant is used as the influent for the reclaimed water plant. The effluent from the reclaimed water plant is supplied to various enterprises in the park for production use according to the principles of reclaimed water demand and supply-demand matching.

[0148] Furthermore, the park is surrounded by a complex and slow-flowing river network with limited environmental capacity, necessitating efforts to minimize the park's negative impact on the surrounding water system. Currently, there are constructed wetlands connected to this river network in the vicinity of the park. Therefore, it is considered to utilize these constructed wetlands to purify the wastewater treatment plant's effluent, and then use the purified water for ecological replenishment of the surrounding rivers. This would reduce the amount of pollutants discharged from the park into the surrounding rivers, increase the flow rate and velocity of the river system, and improve environmental capacity.

[0149] Based on the above analysis, a superstructure model of the water network in the studied park was constructed, such as... Figure 4 As shown.

[0150] S2: Operating cost and carbon emission simulation, including the following:

[0151] Operating Cost Simulation: To evaluate the control strategies for the wastewater treatment process, this invention only considers the operating cost (Cost) incurred by sludge production (SP). The specific evaluation criteria for operating cost are shown in the formula:

[0152] Cost = SP

[0153]

[0154] In the formula, TSS is the concentration of suspended solids, Q is the particulate matter content in the waste sludge, and Q w The flow rate of waste sludge.

[0155] Carbon emission simulation: In terms of time frame, this invention only considers the emissions during the main carbon emission generation phase—the operation and maintenance phase.

[0156] In this invention, all direct and indirect greenhouse gas emissions are primarily calculated using the emission factor method, with the specific calculation method as follows:

[0157] (1) Direct discharge of CH4 from sewage pipe network

[0158]

[0159] In the formula, CH4 carbon emission intensity of wastewater pipe network, kg CO2-eq / m³ 3 ; denoted as CH4 emission factor in the sewage network, kg CH4 / kg COD. Based on the anaerobic reaction process of organic matter, the theoretical value is taken as 0.25 kg CH4 / kg COD. C represents the average concentration of organic matter in the sewage network, i.e., the initial COD value within the calculation boundary range. The average concentration of organic matter entering the municipal sewage network from the High-tech Zone is used as the initial COD value, kg COD / m³. 3 η T denoted as anaerobic conversion rate of organic matter in the sewage network; t is the average hydraulic retention time of sewage within the calculation boundary range, expressed in days.

[0160] (2) Direct discharge of N2O from sewage pipe network

[0161]

[0162] In the formula, The carbon emission intensity of N2O from the wastewater pipe network, expressed as kg CO2-eq / m³ 3 ; TN0 represents the N2O emission factor from the wastewater pipe network, expressed as kg N2O-N / kg N, taken as 0.005 kg N2O-N / kg N; TN0 represents the initial average total nitrogen concentration within the calculation boundary range, expressed as kg N / m³. 3 The average total nitrogen concentration entering the municipal sewage pipe network in the High-tech Zone was used as the initial total nitrogen concentration; TN e To calculate the terminal average total nitrogen concentration within the boundary range, kgN / m 3 The total nitrogen concentration at the influent of the wastewater treatment plant is used as the final total nitrogen concentration.

[0163] (3) Direct discharge of CH4 during wastewater treatment process

[0164]

[0165] In the formula, B in The average influent BOD5 concentration of the wastewater treatment plant is expressed in mg BOD5 / L. The amount of CH4 gas recovered or removed is expressed in kg CH4 / m³. 3 .

[0166] (4) Direct discharge of N2O during wastewater treatment process

[0167]

[0168] In the formula, TN in The average total nitrogen concentration in the influent of the wastewater treatment plant, in mg N / L; The amount of N2O gas recovered or removed is expressed in kgN2O / m³. 3 .

[0169] (5) Direct CO2 emissions during sludge treatment and disposal process

[0170]

[0171] In the formula, CO2 emission intensity from fossil sources, kg CO2-eq / m³ 3 ;MFCF represents the percentage of CO2 emissions from fossil sources; M SS The dry weight of the sludge to be treated is kg dry sludge / a; CF is the carbon content in the dry matter, %; OF is the oxidation factor, %; Q in To evaluate the total amount of domestic sewage treated within the year, m 3 / a.

[0172] (6) Direct emission of N2O during sludge treatment and disposal process

[0173]

[0174] (7) Indirect emissions from electricity consumption

[0175] C d =(E d ·EF d ) / Q

[0176] In the formula, C d Carbon emission intensity from electricity purchased for operation and maintenance, kg CO2-eq / m³ 3 E d Total electricity consumption for operation and maintenance during the year, kWh / a; EF d The regional electricity emission factor is kg CO2-eq / kWh.

[0177] (8) Indirect emissions from material consumption

[0178]

[0179] In the formula, C cl The indirect carbon emission intensity generated by chemicals, etc., consumed in the operation of the water system, expressed as kg CO2-eq / m³. 3 M cl,i To evaluate the total consumption of the i-th agent within a year, kg / a; EF cl,i Let be the emission factor of the i-th agent, kg CO2-eq / kg.

[0180] Step S3, setting control variables, specifically includes:

[0181] Three objective functions: fresh water consumption, total carbon dioxide emissions, and operating costs;

[0182] (1) Minimum use of fresh water

[0183]

[0184] In the formula, Fi represents the fresh water consumption of each enterprise (t / d).

[0185] (2) Minimal system operating cost

[0186]

[0187] In the formula, α represents the proportion of local water to fresh water, Rpi represents the amount of reclaimed water used by each enterprise from the reclaimed water network (t / d), and R f1,i R represents the amount of reclaimed water (t / d) used by each enterprise after treatment at its internal wastewater treatment plant. f2,i S represents the volume of treated wastewater used by each enterprise (t / d), S represents the volume of wastewater treated by the High-tech Zone Wastewater Treatment Plant (t / d), C represents the volume of water treated by the constructed wetland (t / d), and e represents the volume of wastewater treated by the constructed wetland. α e 1-α e Rpi e Rf1,i e Rf2,i e s e c The costs per ton of water treated are as follows (RMB / t): local water, water from other regions, reclaimed water, wastewater treated at in-house wastewater treatment plants, wastewater treatment plants, and constructed wetlands.

[0188] (3) Water networks have the lowest carbon emissions.

[0189]

[0190] In the formula, C dm Let C be the direct greenhouse gas emissions (tCO2-eq / a) of carbon in the m-th calculation unit. inmLet C be the indirect carbon emissions (tCO2-eq / a) of the m-th computational unit. zbm Let t be the carbon sink amount of the m-th calculation unit (tCO2-eq / a).

[0191] Three constraints: water balance, supply and demand balance, and limited water supply.

[0192] (1) Horizontal measurement: The inflow of water to each water-contacting unit = the outflow of water + the loss of water, and the loss rate is denoted as β. i

[0193] (1-β i (F) i Rpi+Rf 1,i +Rf 2,i ) = E 1,i +E 2,i

[0194] In the formula, β i E represents the loss rate. 1,i E represents the volume (t / d) of water discharged from the enterprise's water use process to the in-situ treatment and reuse process. 2,i This refers to the amount of water (t / d) that enterprises discharge into wastewater treatment plants during their water usage process.

[0195] (2) Supply and demand balance: The water consumption of each enterprise is equal before and after the optimization.

[0196] F i +Rpi+Rf 1.i +Rf 2.i =F i,0

[0197] In the formula, F i,0 To optimize the fresh water consumption (t / d) of each enterprise before configuration, F i Rpi represents the fresh water consumption after configuration (t / d), Rf represents the reclaimed water consumption (t / d) from the reclaimed water network used by each enterprise, and Rf represents the fresh water consumption after configuration. 1,i Rf represents the amount of reclaimed water (t / d) used by each enterprise after treatment at its internal wastewater treatment plant. 2,i This refers to the amount of wastewater (t / d) treated by the wastewater treatment plant used by each enterprise.

[0198] (3) Limitations on available water resources:

[0199]

[0200] In the formula, W0 is the maximum water supply of the water diversion project (t / d), W l This represents the available water supply from local water resources (t / d).

[0201] iv. Nonnegativity constraint:

[0202] F i ≥0

[0203] v. Water-carbon relationship index constraint:

[0204]

[0205] In the formula, RW m Carbon emission intensity per cubic meter of water in the m-th water-related unit (kgCO2-eq / m) 3 ), C m For the carbon emissions (tCO2-eq / a) of the m-th water-contact unit, W sm RW represents the water treatment capacity (ten thousand tons / year) of the m-th water-contacting unit. max This represents the historical maximum value of carbon emissions and available water volume for the m-th water-related unit, i.e., the maximum carbon emission intensity per cubic meter of water (kgCO2-eq / m³). 3 ).

[0206] Four decision variables: treatment flow rate Qa during wastewater treatment; influent and effluent flow rates Qwtl during wastewater ecological treatment; and influent flow rate Qf and operating pressure Pf of the ultrafiltration stage during wastewater regeneration treatment.

[0207] Based on the analysis of carbon emission links in the entire water-related process, the decision variables for multi-objective optimization of the park's water network include: treatment flow rate Qa in the wastewater treatment process; influent and effluent flow rates Qwtl in the ecological wastewater treatment process; and influent flow rate Qf in the ultrafiltration stage and operating pressure Pf in the reverse osmosis stage in the wastewater regeneration process.

[0208] Since frequent changes in pump flow rate will increase energy consumption and cause wear and tear on the pump, the frequency of pump flow rate changes should be minimized. Therefore, the time scale for the internal return flow rate Qa in wastewater treatment and the influent and effluent flow rates Qwtl in wastewater ecological treatment is set to 1 hour. In the wastewater regeneration process, the filtration cycle for ultrafiltration and reverse osmosis is set to 0.5 hours; therefore, the time scale for the influent flow rate Qf in the ultrafiltration stage and the operating pressure Pf in the reverse osmosis stage is set to 0.5 hours.

[0209] To simultaneously obtain optimized setpoints for operational variables with different time scales, decision variables that reflect the changes of operational variables over time were designed. The optimization time length was divided according to the time scale of each operational variable, and the operational variables corresponding to each unit of time length became the decision variables. Finally, a set of decision variables consisting of the time series values ​​of multiple operational variables was obtained.

[0210] Step S4, Algorithm and Network Structure Design: Determine the reinforcement learning algorithm and construct the network structure, specifically including:

[0211] The chosen reinforcement learning algorithm - PPO algorithm:

[0212] Network Structure Design of the PPO Algorithm: The PPO algorithm is based on the AC framework and consists of three neural networks: two Actor networks and one Critic network. The Actor networks continuously adjust the policy parameters based on reward expectations to increase the probability of higher rewards. The Critic network obtains the potential value of the current state through the relationship between the environment and rewards. Therefore, the PPO algorithm uses the Actor networks to guide the Critic network, causing it to converge towards higher rewards. Below is the pseudocode for the PPO algorithm used to optimize wastewater treatment processes:

[0213]

[0214]

[0215] The PPO algorithm flowchart is as follows: Figure 5 As shown: First, random action sampling is performed to obtain a set of training data through interaction with the environment, and this data is stored in the experience pool. Next, random sampling is performed from the experience pool, and the data is input into the Critic network to calculate the state value. Then, the advantage function (GAE) is used to calculate the advantage function value and the discounted cumulative reward. Using the discounted cumulative reward as the objective, the Critic network is updated to minimize the error between the Critic network's predictions and the true values. The advantage function (GAE) is used as weights to calculate the policy gradient and update the Actor network. After a certain number of cycles, the parameters of the Actor network are assigned to the old Actor network. This process is repeated until the optimal policy parameters corresponding to the objective function are found.

[0216] Step S5: Construct a Markov chain using reinforcement learning to improve algorithm performance.

[0217] The aforementioned multi-objective optimization problem can be viewed as a non-zero-sum stochastic game involving multiple agents. Multiple agents take actions by selecting values ​​for decision variables, obtaining their respective rewards from the environment (the park's water system) according to the objective function. In reinforcement learning, the agents rely on tuples {S, A, T, R} to interact with the environment and continuously learn through a Markov Decision Process (MDP), thereby optimizing their strategies.

[0218] The system considers the optimal management of objectives and the optimization of wastewater treatment processes, forming a Markov game framework for a multi-objective optimization problem. In this multi-agent system, experience sharing among different agents can improve algorithm performance. Therefore, within this framework, each agent is allowed to observe the actions and rewards of others. Figure 6As shown, water variable information obtained by simulating the park's water system processes forms the state space. These variables are used to calculate the operating cost and carbon emissions in the objective function, forming the reward. After receiving feedback in the Markov game environment, the agent adjusts the control variables by interacting with the environment, thereby improving the optimization objective within the environment. First, a multi-agent decision-making model is built based on the tuple {S, A, T, R}. Figure 6 As shown.

[0219] (1) State S

[0220] The wastewater treatment process involves multiple state variables X. Selecting variables with distinctive characteristics allows the agent to understand and become familiar with its current environment. Therefore, the agent's state at time t can be expressed as:

[0221]

[0222] (2) Action A

[0223] Reinforcement learning agents learn by observing states, deriving optimal choices for the current environment, and optimizing their policies based on independently accumulated rewards. Therefore, the agent's actions are defined as:

[0224]

[0225] (3) Transition probability T

[0226] The state transition probability T is the probability that the agent chooses action A at time t, from S t Transition to the next state S t+1 The probability of [the action / state]. The performance of the agent is evaluated through state transitions, helping the agent converge to the optimal control policy more quickly. For all actions and states that satisfy the constraints in the above optimization problem, we have:

[0227] T(S t+1 |S t A t )>0 and

[0228] (4) Reward R

[0229] For each agent in the model, the common goal is to simultaneously minimize operating costs, fresh water consumption, and greenhouse gas emissions. Therefore, the reward function is set to point to the corresponding objective function, expressed as:

[0230] R1(S t A t )=R2(S t A t ) = r c -w g ·g(X l )-wh ·h(X m )-w j ·j(X n Step S6: Determine a set of training parameter settings through debugging.

[0231] The system parameters set during training are shown in Table 2. In the PPO algorithm, the experience pool size D, number of samples M, Actor learning rate, discount factor, initial exploration rate, final exploration rate, and online neural network update frequency are set to 104, 128, 0.001, 0.9, 1, 0.01, and 10, respectively. The initial exploration rate controls the balance between the agent's exploration and strategy utilization. As the number of training iterations increases, the exploration rate is multiplied by the discount factor, gradually decreasing until the final exploration rate.

[0232] Table 2 PPO Algorithm Training Parameter Settings

[0233] Parameter name numerical values Sample M 128 Experience Pool D 104 Actor Learning Rate 0.001 Discount factor 0.9 Initial exploration rate 1 Termination of exploration rate 0.01 Online neural network update frequency 10

[0234] Step S7: Analyze the algorithm performance using metrics, including: interval and generational distance.

[0235] (1) Spacing (SPA)

[0236] The principle of this index is based on the statistical distribution of nearest neighbor distance. By calculating the Euclidean standard deviation of each solution to other solutions on the Pareto front, SPA measures the uniformity of the solution distribution on the Pareto front. The smaller the SPA value (closer to 0), the more uniformly the solutions are distributed on the Pareto front.

[0237]

[0238] Key characteristics of the SPA metric: Focusing on uniformity rather than convergence, the SPA metric does not concern itself with whether the found solution set closely approximates the true Pareto optimal front. It only evaluates the uniformity of the distribution among solutions within the currently found solution set. Essentially, it is the standard deviation of the distances between solutions. Ideally, all solutions should be equally distributed. However, the SPA metric cannot assess convergence and coverage: a small SPA value for a solution set, representing its uniform distribution, may actually indicate that it lies entirely within a local optimum and does not approach the global Pareto front. It needs to be used in conjunction with Global Distributed Analysis (GD) to comprehensively evaluate the quality of the solution set.

[0239] (2) General Distance (GD)

[0240] Algebraic distance is an indicator for evaluating the convergence of optimization results. It measures the average closeness between the Pareto approximation front found by the algorithm and the true Pareto front. The method is to calculate the "distance" from each point in the approximation front to the true front, and then take the average value. The smaller the value, the better the convergence.

[0241]

[0242] The park optimization results evaluated two important optimization indicators. Based on these, additional parameters such as uniformity, scalability, and average local density were selected. The main performance evaluation indicators are shown in Table 3.

[0243] Table 3 Performance Evaluation of Optimization Indicators

[0244] Serial Number Parameter name numerical values Reference range 1 spacing 0.000722 <0.1 2 Uniformity 0.2194 <0.6 3 Scalability 0.41 <0.6 4 Distance 0.00077 <0.1 5 Average local density 32.85 <40

[0245] The values ​​of the interval and the generation distance are both much less than 0.1, which indicates that the results of the algorithm optimization simulation show a uniform distribution. This means that the Pareto solution set found by the algorithm is very close to the real global Pareto optimal frontier, implying that the optimized solution is very close to the theoretically best possible solution in terms of quality. The convergence, distribution and coverage of the performance indicators all show that the optimization multi-objective optimization model has strong applicability in the study of water and carbon synergy problems in industrial parks.

[0246] Step S8 involves analyzing the optimization results, proposing water-saving and carbon-reduction weight settings for different scenarios, setting optimization weights for the park according to different focuses, and analyzing the optimization results:

[0247] Weighting analysis determines the focus of decision optimization by quantifying the importance of each objective. Based on the development needs of the park, it is divided into several categories such as cost priority, water conservation priority, low carbon priority, and balanced development.

[0248] The weights are evaluated using a weighted sum model:

[0249]

[0250] Where: M is the number of optimization objectives; w k It is the weight of the k-th objective; f norm,k,j Score is the value of the j-th solution on the k-th normalization objective. j It is the overall score of the solution.

[0251] To address the practical issues of park optimization, the following table details the water conservation and carbon reduction weight settings for different scenarios. Table 4 shows the park optimization weights set according to different focuses:

[0252] Table 4 Weight Allocation Table

[0253] Serial Number type Weight 1 Balanced (0.33,0.33,0.34) 2 Water-saving priority (0.6,0.2,0.2) 3 Low carbon priority (0.2,0.6,0.2) 4 Economic priority (0.2,0.2,0.6) 5 Water-saving and carbon-reducing (0.4,0.4,0.2) 6 Water-saving and economical (0.4,0.4,0.2)

[0254] The Pareto optimal solutions with added weights are shown in Table 5:

[0255] Table 5 shows the corresponding optimal solutions under different conditions.

[0256]

[0257] (1) Water-saving priority scheme

[0258] Table 6 shows the specific details of water resource allocation, production costs, and carbon emissions for each system in the water-saving priority scheme.

[0259] Table 6. Water-saving priority parameter configuration

[0260]

[0261] The water-saving priority scheme is suitable for industrial parks with the following main characteristics: large water consumption, a single state-owned enterprise as the main player, low water reuse rate due to early technical problems, and inherently insufficient water resources, which is a major bottleneck restricting social and economic development.

[0262] The water-saving priority scheme achieves a strict water conservation target by setting a water reuse rate of 94.72%. The production system cost structure exhibits characteristics of water-scarce regions, with industrial electricity costs exceeding water costs, and sludge treatment accounting for a significant proportion of these costs. Chemical consumption carbon emissions reached 181,792 (t / a), demonstrating that achieving a high reuse rate requires the addition of chemicals in the treatment process and the use of chemicals in the production process.

[0263] The water-saving-first, water-carbon co-optimization scheme proposed in this invention significantly reduces the amount of fresh water used in industrial parks and improves the reuse rate of water resources. The optimized scheme reduces fresh water usage and water network carbon emissions by 33% and 14% respectively compared to before optimization, at the cost of a 36% increase in economic costs. Within the framework of water-carbon co-optimization technology for industrial parks under low-carbon constraints, the primary objective of water resource allocation is to ensure water resource utilization and efficiency while minimizing carbon emissions throughout the water network process. Economic benefits are secondary, and the benefits of water resource conservation and carbon reduction must be guaranteed at a certain economic cost.

[0264] (2) Low-carbon priority scheme

[0265] Table 7 shows the specific details of water resource allocation, production costs, and carbon emissions for each system in the low-carbon priority scheme.

[0266] Table 7 Low-carbon priority parameter configuration

[0267]

[0268] The low-carbon priority scheme is suitable for industrial parks with the following main characteristics: large carbon emissions, large-scale integrated raw material manufacturing production, and the generation of a large amount of high-concentration wastewater.

[0269] The low-carbon priority scheme, by setting a water reuse rate of 91.01%, an energy share of 0.6990, and replacing coal-fired boilers with natural gas, reduces the park's carbon emissions to 4,141,163 t / a, achieving the low-carbon target under water-saving conditions. In the production system cost structure, water costs are lower than water-saving costs, mainly because the price of water intake is cheaper than reusing the same amount of water. Maintenance costs are slightly higher, while other costs remain basically the same. Chemical consumption carbon emissions are 156,239 t / a. The relatively low water reuse rate reduces the carbon emissions from electricity and chemicals used in the park's wastewater treatment, achieving low-carbon goals while maximizing the park's water reuse objectives.

[0270] The low-carbon-priority water-carbon synergistic optimization scheme proposed in this invention significantly reduces carbon emissions in the industrial park and improves the reuse rate of water resources. The optimized scheme reduces fresh water consumption and water network carbon emissions by 12% and 45% respectively compared to before optimization, at the cost of a 43% increase in economic costs.

[0271] (3) Economic priority scheme

[0272] Table 8 shows the details of water resource allocation, production costs, and carbon emissions for each system in the economy-first scheme.

[0273] Table 8. Parameter Configuration for Economy Priority Model

[0274]

[0275] The key characteristics of the industrial parks that fit the economic priority scheme are small scale, main products are consumer goods, and high production and processing costs. Their system optimization configuration reflects the characteristics under cost constraints.

[0276] The economy-first approach, with a 90.03% water reuse rate, a 0.6064% energy share, and a high proportion of coal-fired boilers, resulted in carbon emissions of 4,603,427 tons per year for the industrial park. This reflects the realistic trade-offs faced by small businesses between economic development and green transformation. Water reuse in the park ensures basic water resource recycling efficiency while avoiding high-cost investment in advanced treatment technologies, aligning with the practical path for small businesses to promote water conservation within limited budgets. In the production system cost structure, costs for electricity purchase, water use, maintenance, sludge treatment, and water treatment have all decreased, primarily due to the change in the proportion of water reuse. Simultaneously, the energy structure dictates an increase in direct carbon emissions from energy combustion, with chemical consumption resulting in carbon emissions of 166,433 tons per year. The relatively low water reuse rate reduces the carbon emissions from electricity and chemicals used in wastewater treatment, achieving low carbon emissions while maximizing the park's water reuse goals.

[0277] The economic-priority water-carbon co-optimization scheme proposed in this invention significantly reduces the economic costs of the industrial park. The optimized scheme reduces fresh water consumption, economic costs, and water network carbon emissions by 6%, 28%, and 15%, respectively, compared to the unoptimized scheme.

[0278] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as herein.

[0279] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A reinforcement learning-based method for the coordinated optimization of water and carbon allocation in industrial parks, characterized in that, Including the following steps: Step S1: Water network superstructure simulation: The unit water network superstructure includes the superstructure of water use units, water treatment units, and regeneration units; Supply and demand point identification includes the water network supply and demand structure and matching probability matrix; The park's water network superstructure includes the park's water balance and the park's water regeneration; Step S2: Simulation of operating costs and carbon emissions, including: Operating cost simulation: Operating costs generated by sludge discharge; Carbon emission simulation: including direct and indirect emissions; Step S3: Set control variables, determine the objective function, constraints, and decision variables, specifically: Three objective functions: fresh water consumption, total carbon dioxide emissions, and operating costs. Three constraints: water balance, supply and demand balance, and limited available water resources. Four decision variables: treatment flow rate Qa during wastewater treatment, influent and effluent flow rates Qwtl during wastewater ecological treatment, and influent flow rate Qf and operating pressure Pf of the ultrafiltration stage during wastewater regeneration treatment. Step S4: Algorithm and network structure design, determine the reinforcement learning algorithm, and construct the network structure; Step S5: Markov chain construction; Step S6: Setting training parameters; Step S7: Optimize performance analysis: Analyze the algorithm performance using metrics, including: interval and generation distance; Step S8: Optimization Result Analysis: Analyze the optimization results, propose water-saving and carbon-reduction weight setting ratios for different scenarios, set the park optimization weights according to different focuses, and analyze the optimization results.

2. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 1, characterized in that, In step S2, the operating cost is the total operating cost based on the operating cost (Cost) generated by the sludge discharge volume (SP). The specific evaluation criteria are shown in the formula: Cost = SP In the formula, TSS is the concentration of suspended solids, X is the particulate matter content in the waste sludge, and Q is the concentration of suspended solids. w denoted as sludge flow rate, T as sludge discharge cycle, and t as start time.

3. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 1, characterized in that, In step S2, the carbon emission simulation uses the emission factor method, and the specific calculation method is as follows: (1) Direct discharge of CH4 from sewage pipe network In the formula, CH4 carbon emission intensity of wastewater pipe network, kg CO2-eq / m³ 3 ; denoted as CH4 emission factor in the sewage network, kgCH4 / kg COD. Based on the anaerobic reaction process of organic matter, the theoretical value is taken as 0.25 kgCH4 / kg COD. C represents the average concentration of organic matter in the sewage network, i.e., the initial COD value within the calculation boundary range. The average concentration of organic matter entering the municipal sewage network from the High-tech Zone is used as the initial COD value, kgCOD / m³. 3 ; η T t represents the anaerobic conversion rate of organic matter in the sewage network; t represents the average hydraulic retention time of sewage within the calculation boundary range, in days. (2) Direct discharge of N2O from sewage pipe network In the formula, The carbon emission intensity of N2O from the wastewater pipe network, expressed as kg CO2-eq / m³ 3 ; TN0 represents the N2O emission factor from the wastewater pipe network, expressed as kg N2O-N / kg N, taken as 0.005 kg N2O-N / kg N; TN0 represents the initial average total nitrogen concentration within the calculation boundary range, expressed as kg N / m³. 3 The average total nitrogen concentration entering the municipal sewage pipe network in the High-tech Zone was used as the initial total nitrogen concentration; TN e To calculate the terminal average total nitrogen concentration within the boundary range, kgN / m 3 The total nitrogen concentration at the influent of the wastewater treatment plant is used as the final total nitrogen concentration. (3) Direct discharge of CH4 during wastewater treatment process In the formula, B in The average influent BOD5 concentration of the wastewater treatment plant is expressed in mg BOD5 / L. The amount of CH4 gas recovered or removed is expressed in kg CH4 / m³. 3 ; (4) Direct discharge of N2O during wastewater treatment process In the formula, TN in The average total nitrogen concentration in the influent of the wastewater treatment plant, in mg N / L; The amount of N2O gas recovered or removed is expressed in kgN2O / m³. 3 ; (5) Direct CO2 emissions during sludge treatment and disposal process In the formula, CO2 emission intensity from fossil sources, kg CO2-eq / m³ 3 ;MFCF represents the percentage of CO2 emissions from fossil sources; M SS The dry weight of the sludge to be treated is kg dry sludge / a; CF is the carbon content in the dry matter, %; OF is the oxidation factor, %; Q in To evaluate the total amount of domestic sewage treated within the year, m 3 / a; (6) Direct emission of N2O during sludge treatment and disposal process (7) Indirect emissions from electricity consumption C d =(E d ·EF d ) / Q In the formula, C d Carbon emission intensity from electricity purchased for operation and maintenance, kg CO2-eq / m³ 3 E d Total electricity consumption for operation and maintenance during the year, kWh / a; EF d The regional electricity emission factor is kg CO2-eq / kWh; Q is the total water treatment volume in the park, m³ / h. 3 / a; (8) Indirect emissions from material consumption In the formula, C cl Indirect carbon emission intensity from chemicals consumed in the operation of the water system, kg CO2-eq / m³ 3 ; M cl.i To evaluate the total consumption of the i-th agent within a year, kg / a; EF cl,i Let be the emission factor of the i-th agent, kg CO2-eq / kg.

4. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 1, characterized in that, The three objective functions in step S3 are: fresh water usage, total carbon dioxide emissions, and operating costs, as detailed below: (1) The amount of fresh water used, f1, is the smallest: In the formula, F i Let i be the fresh water consumption of each enterprise, t / d, i be the i-th enterprise, and n be the total number of enterprises in the park; (2) Minimizes operating cost f2: In the formula, α represents the proportion of local water to fresh water, Rpi represents the amount of reclaimed water used by each enterprise from the reclaimed water network (t / d), and Rf represents the amount of reclaimed water used by each enterprise. 1,i The amount of reclaimed water used by each enterprise after treatment at its internal wastewater treatment plant, in t / d, Rf 2,i S represents the volume of treated wastewater used by each enterprise (t / d), S represents the volume of wastewater treated by the High-tech Zone Wastewater Treatment Plant (t / d), C represents the volume of water treated by the constructed wetland (t / d), and e represents the volume of wastewater treated by the constructed wetland (t / d). α e 1-α e Rpi e Rf1,i e Rf2,i e s e c The costs per ton of water treated are local water, water from other regions, reclaimed water, wastewater treated at in-house wastewater treatment plants, wastewater treatment plants, and constructed wetlands, respectively, in yuan / t. (3) The total carbon dioxide emissions f3 are the smallest: In the formula, m represents the m-th calculation unit. Let tCO2-eq / a be the direct greenhouse gas emissions of carbon in the m-th calculation unit. Let tCO2-eq / a be the indirect carbon emissions of the m-th calculation unit. Let tCO2 be the carbon sink of the m-th calculation unit, tCO2-eq / a; The three constraints are: water balance, supply and demand balance, and water resource availability limits, specifically: (1) Horizontal measurement: The inflow of water to each water-contacting unit = the outflow of water + the loss of water, and the loss rate is denoted as β. i , (1-β i )(F i Rpi+Rf 1,i +Rf 2,i )=E 1,i +E 2,i In the formula, E 1,i The volume of water discharged from the enterprise's water use process to the in-situ treatment and reuse process, in t / d, E 2,i The amount of water discharged from the enterprise's water usage process to the wastewater treatment plant, in t / d, F i The amount of fresh water used after preparation, t / d, Rf 1,i The amount of reclaimed water used by each enterprise after treatment at its internal wastewater treatment plant, in t / d, Rf 2,i Rpi represents the amount of wastewater treated by the wastewater treatment plant used by each enterprise, in t / d; Rpi represents the amount of reclaimed water used by each enterprise from the reclaimed water network, in t / d. (2) Supply and demand balance: The water consumption of each enterprise is equal before and after the optimization. F i +Rpi+rf 1.i +Rf 2.i =F i,0 In the formula, F i,0 To optimize the fresh water consumption of each enterprise before configuration, t / d; (3) Limitations on available water resources: In the formula, W0 is the maximum water supply of the water diversion project, t / d, W i The local water resource availability is expressed in t / d, and α represents the proportion of local water to fresh water. i. Nonnegativity constraint: f i ≥0 That is, the amount of fresh water used should not be less than 0; ii. Water-carbon relationship index constraint: The carbon emission intensity per cubic meter of water in each water-related unit shall not exceed the maximum value; In the formula, RW m Carbon emission intensity per cubic meter of water in the m-th water-contacting unit, kgCO2-eq / m 3 C m Let tCO2-eq / a be the carbon emissions of the m-th water-contact unit. RW represents the water treatment capacity of the m-th water-contacting unit, in ten thousand tons per year. max This represents the historical maximum value of carbon emissions and available water volume for the m-th water-related unit, i.e., the maximum carbon emission intensity per cubic meter of water, expressed as kgCO2-eq / m³. 3 ; 4 Decision variables: Treatment flow rate Qa during wastewater treatment, influent and effluent flow rates Qwtl during ecological wastewater treatment, and influent flow rate Qf and operating pressure Pf of the ultrafiltration stage during wastewater regeneration treatment. Specifically: In order to obtain optimized setpoints for operational variables with different time scales simultaneously, decision variables reflecting the changes of operational variables over time were designed. The optimization time length was divided according to the time scale of each operational variable, and each operational variable corresponding to a unit time length became the decision variable. Finally, a set of decision variables composed of time series values ​​of multiple operational variables was obtained.

5. The reinforcement learning-based method for coordinated water-carbon allocation in industrial parks according to claim 1, characterized in that, The algorithm and network structure design in step S4 specifically include: Choose the reinforcement learning algorithm—PPO algorithm; PPO algorithm network structure design: Based on the AC framework, the algorithm has three neural networks, two of which are Actor networks and one is a Critic network.

6. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 6, characterized in that, The PPO algorithm process is as follows: First, random action sampling is performed to obtain a set of training data through interaction with the environment, and the data is stored in the experience pool. Next, random sampling is performed in the experience pool, and the data is input into the Critic network to calculate the state value. Then, the advantage function (GAE) is used to calculate the advantage function value and the discounted cumulative reward. The discounted cumulative reward is used as the objective to update the Critic network and reduce the error between the Critic network's predicted value and the true value. The advantage function (GAE) is used as the weight to calculate the policy gradient and update the Actor network. After the Actor network has been updated for a certain period, the parameters of the Actor network are assigned to the old Actor network. The above steps are repeated until the optimal policy parameters corresponding to the objective function are found.

7. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 1, characterized in that, Step S5 involves constructing a reinforcement learning Markov chain to improve algorithm performance. Specifically: Multi-objective optimization problems can be viewed as non-zero-sum random games involving multiple agents, where multiple agents take actions by choosing the values ​​of decision variables and obtain their respective rewards from the park's water system according to the objective function. In reinforcement learning, agents rely on tuples {S, A, T, R} to interact with the environment and learn continuously through Markov decision-making (MDP) processes, thereby optimizing their strategies. The system considers the optimal management of objectives and the optimization of wastewater treatment processes, forming a Markov game framework for a multi-objective optimization problem. In this multi-agent system, experience sharing among different agents can improve algorithm performance. Therefore, within the Markov game framework, each agent is allowed to observe the actions and rewards of others. Water variable information obtained by simulating the park's water system processes is used to construct the state space. The operating cost and carbon emissions in the objective function are calculated using this water variable information, forming the reward. After receiving feedback from the Markov game environment, the agents adjust control variables by interacting with the environment, thereby improving the optimization objective in the park's water system. First, a multi-agent decision-making model is built based on the tuple {S, A, T, R}. (1) State S The water-carbon system process involves multiple state variables x. Selecting variables with distinctive characteristics allows the agent to understand and become familiar with its current environment. Therefore, the agent's state at time t can be expressed as: In the formula, x n State variables representing the corresponding water-related processes in the park, including flow rate, water quality, carbon emissions, and cost, x in This represents the state variable of the Markov matrix in the i-th dimension, used for its state transition; (2) Action A Reinforcement learning agents learn by observing states, deriving optimal choices for the current environment, and optimizing their policies based on independently accumulated rewards. The agent's actions are defined as follows: In the formula, K L a5 represents the parameter simulating dissolved oxygen concentration. For the parameter range, Q int ξ represents the parameter range for simulating nitrogen oxide concentration. (3) Transition probability T The state transition probability T is the probability that the agent chooses action A at time t and transitions from state St to the next state St+1. Evaluating the agent's performance through state transitions helps the agent converge to the optimal control policy more quickly. For all actions and states that satisfy the constraints in the above optimization problem, we have: T(S t+1 |S t A t )>0 and (4) Reward R The common goal of all agents in the model is to simultaneously minimize operating costs, fresh water consumption, and greenhouse gas emissions. Therefore, the reward function is set to point to the corresponding objective function, expressed as: R1(S t ,A t )=R2(S t ,A t )=r c -w g ·g(X l )-w h ·h(X m )-w j ·j(X n )。 In the formula, r c The penalty for an agent when water quality limits are exceeded; w g w h w j These represent the weights of operating costs, fresh water consumption, and total carbon dioxide emissions in the multi-objective optimization.

8. The reinforcement learning-based method for coordinated water-carbon allocation in industrial parks according to claim 1, characterized in that, Step S6 involves determining a set of training parameter settings through debugging: The system parameters set during training include the experience pool capacity D, the number of samples M, the actor learning rate, the discount factor, the initial exploration rate, the final exploration rate, and the online neural network update frequency in the PPO algorithm. The initial exploration rate controls the balance between the agent's exploration and strategy utilization. As the number of training iterations increases, the exploration rate is multiplied by the discount factor and gradually decreases until the final exploration rate.

9. The reinforcement learning-based water-carbon synergistic optimization allocation method for industrial parks according to claim 1, characterized in that, Step S7 analyzes the algorithm performance using interval and generation distance metrics, specifically as follows: (1) Spacing (SPA) The principle of the Spacing Index (SPA) is based on the statistical distribution of nearest neighbor distances. By calculating the Euclidean standard deviation of each solution to other solutions in the Pareto front, SPA measures the uniformity of solution distribution on the Pareto front. The smaller the SPA value, the closer it is to 0, indicating a more uniform distribution of solutions on the Pareto front. Where PF represents the optimized Pareto solution space, and di represents the minimum distance from the i-th solution to other solutions in PF. This represents the average value of all solutions; (2) General Distance (GD) Algebraic distance is a metric for evaluating the convergence of optimization results. It measures the average closeness between the Pareto approximation front found by the algorithm and the true Pareto front. It calculates the "distance" from each point in the approximation front to the true front and then takes the average. A smaller value indicates better convergence.

10. The reinforcement learning-based method for coordinated water-carbon allocation in industrial parks according to claim 1, characterized in that, Step S8 specifically involves: Weighting analysis quantifies the importance of each objective to determine the priority of decision optimization, including cost priority, water conservation priority, low carbon priority, and balanced development. The weights are evaluated using a weighted sum model. Where: M is the number of optimization objectives; w k It is the weight of the k-th objective; f norm,k,j Score is the value of the j-th solution on the k-th normalization objective. j It is the overall score of the solution.