A membrane pool optimization control method, system, equipment and medium based on multi-agent collaborative decision-making
Through the multi-agent collaborative decision-making method, the independent control and dynamic imbalance problems of the water plant membrane pool cluster were solved, water production efficiency, energy savings and membrane life extension were achieved, and the water plant was promoted to develop towards intelligence and sustainability.
Patent Information
- Application Number
- CN202510929789.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-07
AI Technical Summary
The existing water plant membrane pool cluster water production process has problems of independent control, fixed timing control rules and multi-objective dynamic imbalance, which makes it difficult to meet the global optimization needs under dynamic working conditions and lacks cluster collaborative control capabilities.
A multi-agent collaborative decision-making method is adopted to abstract the membrane pool into an independent decision-making agent, construct the state space and generate continuous control variables, create a global reward function, impose physical boundary hard constraints and trans-membrane pressure difference soft constraints, embed the agent relationship analysis layer, and optimize the strategy parameters through offline pre-training and online learning collaborative training framework.
It has achieved precise compliance of water production, optimal energy consumption distribution and balanced trans-membrane pressure difference of the membrane pool cluster under dynamic water quality and water quantity requirements, promoted the intelligent and sustainable development of water plant operations, saved energy and reduced emissions, and extended the life of membrane components.
Smart Images

Figure CN120428576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water plant operation control, and in particular to a membrane pool optimization control method, system, equipment and medium based on multi-agent collaborative decision-making. Background Art
[0002] In the water production process of the membrane pool cluster in the water plant, the traditional method has the following limitations: first, independent control, that is, each membrane pool group operates independently, lacks a global perspective, and is prone to overcapacity or undercapacity; second, static adjustment, that is, the operation of the membrane pool is based on fixed rules or manual experience adjustment, which is difficult to cope with dynamic changes in water quality and water volume. Taking Water Plant A as an example, the water production control of the membrane pool operates according to the following fixed logic: water production for 60 minutes, air flushing for 90 seconds, water production for 60 minutes, mixed flushing for 120 seconds, and then this process is repeated continuously; third, resource waste, that is, it is impossible to achieve a dynamic balance between goals such as energy consumption, membrane life, and water production. Since the parameter performance of membranes made of different materials from different manufacturers is different, even the same membrane will have differences due to different working conditions after long-term use. Control with traditional methods is inefficient.
[0003] Current research focuses primarily on optimizing single process steps, such as using deep reinforcement learning for water withdrawal risk warnings (e.g., CN117787631A) or employing multi-agent approaches for energy-efficient pump station scheduling (e.g., CN115544899A). Other studies have attempted to construct membrane fouling decision models using fuzzy rules (e.g., CN113283481A), but these efforts have yet to address the dynamic parameter optimization of membrane pool water production processes. Existing solutions for multi-objective collaborative optimization of membrane pool clusters generally suffer from a single control dimension and weak parameter adaptability, making it difficult to achieve global optimal control under complex operating conditions. Summary of the Invention
[0004] (1) Technical issues to be resolved
[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a membrane pool optimization control method, system, equipment and medium based on multi-agent collaborative decision-making, which solves the problem that the water production process of the existing water plant membrane pool cluster is difficult to meet the global optimization needs under dynamic working conditions due to the independent operation of the membrane group, fixed timing control rules and multi-objective dynamic imbalance, and the technical problem that the existing technology focuses on the optimization of a single link and lacks cluster collaborative control capabilities.
[0006] (2) Technical solution
[0007] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:
[0008] In a first aspect, an embodiment of the present invention provides a membrane pool optimization control method based on multi-agent collaborative decision-making, comprising:
[0009] Each membrane pool is abstracted into an independent decision-making intelligent agent, and a state space containing membrane pool operating parameters, environmental perception parameters, and global demand parameters is established. Continuous control variables for water production operations are generated as an action space for continuous regulation.
[0010] Create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective based on the transmembrane pressure status;
[0011] The physical operation boundary hard constraint and the dynamic soft constraint driven by the transmembrane pressure difference are simultaneously applied to the action space to obtain the restricted action space.
[0012] An agent relationship analysis layer is embedded in the evaluation network. The behavioral characteristics of each agent and the distribution of cluster collaboration intensity are obtained through high-dimensional feature mapping and multi-head attention interaction mechanism. The global value evaluation is generated by combining the rewards of the global reward function and adaptively correcting the collaboration intensity distribution.
[0013] A collaborative training framework of offline pre-training and online learning is adopted to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value evaluation, the strategy parameters of each intelligent agent are simultaneously optimized to drive the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0014] Optionally, each membrane pool is abstracted into an independent decision-making intelligent agent, and a state space containing membrane pool operating parameters, environmental perception parameters, and global demand parameters is established. Continuous control variables for water production operations are generated as an action space for continuous regulation, including:
[0015] Map each membrane pool in the membrane pool cluster into an independent decision-making intelligent agent, which generates continuous control actions based on input information;
[0016] Real-time acquisition of membrane pool operating parameters including real-time water production flow, daily cumulative water production time, and dynamic average of transmembrane pressure difference;
[0017] Synchronously obtain environmental sensing parameters including water temperature, inlet water turbidity and backwash pressure data of the membrane pool working environment;
[0018] Automatically or in response to user input instructions, introduce global demand parameters including total water demand target value and remaining dispatchable time window;
[0019] The membrane pool operation parameters, environmental perception parameters and global demand parameters are integrated according to the preset dimensions to construct a state space that represents the local observation state of the intelligent agent;
[0020] A continuous action space of water production operation control variables including a water production flow rate set value and a single water production time set value is configured for each intelligent agent. The initial control range of the water production flow rate set value and the initial duration range of the single water production time set value are determined based on a preset initial control threshold, forming the initial boundary of the continuous action space.
[0021] Optionally, a global reward function is created that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and the weight coefficients of each objective are dynamically adjusted according to the transmembrane pressure state, including:
[0022] Determine a reasonable deviation range of water production based on the total water demand target value, wherein the lower limit of the reasonable deviation range is a first preset proportion value of the total water demand, and the upper limit is a second preset proportion value of the total water demand;
[0023] When the actual total water production is equal to the total water demand target value, the highest reward benchmark is determined;
[0024] When the actual total water production is within a reasonable deviation range and deviates from the target value, the reward attenuation value is calculated based on the highest reward benchmark according to the quadratic exponential decay relationship between the degree of deviation and the target value, so that the reward value decreases at an accelerated rate as the deviation increases;
[0025] When the actual total water production is outside the reasonable deviation range, the reward attenuation value is converted into a negative penalty value, and the absolute value of the penalty value increases exponentially with the degree of deviation, forming a two-way penalty gradient for insufficient or excessive water production;
[0026] Generate a global continuous and derivable water tracking reward item based on the highest reward benchmark, the decay rule within the deviation range, and the bidirectional penalty gradient outside the deviation range;
[0027] Combined calculation of the total energy consumption of the water production process pump and backwash operation energy consumption as the total energy consumption indicator, and the energy consumption benchmark reference value is determined based on historical operating data or expert experience;
[0028] Compare the actual total energy consumption in the current water production cycle with the benchmark reference value to generate an energy consumption penalty item, and the penalty intensity of the energy consumption penalty item is incrementally amplified as the actual total energy consumption exceeds the benchmark reference value;
[0029] The transmembrane pressure difference values of each membrane pool under the water production state are normalized to the global extreme value to eliminate the hardware differences of the membrane components and generate a standardized pressure difference series;
[0030] Calculate the dynamic change of the standardized pressure difference value in adjacent time periods, execute the segmented equalization adjustment strategy according to the numerical range of the dynamic change, and output the transmembrane pressure difference equalization item that integrates the segmented control results:
[0031] Based on the water tracking reward item, energy consumption penalty item, and transmembrane pressure difference balance item, a global reward function is constructed, and initial weight coefficients are assigned to the water tracking reward item, energy consumption suppression penalty item, and transmembrane pressure difference balance item. The initial weight of the transmembrane pressure difference balance item is lower than the weight of the water tracking item.
[0032] When it is detected that the mean normalized transmembrane pressure difference exceeds the preset warning threshold, the weight coefficient of the transmembrane pressure difference balance term is increased by a preset ratio, while the weights of other terms are reduced by the same ratio, so that the control priority of the balance term in the global reward function is enhanced;
[0033] When it is detected that the mean of the normalized transmembrane pressure difference falls below the warning threshold and remains stable, the initial weight coefficient allocation of each target item is restored.
[0034] Optionally, the segmented balancing adjustment strategy is:
[0035] When the dynamic change is negative or zero, a fixed positive incentive is given;
[0036] When the dynamic change amount is within the first preset positive threshold range, the reward value is decreased according to the nonlinear attenuation rule;
[0037] When the dynamic change exceeds a first preset positive threshold, a linearly increasing penalty positively correlated with the degree of deviation is triggered.
[0038] Optionally, physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference are simultaneously applied to the action space, and the restricted action space obtained includes:
[0039] Based on the physical operation rules of the water plant membrane pool, preset minimum and maximum hard clipping constraints are imposed on the water production flow set value and the single water production time set value respectively. When the set value output by the intelligent agent exceeds the corresponding range, it is forced to be constrained to the boundary value;
[0040] The membrane pool operation status is divided into three control intervals based on the normalized mean transmembrane pressure difference obtained in real time;
[0041] The maximum allowable set values of the water production flow rate and single water production time are dynamically adjusted according to the interval to which the current mean transmembrane pressure difference belongs. The set values after hard constraint clipping are superimposed and verified with the allowable range of dynamic soft constraint restrictions, and the restricted action space is finally output;
[0042] Among them, the three-level control range is:
[0043] In the first control range, the maximum allowable water production flow rate setting value and the single water production time setting value are the maximum values of the preset hard constraints;
[0044] In the second control range, the flow rate setting upper limit is reduced according to the first preset ratio and the time setting upper limit is reduced according to the second preset ratio;
[0045] In the third control interval, the flow rate setting value upper limit is further reduced according to the third preset ratio and the time setting value upper limit is simultaneously reduced according to the fourth preset ratio;
[0046] The first preset ratio is greater than the second preset ratio, the third preset ratio is greater than the first preset ratio, and the fourth preset ratio is greater than the second preset ratio.
[0047] Optionally, an agent relationship analysis layer is embedded in the evaluation network. The behavioral characteristics of each agent and the distribution of cluster collaboration intensity are obtained through high-dimensional feature mapping and multi-head attention interaction mechanism. Combined with the rewards of the global reward function and through adaptive correction of the collaboration intensity distribution, a global value assessment is generated, including:
[0048] After embedding the agent relationship analysis layer, which includes a feature embedding module, a multi-head attention interaction module, and a feature fusion module, between the input layer and the fully connected layer of the pre-trained, centrally deployed evaluation network, the feature embedding module performs feature concatenation on the state observation data received by each agent and the output continuous control action. The concatenated feature vector is then mapped to a high-dimensional embedding space through a linear transformation to generate a high-dimensional feature vector that represents the behavioral characteristics of each agent.
[0049] In the high-dimensional embedding space, a multi-head attention interaction module with parallel computing is used to process the high-dimensional feature vectors of each agent in parallel, and the query vector, key vector, and value vector corresponding to each attention head are generated through independent linear projection.
[0050] Perform a scaled dot product operation on the query vector and the key vector, generate an attention weight matrix through normalization, use the attention weight matrix to perform weighted aggregation on the corresponding value vector, and output the local interaction features of a single attention head;
[0051] The feature fusion module concatenates and linearly fuses the local interaction features of multiple attention heads to obtain a global collaborative feature encoding that characterizes the collaborative relationship between membrane pool clusters, and extracts the cluster interaction intensity distribution information in the attention weight matrix;
[0052] The global interaction feature vector is input into the fully connected layer of the evaluation network, and the cumulative reward prediction benefit is based on the global reward function in historical experience. The prediction deviation is corrected using the cluster interaction intensity distribution information to generate a global value evaluation signal of the fusion membrane pool cluster interaction characteristics.
[0053] Optionally, a collaborative training framework of offline pre-training and online learning is used to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the policy parameters of each agent are simultaneously optimized to drive the coordinated convergence of water production, energy consumption, and transmembrane pressure balance of the membrane pool cluster, including:
[0054] Establish a digital twin model that is updated synchronously with the membrane pool cluster, and initialize the strategy network parameters deployed separately on each intelligent agent based on the historical data of the membrane pool cluster's water production, energy consumption, and transmembrane pressure balance;
[0055] In the offline training phase, the policy network is pre-trained by iteratively generating initial action instructions and simulating the state transition process;
[0056] During the online learning phase, each agent generates continuous control actions based on the current strategy output by the policy network, and injects mean-reversion exploration noise into the restricted action space. The noise amplitude is dynamically attenuated along with stability indicators including the change in water production, the rate of change in energy consumption, and the fluctuation amplitude of the transmembrane pressure difference.
[0057] The difference in the membrane pool's operating state before and after executing the control action is recorded as state transition data, and the current global reward value output by the global reward function is received to form an experience sample associated with the control action and store it in the experience replay pool;
[0058] Based on the data in the experience replay pool and the obtained global value assessment, the corresponding cluster collaboration strength distribution is back-propagated in the policy network to generate a collaboration correction gradient with the gradient direction pointing to enhanced collaboration strength;
[0059] The initial local policy gradients output by each agent's policy network are hierarchically blended with the collaborative correction gradients. A cross-temporal collaborative constraint is embedded in the blending process to generate a composite collaborative gradient vector. The cross-temporal collaborative constraint generates a penalty signal to suppress policy mutations by comparing the consistency of historical collaborative patterns with the current gradient direction.
[0060] Subspace projection is used to map the composite gradient vector to non-conflicting directions for optimizing water production, energy consumption, and transmembrane pressure difference, generating Pareto equilibrium update gradients to achieve multi-objective collaborative convergence of the membrane pool cluster.
[0061] A parameter server architecture is used to aggregate the gradients of each agent. After verifying the gradient consistency, global update instructions are broadcast. For transmission abnormal nodes, the sliding average interpolation of the gradients of adjacent agents is used to complete the task and mark the faulty nodes. All online agent strategy parameters are updated synchronously.
[0062] In a second aspect, an embodiment of the present invention provides a membrane pool optimization control method system based on multi-agent collaborative decision-making, comprising:
[0063] The agent building module is used to abstract each membrane pool into an independent decision-making agent, establish a state space containing membrane pool operating parameters, environmental perception parameters, and global demand parameters, and generate continuous control variables for water production operations as an action space for continuous regulation;
[0064] A reward function building module is used to create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective according to the transmembrane pressure state;
[0065] The action space restriction module is used to synchronously impose physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference on the action space to obtain a restricted action space;
[0066] The value assessment module is used to embed the agent relationship analysis layer in the evaluation network. It obtains the behavioral characteristics of each agent and the distribution of cluster collaboration strength through high-dimensional feature mapping and multi-head attention interaction mechanism. It combines the rewards of the global reward function and adaptively corrects the distribution of collaboration strength to generate a global value assessment.
[0067] The collaborative optimization module is used to adopt a collaborative training framework of offline pre-training and online learning to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the strategy parameters of each intelligent agent are simultaneously optimized to drive the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0068] In the third aspect, an embodiment of the present invention provides a membrane pool optimization control method and device based on multi-agent collaborative decision-making, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the membrane pool optimization control method based on multi-agent collaborative decision-making as described above.
[0069] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having computer-executable instructions stored thereon. When the executable instructions are executed by a processor, the membrane pool optimization control method based on multi-agent collaborative decision-making as described above is implemented.
[0070] (3) Beneficial effects
[0071] The beneficial effects of the present invention are: first, by abstracting each membrane pool into an intelligent entity with independent decision-making and constructing a three-dimensional state space that integrates operating parameters, environmental perception and global needs, the problem of lack of global perspective caused by independent control of traditional membrane pool clusters is solved, laying a perception foundation for the collaborative optimization of multiple membrane pools; on this basis, a multi-objective global reward function of water volume tracking, energy consumption suppression and pressure difference balance is designed, and a dynamic weight adjustment mechanism driven by the trans-membrane pressure difference state is introduced to break through the limitations of single-objective optimization and achieve a dynamic balance between water production efficiency, energy consumption savings and membrane life extension.
[0072] Furthermore, the simultaneous application of physical boundary hard constraints and pressure differential feedback soft constraints to the action space not only ensures the operational safety of water production flow and time, but also adaptively suppresses membrane module losses under abnormal operating conditions through a pressure differential graded production limit mechanism, forming a constrained action space that provides both safety and efficiency guarantees. Simultaneously, an agent relationship analysis layer is embedded in the evaluation network, utilizing high-dimensional feature mapping and a multi-head attention mechanism to analyze the distribution of cluster collaboration intensity. Combining the long-term return prediction of the global reward function with adaptive correction of collaboration intensity, this generates an evaluation signal that accurately reflects the value of multi-objective collaboration, addressing the shortcomings of traditional methods in their inability to model complex interactive relationships.
[0073] Ultimately, through a two-stage collaborative training framework combining offline pre-training and online learning, combined with controllable noise exploration within a constrained action space, efficient optimization of strategy parameters was achieved. This enabled the membrane pool cluster to achieve a combined optimal state of precise water production, optimal energy consumption allocation, and balanced transmembrane pressure differential regulation under dynamic water quality and quantity demands.
[0074] Compared with traditional control methods, the present invention achieves the goals of energy conservation and emission reduction, cost reduction and efficiency improvement, reasonable allocation of production capacity and extension of the average life of membrane groups, and significantly promotes the development of water plant operations towards intelligence and sustainability. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 A schematic flow chart of a method provided in an embodiment of the present invention;
[0076] Figure 2 A schematic diagram of a specific flow chart of step S1 of the method provided in an embodiment of the present invention;
[0077] Figure 3 A schematic diagram of a specific flow chart of step S2 of the method provided in an embodiment of the present invention;
[0078] Figure 4 A schematic diagram of a water production reward curve for the method provided in an embodiment of the present invention;
[0079] Figure 5 A schematic diagram of a transmembrane pressure difference balance term curve of the method provided in an embodiment of the present invention;
[0080] Figure 6 A schematic diagram of a specific flow chart of step S3 of the method provided in an embodiment of the present invention;
[0081] Figure 7 A schematic diagram of a specific flow chart of step S4 of the method provided in an embodiment of the present invention;
[0082] Figure 8 This is a schematic diagram of a specific flow chart of step S5 of the method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0083] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0084] like Figure 1 As shown, an embodiment of the present invention proposes a membrane pool optimization control method based on multi-agent collaborative decision-making, including: abstracting each membrane pool into an independent decision-making agent, establishing a state space including membrane pool operation parameters, environmental perception parameters and global demand parameters, and generating continuous control variables for water production operation as an action space for continuous regulation; creating a global reward function including water volume tracking, energy consumption suppression and transmembrane pressure difference balance, and dynamically adjusting the weight coefficients of each target according to the transmembrane pressure difference state; synchronously applying physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference to the action space to obtain a restricted action space; embedding an agent relationship analysis layer in the evaluation network, obtaining the behavioral characteristics of each agent and the cluster collaboration strength distribution through high-dimensional feature mapping and multi-head attention interaction mechanism, combining the rewards of the global reward function and generating a global value evaluation through adaptive correction of the collaboration strength distribution; adopting an offline pre-training and online learning collaborative training framework to inject controllable exploration noise into the restricted action space, synchronously optimizing the strategy parameters of each agent based on the empirical data collected during the training process and the global value evaluation, and driving the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0085] First, by abstracting each membrane pool into an intelligent entity with independent decision-making and constructing a three-dimensional state space that integrates operating parameters, environmental perception and global needs, the problem of lack of global perspective caused by independent control of traditional membrane pool clusters is solved, laying a perception foundation for the collaborative optimization of multiple membrane pools; on this basis, a multi-objective global reward function of water volume tracking, energy consumption suppression and pressure difference balance is designed, and a dynamic weight adjustment mechanism driven by the trans-membrane pressure difference state is introduced to break through the limitations of single-objective optimization and achieve a dynamic balance between water production efficiency, energy consumption savings and membrane life extension.
[0086] Furthermore, the simultaneous application of physical boundary hard constraints and pressure differential feedback soft constraints to the action space not only ensures the operational safety of water production flow and time, but also adaptively suppresses membrane module losses under abnormal operating conditions through a pressure differential graded production limit mechanism, forming a constrained action space that provides both safety and efficiency guarantees. Simultaneously, an agent relationship analysis layer is embedded in the evaluation network, utilizing high-dimensional feature mapping and a multi-head attention mechanism to analyze the distribution of cluster collaboration intensity. Combining the long-term return prediction of the global reward function with adaptive correction of collaboration intensity, this generates an evaluation signal that accurately reflects the value of multi-objective collaboration, addressing the shortcomings of traditional methods in their inability to model complex interactive relationships.
[0087] Ultimately, through a two-stage collaborative training framework combining offline pre-training and online learning, combined with controllable noise exploration within a constrained action space, efficient optimization of strategy parameters was achieved. This enabled the membrane pool cluster to achieve a combined optimal state of precise water production, optimal energy consumption allocation, and balanced transmembrane pressure differential regulation under dynamic water quality and quantity demands.
[0088] Compared with traditional control methods, the present invention achieves the goals of energy conservation and emission reduction, cost reduction and efficiency improvement, reasonable allocation of production capacity and extension of the average life of membrane groups, and significantly promotes the development of water plant operations towards intelligence and sustainability.
[0089] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0090] Specifically, an embodiment of the present invention provides a membrane pool optimization control method based on multi-agent collaborative decision-making, including:
[0091] S1. Abstract each membrane pool into an intelligent agent with independent decision-making, establish a state space containing membrane pool operating parameters, environmental perception parameters and global demand parameters, and generate continuous control variables for water production operations as an action space for continuous regulation.
[0092] Further, if Figure 2 As shown, step S1 includes:
[0093] S11. Map each membrane pool in the membrane pool cluster into an intelligent agent that makes independent decisions. The intelligent agent generates continuous control actions based on input information.
[0094] S12. Real-time acquisition of membrane pool operating parameters including real-time water production flow rate, cumulative water production time for the day, and dynamic average of transmembrane pressure difference. This data can be obtained through sensors such as flow meters, timers, and pressure difference sensors.
[0095] S13. Synchronously acquire environmental sensing parameters including water temperature, inlet water turbidity, and backwash pressure data of the membrane pool working environment. This data can be acquired through a temperature sensor, turbidity meter, and pressure gauge.
[0096] S14. Automatically or in response to user input, global demand parameters including the total water demand target value and the remaining schedulable time window are introduced. In the automatic phase, the global demand parameters are automatically introduced based on the daily production plan provided by the water plant scheduling system.
[0097] S15. The membrane pool operation parameters, environmental perception parameters and global demand parameters are integrated according to preset dimensions to construct a state space that represents the local observation state of the intelligent agent.
[0098] Therefore, the local observation state of each membrane pool agent s i Include:
[0099] Membrane pool status: real-time water production flow ( )、Today's cumulative water production time( ), average transmembrane pressure difference under today's water production status ( )
[0100] Environmental conditions: water temperature ( ), inlet water turbidity ( )、Backwash water pressure( )
[0101] Global demand status: total water demand ( ), remaining time window ( )
[0102] By fusing the above three types of parameters according to the preset dimensions, the mathematical expression of the state space representing the local observation state of the intelligent agent is obtained:
[0103] Among them, each parameter is normalized to eliminate dimensional differences and ensure the numerical stability of the state space.
[0104] S16. Configure a continuous action space for each intelligent agent, including a water production flow rate set value and a single water production time set value, for the water production operation control variable, and determine the initial control range of the water production flow rate set value and the initial duration range of the single water production time set value based on a preset initial control threshold value, thereby forming an initial boundary of the continuous action space.
[0105] In the action space, each agent outputs a continuous action vector ,in: : Water production flow setting value (m 3 / h) : Single water production time setting value (minutes).
[0106] S2. Create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure difference balance, and dynamically adjust the weight coefficients of each objective according to the transmembrane pressure difference status.
[0107] Furthermore, if Figure 3 As shown, step S2 includes:
[0108] S21. Determine a reasonable deviation range for water production based on the total water demand target value, where the lower limit of the reasonable deviation range is a first preset proportion value of the total water demand, and the upper limit is a second preset proportion value of the total water demand; when the actual total water production is equal to the total water demand target value, determine the highest reward benchmark.
[0109] S22. When the actual total water production is within a reasonable deviation range and deviates from the target value, the reward attenuation value is calculated based on the highest reward benchmark according to the quadratic exponential decay relationship between the degree of deviation and the target value, so that the reward value decreases at an accelerated rate as the deviation increases.
[0110] S23. When the actual total water production is outside the reasonable deviation range, the reward attenuation value is converted into a negative penalty value, and the absolute value of the penalty value increases exponentially with the degree of deviation, forming a two-way penalty gradient for insufficient or excessive water production.
[0111] S24. Generate a global continuous and derivable water tracking reward item based on the highest reward benchmark, the attenuation rule within the deviation interval, and the bidirectional penalty gradient outside the deviation interval.
[0112] Specifically, this involves a water tracking reward program, the core goal of which is to drive the water output of the membrane pool cluster to accurately match the total water demand target value, while taking into account the flexible control requirements under dynamic conditions. The specific implementation method is as follows:
[0113] Total water demand target value As the reference center, the reasonable deviation range is defined as , that is, the water production is allowed to fluctuate within the range of 70% to 130% of the target value. Figure 4 As shown, when the actual total water production equal When the maximum reward base value is given , forming a strong incentive for precise achievement of standards.
[0114] when When the reward value is within the reasonable deviation range but deviates from the target value, the reward value decreases rapidly with the degree of deviation. This is achieved through a quadratic exponential decay function:
[0115] ;
[0116] when When , the exponential term is zero and the reward is maximized to 1; when it deviates to the interval boundary (such as or ), the function value decays to zero; when When the deviation exceeds the reasonable range, the reward value turns into a negative penalty, and the penalty intensity increases exponentially with the deviation. When , the reward value drops to -0.486, forming a gradient penalty for overcapacity; similarly, when production is insufficient (such as ) Penalty increases symmetrically. Through the smoothness of the exponential function and the symmetry of the square term, we can ensure exist The entire domain is continuously differentiable, avoiding gradient mutations or oscillations during strategy optimization.
[0117] S25. Calculate the total energy consumption of the water pump during the water production process and the energy consumption of the backwash operation as the total energy consumption indicator, and determine the energy consumption benchmark reference value based on historical operating data or expert experience.
[0118] S26. Compare the actual total energy consumption in the current water production cycle with the benchmark reference value to generate an energy consumption penalty item, and the penalty intensity of the energy consumption penalty item is incrementally amplified as the actual total energy consumption exceeds the benchmark reference value.
[0119] Specifically, the energy consumption penalty items designed here include water production energy consumption and backwashing energy consumption, which are used to quantitatively evaluate the comprehensive energy consumption efficiency of water production and backwashing operations of the membrane pool cluster, with the average energy consumption within a water production cycle as the reference energy consumption value. , or a reasonable threshold set based on experience, and then the energy consumption penalty term is obtained:
[0120] ;
[0121] Where, It is the comprehensive energy consumption value of the current water production cycle, which is composed of water production energy consumption and backwash energy consumption;
[0122] ;
[0123] in, is the energy consumption value, is the water pump power coefficient (kWh / m 3 ), Backwash water volume (m 3 ), is the power coefficient of the recoil pump (kWh / m 3 ), It is the recoil energy consumption conversion coefficient, which needs to be calibrated experimentally. The default value can be 0.15.
[0124] S27. Perform global extreme value normalization on the transmembrane pressure differential values in each membrane pool under water production conditions to eliminate differences in membrane component hardware and generate a standardized pressure differential sequence. Calculate the dynamic variation of the standardized pressure differential values in adjacent time periods. Based on the numerical range of the dynamic variation, execute a segmented balancing adjustment strategy and output a transmembrane pressure differential balancing term that integrates the segmented control results. The segmented balancing adjustment strategy is as follows: when the dynamic variation is negative or zero, a fixed positive incentive is assigned; when the dynamic variation is within a first preset positive threshold, the reward value is decreased according to a nonlinear attenuation rule; and when the dynamic variation exceeds the first preset positive threshold, a linearly increasing penalty that is positively correlated with the degree of deviation is triggered.
[0125] In specific embodiments, the transmembrane pressure (TP), the pressure difference between the inlet and outlet sides of a membrane filtration process, directly affects the water production rate and membrane fouling rate. Excessively high TMPs can accelerate membrane fouling and shorten membrane life, while excessively low TMPs can lead to insufficient water production efficiency.
[0126] In order to eliminate the influence of transmembrane pressure difference between membrane pools of different manufacturers and different life cycles, the transmembrane pressure difference of each membrane pool in the water production state is normalized to the global extreme value of 0-1. The normalization calculation method is as follows:
[0127] ;
[0128] Where, Indicates the i The transmembrane pressure difference of each membrane pool, The minimum pressure difference allowed for membrane pool design, The maximum pressure difference allowed for membrane pool design.
[0129] The pressure difference change in adjacent time periods is: After standardization ∈[0,1], the maximum change in adjacent time periods | ∣≤1∣.
[0130] The normalized transmembrane pressure balance term is as follows:
[0131] ;
[0132] Where, λ It is a linear penalty coefficient that suppresses the rapid increase of pressure difference. The larger the absolute value of the coefficient, the more severe the penalty for excessive growth. λ The value is -5. is the absolute value penalty term of pressure difference, κ =2, the penalty is stronger at high pressure difference, which suppresses the risk of membrane component overload and prolongs its life.
[0133] like Figure 5As shown in the figure, this design can ensure that ① the transmembrane pressure difference balance term is greater than 0 when the average value of the transmembrane pressure difference under the standardized daily water production state decreases or increases by less than 0.1; ② the transmembrane pressure difference balance term is negative when the average value of the transmembrane pressure difference under the standardized daily water production state increases by more than 0.1, that is, a penalty is imposed.
[0134] S28. Based on the water tracking reward item, energy consumption penalty item and transmembrane pressure difference balance item, a global reward function is constructed, and initial weight coefficients are assigned to the water tracking reward item, energy consumption suppression penalty item and transmembrane pressure difference balance item, where the initial weight of the transmembrane pressure difference balance item is lower than the weight of the water tracking item.
[0135] S29. When it is detected that the mean value of the normalized transmembrane pressure difference exceeds the preset warning threshold, the weight coefficient of the transmembrane pressure difference balance item is increased according to the preset ratio, and the weights of other items are reduced in proportion, so that the control priority of the balance item in the global reward function is enhanced; when it is detected that the mean value of the normalized transmembrane pressure difference falls below the warning threshold and remains stable, that is, the fluctuation is within the preset range, the initial weight coefficient distribution of each target item is restored.
[0136] Therefore, the embodiment of the present invention constructs a global reward function for multi-objective collaborative optimization, which balances the three core objectives of achieving water production targets, minimizing energy consumption, and balancing transmembrane pressure difference through dynamic weight coefficients. The specific implementation method is as follows: ; Among them, the water tracking reward coefficient α =0.6, the highest weight, giving priority to ensuring that water production meets the standard; energy consumption penalty coefficient , negative value indicates the suppression of high energy consumption; transmembrane pressure difference balance term coefficient γ =0.1, the initial weight is the lowest, taking into account the life of the membrane component.
[0137] When the daily average of the standardized transmembrane pressure difference is ≥0.8 (the threshold is configurable), it is determined to be a high pressure difference risk state, triggering weight adjustment:
[0138] The weight of the pressure difference balance item is increased: γ Increase from 0.1 to 0.3, and reduce the weight of other items proportionally:
[0139] α Adjust to ; β Adjust to ;Adjusted total weight α +∣ β ∣+ γ =0.42+0.21+0.3=0.93, some flexibility is allowed here.
[0140] S3. Physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference are synchronously applied to the action space to obtain a restricted action space.
[0141] Furthermore, if Figure 6 As shown, step S3 includes:
[0142] S31. Based on the physical operation rules of the water plant membrane pool, preset minimum and maximum hard clipping constraints are imposed on the water production flow set value and the single water production time set value respectively. When the set value output by the intelligent agent exceeds the corresponding range, it is forced to be constrained to the boundary value.
[0143] S32. Based on the normalized mean transmembrane pressure difference obtained in real time, the operating state of the membrane pool is divided into three-level control intervals; wherein, the three-level control intervals are: in the first control interval, the maximum allowable water production flow rate setting value and the single water production time setting value are the maximum values of the preset hard constraints; in the second control interval, the upper limit of the flow rate setting value is reduced according to the first preset ratio and the upper limit of the time setting value is reduced according to the second preset ratio; in the third control interval, the upper limit of the flow rate setting value is further reduced according to the third preset ratio and the upper limit of the time setting value is simultaneously reduced according to the fourth preset ratio; the first preset ratio is greater than the second preset ratio, the third preset ratio is greater than the first preset ratio, and the fourth preset ratio is greater than the second preset ratio.
[0144] S33. Dynamically adjust the maximum allowable set values of the water production flow rate and the single water production time according to the interval to which the current mean transmembrane pressure difference belongs, and superimpose and verify the set values after hard constraint clipping with the allowable range of dynamic soft constraint restrictions, and finally output the restricted action space.
[0145] In a specific embodiment, in order to avoid invalid actions and enhance the safety and effectiveness of the model, the present invention optimizes the continuous action space and introduces a physical constraint module in the action output layer, including hard constraints based on the actual rules of the water plant membrane pool and dynamic constraints based on the transmembrane pressure difference:
[0146] (1) Enforcement of hard constraints
[0147] Flow setpoint tailoring: ;
[0148] Time setting value clipping: ;
[0149] Indicates when Less than hour, ,when Greater than hour ;
[0150] Here Indicates the minimum and maximum flow setting values of the membrane pool intelligent body, Indicates the minimum and maximum values of the single water production time setting value of the membrane pool intelligent body.
[0151] (2) Dynamic constraints driven by transmembrane pressure difference
[0152] Transmembrane pressure difference graded production limit:
[0153] ;
[0154] Where, Indicates the maximum allowable flow setting value, Indicates the maximum permissible single water production time setting value.
[0155] The final output action instructions of the embodiment of the present invention must meet both hard constraints and dynamic soft constraints. Through the collaborative design of hard constraints and dynamic soft constraints, safe and efficient regulation of the action space of the membrane pool cluster can be achieved, which meets the comprehensive requirements of the smart water plant for stability, economy and equipment life.
[0156] S4. Embed the agent relationship analysis layer in the evaluation network, obtain the behavioral characteristics of each agent and the distribution of cluster collaboration intensity through high-dimensional feature mapping and multi-head attention interaction mechanism, combine the rewards of the global reward function and adaptively correct the distribution of collaboration intensity to generate a global value evaluation.
[0157] Further, if Figure 7 As shown, step S4 includes:
[0158] S41. After embedding the agent relationship analysis layer containing the feature embedding module, the multi-head attention interaction module and the feature fusion module between the input layer and the fully connected layer of the pre-trained centrally deployed evaluation network, the state observation data received by each agent and the output continuous control action are feature spliced through the feature embedding module, and the spliced feature vector is mapped to the high-dimensional embedding space through linear transformation to generate a high-dimensional feature vector that represents the behavioral characteristics of a single agent.
[0159] First, within the digital twin environment, a large number of state-action pairs and corresponding global reward values are generated using historical operational data or random strategies to construct a pre-training dataset. An evaluation network is trained using supervised learning or reinforcement learning to predict the cumulative value of global rewards. A relationship analysis layer is then inserted between the input layer and the fully connected layer of the pre-trained network to enhance the modeling of collaborative relationships between agents. The weights of the input and fully connected layers are also locked to prevent corruption of existing knowledge during subsequent embedding.
[0160] S42. Within the high-dimensional embedding space, a multi-head attention interaction module employs parallel computing to process the high-dimensional feature vectors of each agent in parallel, generating the query vector, key vector, and value vector corresponding to each attention head through independent linear projection. "High-dimensional" refers to the use of linear transformations to map the original feature vectors to a continuous vector space with significantly higher dimensionality than the input. Within the high-dimensional embedding space, the high-dimensional features of each agent are processed in parallel, generating the query vector (for indexing guidance), key vector (for matching relevance), and value vector (carrying actual information) corresponding to each attention head through independent linear transformations. A scaled dot product operation is performed on the query vector and key vector to measure the strength of interaction between agents, and an attention weight matrix is generated through normalization to quantify the influence of different agents on the current decision.
[0161] S43. Perform a scaled dot product operation on the query vector and the key vector, generate an attention weight matrix through normalization, use the attention weight matrix to perform weighted aggregation on the corresponding value vector, and output the local interaction features of a single attention head. The attention weight matrix dynamically reflects the priority dependencies between agents during the operation of the membrane pool cluster, and achieves adaptive optimization of the intensity distribution through an end-to-end training process.
[0162] S44. The feature fusion module concatenates and linearly fuses the local interaction features of multiple attention heads to obtain a global collaborative feature encoding that characterizes the collaborative relationship of the membrane pool cluster, and extracts the cluster interaction intensity distribution information in the attention weight matrix. The local interaction features of all attention heads are concatenated by dimension to form a composite feature vector containing multi-granular collaborative information. The concatenated composite features are then reduced and integrated through a linear fusion layer to generate a global collaborative feature encoding that characterizes the overall collaborative relationship of the membrane pool cluster. At the same time, the interaction intensity distribution information between agents (such as the primary and secondary collaborative relationship) is extracted from the attention weight matrix as a basis for subsequent deviation correction.
[0163] S45. Input the global interaction feature vector into the fully connected layer of the evaluation network, predict the benefits based on the cumulative rewards of the global reward function in historical experience, and use the cluster interaction intensity distribution information to correct the prediction deviation to generate a global value evaluation signal of the fusion membrane pool cluster interaction characteristics.
[0164] In another specific embodiment, the observed states of all membrane pool agents are spliced as ,in Contains parameters such as flow rate, time, and pressure difference;
[0165] Joint action vector: The action instructions of all agents are concatenated into ,in .
[0166] After embedding the agent relationship analysis layer containing the feature embedding module, the multi-head attention interaction module and the feature fusion module between the input layer and the fully connected layer of the pre-trained centralized evaluation network, the state of each agent is analyzed through the feature embedding module. and action Splicing to generate high-dimensional feature vectors .
[0167] Next, each attention head in the multi-head attention interaction module generates an independent query (Query), key (Key), and value (Value):
[0168] ;
[0169] Where, Q 、 K 、 V Represents query, key, and value respectively. For each attention head h , in the formula Representing the h attention head query, key, and value vectors, 、 、 Represent the weight matrices that map the input vector to the query vector, key vector, and value vector respectively.
[0170] Then, a dot product operation is performed on each agent's query vector (Query) and key vector (Key) to quantify the interaction correlation between different agents and obtain a raw attention score. The raw attention score is divided by the square root of the vector dimension to prevent gradient explosion or vanishing during training and improve numerical stability. The Softmax function is applied to the scaled score and normalized along the agent dimension to generate an attention weight matrix. Each element of this matrix represents the priority of one agent's dependence on another agent's decision. The attention weight matrix is used to perform a weighted summation of the value vector (Value): the value vector with a high weight dominates the aggregation result, forming the local interaction feature of a single attention head:
[0171] ;
[0172] Where, For the h The output of the attention head, that is, h The weight of the attention head, for K h the transpose of the key vector, is the key vector dimension, used for scaling gradient stability.
[0173] It should be noted that the attention weight matrix is automatically adjusted through end-to-end training: if an agent frequently becomes the core of collaboration, its weight gradually increases during training; conversely, the weight of the secondary node decays.
[0174] The outputs of multiple attention heads are concatenated, linearly transformed, and dimensionally reduced according to the channel dimension to generate a compact global collaborative feature encoding to represent the overall collaborative situation of the cluster:
[0175] ;
[0176] Where, H represents the number of attention heads, Represents the output weight matrix, which is used to map the concatenated multi-head attention results into the final feature vector. Represents a complete multi-head attention module.
[0177] At the same time, the row / column mean or entropy value is extracted from the weight matrix of each attention head to quantify the collaborative influence of each agent (such as agent i The interaction intensity of the cluster is calculated as follows:
[0178] Furthermore, the feature vectors after attention encoding are fed into a fully connected network. Based on the cumulative rewards of the global reward function from historical experience, future long-term cumulative returns are predicted. The distribution of cluster interaction intensity is used to correct for prediction biases. For example, based on the real-time interaction intensity distribution, predictions from highly influential agents (e.g., with an intensity value > 0.6) are given higher weight. If an agent's interaction intensity suddenly drops (e.g., from 0.7 to 0.2), its historical reward contribution is reduced to avoid interference from outdated data.
[0179] Finally, the revised predicted values and real-time collaboration intensity information are combined to generate a global value assessment signal that incorporates the interactive characteristics of the membrane pool cluster. This signal not only reflects the static benefits of water production and energy consumption, but also encodes the dynamic efficiency of cluster collaboration. These steps enable precise and adaptive value assessment of membrane pool clusters, providing a reliable decision-making basis for collaborative control in complex dynamic scenarios.
[0180] S5. Adopting a collaborative training framework of offline pre-training and online learning, controllable exploration noise is injected into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the strategy parameters of each intelligent agent are simultaneously optimized to drive the coordinated convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0181] Further, if Figure 8 As shown, step S5 includes:
[0182] S51. Establish a digital twin model that is updated synchronously with the membrane pool cluster, and initialize the strategy network parameters deployed separately on each intelligent body based on the historical data of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0183] S52. In the offline training phase, the policy network is pre-trained by iteratively generating initial action instructions and simulating the state transition process.
[0184] S53. During the online learning phase, each intelligent agent generates continuous control actions based on the current strategy output by the strategy network, and injects mean-reversion exploration noise into the restricted action space. The noise amplitude dynamically decays along with the stability indicators including the change in water production, the rate of change in energy consumption, and the fluctuation amplitude of the transmembrane pressure difference.
[0185] Specifically, each agent is configured with a mean-reverting (OU) noise generator. Initial noise amplitude, mean-reverting rate, and random fluctuation intensity parameters are set to ensure temporal continuity. Changes in water production, energy consumption, and differential pressure fluctuations are monitored in real time, and a comprehensive stability index is calculated. The higher the stability (e.g., smaller fluctuations), the greater the noise attenuation coefficient, and vice versa. The OU noise amplitude is then dynamically scaled based on the preset attenuation coefficient: when stability is high, the noise amplitude is reduced (e.g., by 50%), limiting the exploration range; when stability is low, the noise amplitude is increased (e.g., to 120%), enhancing the exploration intensity.
[0186] Dynamically adjusted OU noise is superimposed on the basic actions output by the policy network to generate exploratory action instructions. A physical constraint module is then applied to clip the actions to a preset safety range. The mean-reversion properties of OU noise are automatically reduced over time to ensure smooth transitions in action sequences and avoid sudden changes between adjacent actions. If persistent out-of-bounds actions or system instability are detected, noise injection is temporarily frozen and the OU state is reset. Exploration is reactivated once stability is restored.
[0187] Here, exploration noise is added to the agent's action generation stage to help the agent try new and unknown actions to discover potentially better strategies. The OU noise method used here, also known as the mean reversion process, is suitable for application scenarios that require continuous actions and hope for a certain degree of continuity between actions.
[0188] S54. Record the difference in the membrane pool operating state before and after executing the control action as state transition data, and receive the current global reward value output by the global reward function to form an experience sample associated with the control action and store it in the experience replay pool.
[0189] S55. Based on the data in the experience replay pool and the obtained global value assessment, the corresponding cluster collaboration strength distribution is back-propagated through the policy network to generate a collaboration correction gradient that points to enhanced collaboration strength. Based on the state-action pairs in the experience replay pool and the generated global value assessment, the cluster collaboration strength distribution is analyzed using the attention weight matrix. The collaboration strength distribution is back-propagated to the policy network as a guiding signal to generate a correction gradient. Using the collaboration strength distribution to guide gradient generation solves the problem that traditional single-agent gradients lack a collaborative perspective.
[0190] S56. The initial local policy gradients output by the policy network of each agent are hierarchically mixed with the collaborative correction gradients, and a cross-time collaborative constraint term is embedded in the mixing process to generate a composite collaborative gradient vector. Specifically, the local policy gradients of each agent are mixed with the collaborative correction gradients according to a preset weight ratio (e.g., a weight ratio of 7:3). The mixing weights are dynamically adjusted according to the degree of transmembrane pressure difference balance. At the same time, a cross-time constraint term is introduced. A historical collaborative pattern is generated based on the similarity of historical action sequences and the coordinated fluctuation of pressure differences. The degree of deviation between the current gradient direction and the historical collaborative pattern is compared. If the deviation is significant, a gradient amplitude penalty including a reduction in the update step size is applied. The local policy gradients and the collaborative correction gradients are mixed, and the constraint penalty term is superimposed to form a composite collaborative gradient vector, ensuring that the optimization direction takes into account both individual actions and global collaboration.
[0191] Furthermore, the cross-time collaboration constraint item can generate a penalty signal to suppress strategy mutation by comparing historical collaboration patterns. For example, it can generate a penalty signal based on the consistency between the action sequence similarity matrix and the pressure difference collaborative fluctuation coefficient and the current gradient direction. The formula is:
[0192] ;
[0193] Where, is the current policy gradient direction (the vector direction after the local gradient and the collaborative gradient are mixed), : The direction of the historical collaborative reference mode (calculated based on the historical action sequence and the pressure difference collaborative fluctuation characteristics), the similarity of the historical action sequence is: If the two agents have the same action change direction at time step t (e.g., increasing or decreasing the water flow rate at the same time), then 1 point is awarded, otherwise 0 point is awarded. The pressure difference cooperative fluctuation characteristics are: , record the direction of pressure difference change between the two membrane pools (up / down), and count the proportion of the number of times with the same direction in a certain period of time to the total number of times. i and j In 7 of the 10 observations, the pressure difference changes in the same direction, so . SimilarityIt is a directional consistency measurement function, and cosine similarity is generally used. If Similarity≈1, it means the direction is consistent and the penalty is close to zero. If Similarity≈0, it means the direction deviates and the penalty is increased. λ ( t ) is the dynamic penalty coefficient, , ΔTMP is the pressure difference fluctuation amplitude, calculated as the standard deviation of the mean pressure difference in the current period, is the basic penalty weight, .
[0194] S57. Use subspace projection to map the composite gradient vector to non-conflicting directions for optimizing water production, energy consumption, and transmembrane pressure drop, generating Pareto equilibrium update gradients to achieve multi-objective collaborative convergence of the membrane pool cluster. Decompose the composite gradient vector into three orthogonal subspaces: water production optimization, energy consumption reduction, and pressure drop equilibrium. Projection removes conflicting components in the subspaces, retains the collaborative optimization direction, and generates Pareto equilibrium update gradients.
[0195] S58: A parameter server architecture is used to aggregate the gradients of each agent. After verifying the gradient consistency, a global update command is broadcast. For nodes with transmission anomalies, a sliding average interpolation of the gradients of adjacent agents is used to complete the update. The faulty node is marked as faulty, and the policy parameters of all online agents are updated synchronously. In this step, for nodes with transmission anomalies (such as timeouts or abnormal gradient norms), a sliding average interpolation of the historical gradients of adjacent nodes (such as the average of the last five gradients) is used to complete the update. The faulty node is marked and its impact is isolated, triggering an operations and maintenance alarm. Finally, the verified global gradient is broadcast to all online agents, and the policy network parameters are updated synchronously.
[0196] Additionally, an embodiment of the present invention provides a membrane pool optimization control method system based on multi-agent collaborative decision-making, comprising:
[0197] The intelligent agent construction module is used to abstract each membrane pool into an independent decision-making intelligent agent, establish a state space containing membrane pool operating parameters, environmental perception parameters and global demand parameters, and generate continuous control variables for water production operations as an action space for continuous regulation.
[0198] A reward function building module is used to create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective according to the transmembrane pressure state;
[0199] The action space restriction module is used to synchronously impose physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference on the action space to obtain a restricted action space.
[0200] The value assessment module is used to embed the agent relationship analysis layer in the evaluation network. It obtains the behavioral characteristics of each agent and the distribution of cluster collaboration intensity through high-dimensional feature mapping and multi-head attention interaction mechanism. It combines the rewards of the global reward function and generates a global value assessment through adaptive correction of the collaboration intensity distribution.
[0201] The collaborative optimization module is used to adopt a collaborative training framework of offline pre-training and online learning to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the strategy parameters of each intelligent agent are simultaneously optimized to drive the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
[0202] Furthermore, an embodiment of the present invention provides a membrane pool optimization control method and device based on multi-agent collaborative decision-making, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the membrane pool optimization control method based on multi-agent collaborative decision-making as described above.
[0203] Furthermore, an embodiment of the present invention provides a computer-readable storage medium having computer-executable instructions stored thereon. When the executable instructions are executed by a processor, the membrane pool optimization control method based on multi-agent collaborative decision-making as described above is implemented.
[0204] In summary, the embodiments of the present invention provide a membrane pool optimization control method, system, device, and medium based on multi-agent collaborative decision-making. Based on a multi-agent learning framework, the method dynamically sets the water production flow rate and water production time of each membrane pool to achieve the three core goals of total water production compliance, minimizing energy consumption, and balancing the transmembrane pressure difference. Its innovative implementation process can be summarized as follows:
[0205] First, each membrane pool corresponds to an independent intelligent agent, which realizes distributed decision-making through local observation space (real-time flow, pressure difference, efficiency) and action space (continuous flow / time set value), quickly responds to changes in local working conditions, and avoids the communication delay bottleneck of centralized decision-making.
[0206] Second, physical hard constraints are imposed on the flow and time setpoints, forcing out-of-bounds actions to return to a safe range through a clip function. A soft constraint driven by pressure difference is also employed across three levels of control.
[0207] Second, a multimodal composite reward function is designed for the three major optimization objectives: (1) The water volume tracking reward item adopts an exponential decay mechanism, giving a peak incentive when the total water production approaches the demand value, and accelerating the decay according to a quadratic function when it deviates; (2) The energy consumption penalty item integrates the water production energy consumption and backwash energy consumption, sets a dynamic baseline value based on the historical average, and triggers a nonlinear increasing penalty when the limit is exceeded; (3) The transmembrane pressure difference balance item eliminates hardware differences through normalization processing, and performs segmented regulation of pressure difference fluctuations, giving positive rewards when the fluctuation is stable, and applying progressive penalties when the fluctuation is abnormal. The three rewards are integrated through dynamic weight coefficients, and the weight value is automatically adjusted according to the pressure difference balance degree.
[0208] Third, an agent relationship analysis layer is embedded in the evaluation network. The behavioral characteristics of each agent and the distribution of cluster collaboration intensity are obtained through high-dimensional feature mapping and multi-head attention interaction mechanism. Based on the reward sequence in historical experience and the real-time collaboration intensity distribution, the long-term value prediction deviation is dynamically corrected to generate a global value assessment.
[0209] Fourth, a collaborative training framework of offline pre-training and online learning is adopted to inject controllable exploratory noise into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the policy parameters of each agent are simultaneously optimized. The local policy gradient and the collaborative correction gradient are dynamically proportional according to the pressure difference balance, and cross-temporal constraints are embedded to suppress policy mutations. Subspace projection technology is used to eliminate conflicts in multi-objective optimization directions. The resulting Pareto equilibrium gradient drives the agent's policy iteration, enabling the cluster to achieve dynamic balance between objectives such as water volume, energy consumption, and pressure difference. At the same time, a parameter server architecture is used to implement distributed gradient aggregation, and a sliding average interpolation mechanism is used to ensure seamless recovery of abnormal nodes, thereby improving system robustness.
[0210] The multi-agent collaborative decision-making framework of the present invention deeply integrates reinforcement learning and domain knowledge, overcomes the technical difficulty of balancing global optimization and local response in membrane pool cluster control, and provides a core control solution with autonomous evolution capability for smart water plants.
[0211] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concepts. Therefore, the technical solutions should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0212] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the technical solution of the present invention and its equivalents, the present invention shall also include such modifications and variations.
Claims
1. A membrane pool optimization control method based on multi-agent collaborative decision-making, characterized in that: include: Each membrane pool is abstracted into an intelligent agent with independent decision-making, and a state space including membrane pool operating parameters, environmental perception parameters and global demand parameters is established. Continuous control variables for water production operations are generated as action space for continuous regulation, including: mapping each membrane pool in the membrane pool cluster into an intelligent agent with independent decision-making, and the intelligent agent generates continuous control actions based on input information; obtaining membrane pool operating parameters including real-time water production flow, cumulative water production time for the day and dynamic mean of transmembrane pressure difference in real time; and synchronously obtaining environmental perception parameters including water temperature, inlet water turbidity and backwash pressure data of the membrane pool working environment. ; Automatically or in response to user input instructions, introduce global demand parameters including the total water demand target value and the remaining schedulable time window; fuse the membrane pool operation parameters, environmental perception parameters and global demand parameters according to preset dimensions to construct a state space that characterizes the local observation state of the intelligent agent; configure a continuous action space of water production operation control variables including a water production flow set value and a single water production time set value for each intelligent agent, and determine the initial control range of the water production flow set value and the initial duration range of the single water production time set value based on a preset initial control threshold to form an initial boundary of the continuous action space; Create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective based on the transmembrane pressure status; The physical operation boundary hard constraint and the dynamic soft constraint driven by the transmembrane pressure difference are simultaneously applied to the action space to obtain the restricted action space. An agent relationship analysis layer is embedded in the evaluation network. The behavioral characteristics of each agent and the distribution of cluster collaboration intensity are obtained through high-dimensional feature mapping and multi-head attention interaction mechanism. The global value evaluation is generated by combining the rewards of the global reward function and adaptively correcting the collaboration intensity distribution. A collaborative training framework of offline pre-training and online learning is adopted to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value evaluation, the strategy parameters of each intelligent agent are simultaneously optimized to drive the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
2. The membrane pool optimization control method based on multi-agent collaborative decision-making according to claim 1 is characterized in that: Create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective according to the transmembrane pressure state, including: Determine a reasonable deviation range of water production based on the total water demand target value, wherein the lower limit of the reasonable deviation range is a first preset proportion value of the total water demand, and the upper limit is a second preset proportion value of the total water demand; When the actual total water production is equal to the total water demand target value, the highest reward benchmark is determined; When the actual total water production is within a reasonable deviation range and deviates from the target value, the reward attenuation value is calculated based on the highest reward benchmark according to the quadratic exponential decay relationship between the degree of deviation and the target value, so that the reward value decreases at an accelerated rate as the deviation increases; When the actual total water production is outside the reasonable deviation range, the reward attenuation value is converted into a negative penalty value, and the absolute value of the penalty value increases exponentially with the degree of deviation, forming a two-way penalty gradient for insufficient or excessive water production; Generate a global continuous and derivable water tracking reward item based on the highest reward benchmark, the decay rule within the deviation range, and the bidirectional penalty gradient outside the deviation range; Combined calculation of the total energy consumption of the water production process pump and backwash operation energy consumption as the total energy consumption indicator, and the energy consumption benchmark reference value is determined based on historical operating data or expert experience; Compare the actual total energy consumption in the current water production cycle with the benchmark reference value to generate an energy consumption penalty item, and the penalty intensity of the energy consumption penalty item is incrementally amplified as the actual total energy consumption exceeds the benchmark reference value; The transmembrane pressure difference values of each membrane pool under the water production state are normalized to the global extreme value to eliminate the hardware differences of the membrane components and generate a standardized pressure difference series; Calculate the dynamic change of the standardized pressure difference value in adjacent time periods, execute the segmented equalization adjustment strategy according to the numerical range of the dynamic change, and output the transmembrane pressure difference equalization item that integrates the segmented control results: Based on the water tracking reward item, energy consumption penalty item, and transmembrane pressure difference balance item, a global reward function is constructed, and initial weight coefficients are assigned to the water tracking reward item, energy consumption suppression penalty item, and transmembrane pressure difference balance item. The initial weight of the transmembrane pressure difference balance item is lower than the weight of the water tracking item. When it is detected that the mean normalized transmembrane pressure difference exceeds the preset warning threshold, the weight coefficient of the transmembrane pressure difference balance term is increased by a preset ratio, while the weights of other terms are reduced by the same ratio, so that the control priority of the balance term in the global reward function is enhanced; When it is detected that the mean of the normalized transmembrane pressure difference falls below the warning threshold and remains stable, the initial weight coefficient allocation of each target item is restored.
3. The membrane pool optimization control method based on multi-agent collaborative decision-making according to claim 2 is characterized in that: The segmented balancing adjustment strategy is: When the dynamic change is negative or zero, a fixed positive incentive is given; When the dynamic change amount is within the first preset positive threshold range, the reward value is decreased according to the nonlinear attenuation rule; When the dynamic change exceeds a first preset positive threshold, a linearly increasing penalty positively correlated with the degree of deviation is triggered.
4. The membrane pool optimization control method based on multi-agent collaborative decision-making according to any one of claims 1 to 3, characterized in that: By applying physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference to the action space simultaneously, the restricted action space is obtained, including: Based on the physical operation rules of the water plant membrane pool, preset minimum and maximum hard clipping constraints are imposed on the water production flow set value and the single water production time set value respectively. When the set value output by the intelligent agent exceeds the corresponding range, it is forced to be constrained to the boundary value; The membrane pool operation status is divided into three control intervals based on the normalized mean transmembrane pressure difference obtained in real time; The maximum allowable set values of the water production flow rate and single water production time are dynamically adjusted according to the interval to which the current mean transmembrane pressure difference belongs. The set values after hard constraint clipping are superimposed and verified with the allowable range of dynamic soft constraint restrictions, and the restricted action space is finally output; Among them, the three-level control range is: In the first control range, the maximum allowable water production flow rate setting value and the single water production time setting value are the maximum values of the preset hard constraints; In the second control range, the flow rate setting upper limit is reduced according to the first preset ratio and the time setting upper limit is reduced according to the second preset ratio; In the third control interval, the flow rate setting value upper limit is further reduced according to the third preset ratio and the time setting value upper limit is simultaneously reduced according to the fourth preset ratio; The first preset ratio is greater than the second preset ratio, the third preset ratio is greater than the first preset ratio, and the fourth preset ratio is greater than the second preset ratio.
5. The membrane pool optimization control method based on multi-agent collaborative decision-making according to any one of claims 1 to 3, characterized in that: An agent relationship analysis layer is embedded in the evaluation network. The behavioral characteristics of each agent and the distribution of cluster collaboration intensity are obtained through high-dimensional feature mapping and multi-head attention interaction mechanism. Combined with the rewards of the global reward function and through adaptive correction of the collaboration intensity distribution, a global value assessment is generated, including: After embedding the agent relationship analysis layer, which includes a feature embedding module, a multi-head attention interaction module, and a feature fusion module, between the input layer and the fully connected layer of the pre-trained, centrally deployed evaluation network, the feature embedding module performs feature concatenation on the state observation data received by each agent and the output continuous control action. The concatenated feature vector is then mapped to a high-dimensional embedding space through a linear transformation to generate a high-dimensional feature vector that represents the behavioral characteristics of each agent. In the high-dimensional embedding space, a multi-head attention interaction module with parallel computing is used to process the high-dimensional feature vectors of each agent in parallel, and the query vector, key vector, and value vector corresponding to each attention head are generated through independent linear projection. Perform a scaled dot product operation on the query vector and the key vector, generate an attention weight matrix through normalization, use the attention weight matrix to perform weighted aggregation on the corresponding value vector, and output the local interaction features of a single attention head; The feature fusion module concatenates and linearly fuses the local interaction features of multiple attention heads to obtain a global collaborative feature encoding that characterizes the collaborative relationship between membrane pool clusters, and extracts the cluster interaction intensity distribution information in the attention weight matrix; The global interaction feature vector is input into the fully connected layer of the evaluation network, and the cumulative reward prediction benefit is based on the global reward function in historical experience. The prediction deviation is corrected using the cluster interaction intensity distribution information to generate a global value evaluation signal of the fusion membrane pool cluster interaction characteristics.
6. The membrane pool optimization control method based on multi-agent collaborative decision-making according to any one of claims 1 to 3, characterized in that: Using a collaborative training framework of offline pre-training and online learning, controllable exploration noise is injected into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the strategy parameters of each agent are simultaneously optimized to drive the coordinated convergence of water production, energy consumption, and transmembrane pressure balance of the membrane pool cluster. The following are included: Establish a digital twin model that is updated synchronously with the membrane pool cluster, and initialize the strategy network parameters deployed separately on each intelligent agent based on the historical data of the membrane pool cluster's water production, energy consumption, and transmembrane pressure balance; In the offline training phase, the policy network is pre-trained by iteratively generating initial action instructions and simulating the state transition process; During the online learning phase, each agent generates continuous control actions based on the current strategy output by the policy network, and injects mean-reversion exploration noise into the restricted action space. The noise amplitude is dynamically attenuated along with stability indicators including the change in water production, the rate of change in energy consumption, and the fluctuation amplitude of the transmembrane pressure difference. The difference in the membrane pool's operating state before and after executing the control action is recorded as state transition data, and the current global reward value output by the global reward function is received to form an experience sample associated with the control action and store it in the experience replay pool; Based on the data in the experience replay pool and the obtained global value assessment, the corresponding cluster collaboration strength distribution is back-propagated in the policy network to generate a collaboration correction gradient with the gradient direction pointing to enhanced collaboration strength; The initial local policy gradients output by each agent's policy network are hierarchically blended with the collaborative correction gradients. A cross-temporal collaborative constraint is embedded in the blending process to generate a composite collaborative gradient vector. The cross-temporal collaborative constraint generates a penalty signal to suppress policy mutations by comparing the consistency of historical collaborative patterns with the current gradient direction. Subspace projection is used to map the composite gradient vector to non-conflicting directions for optimizing water production, energy consumption, and transmembrane pressure difference, generating Pareto equilibrium update gradients to achieve multi-objective collaborative convergence of the membrane pool cluster. A parameter server architecture is used to aggregate the gradients of each agent. After verifying the gradient consistency, global update instructions are broadcast. For transmission abnormal nodes, the sliding average interpolation of the gradients of adjacent agents is used to complete the task and mark the faulty nodes. All online agent strategy parameters are updated synchronously.
7. A membrane pool optimization control method system based on multi-agent collaborative decision-making, characterized in that: include: The intelligent agent construction module is used to abstract each membrane pool into an intelligent agent with independent decision-making, establish a state space including membrane pool operating parameters, environmental perception parameters and global demand parameters, and generate continuous control variables for water production operation as an action space for continuous regulation, including: mapping each membrane pool in the membrane pool cluster into an intelligent agent with independent decision-making, and the intelligent agent generates continuous control actions based on input information; obtaining membrane pool operating parameters including real-time water production flow, cumulative water production time of the day and dynamic average of transmembrane pressure difference in real time; synchronously obtaining environmental data including water temperature, inlet turbidity and backwash pressure of the membrane pool working environment Environmental perception parameters; automatically or in response to user input instructions, introduce global demand parameters including the total water demand target value and the remaining schedulable time window; fuse the membrane pool operation parameters, environmental perception parameters and global demand parameters according to preset dimensions to construct a state space that represents the local observation state of the intelligent agent; configure a continuous action space of water production operation control variables including a water production flow set value and a single water production time set value for each intelligent agent, and determine the initial control range of the water production flow set value and the initial duration range of the single water production time set value based on a preset initial control threshold, forming an initial boundary of the continuous action space; A reward function building module is used to create a global reward function that includes water volume tracking, energy consumption suppression, and transmembrane pressure balance, and dynamically adjust the weight coefficients of each objective according to the transmembrane pressure state; The action space restriction module is used to synchronously impose physical operation boundary hard constraints and dynamic soft constraints driven by transmembrane pressure difference on the action space to obtain a restricted action space; The value assessment module is used to embed the agent relationship analysis layer in the evaluation network. It obtains the behavioral characteristics of each agent and the distribution of cluster collaboration strength through high-dimensional feature mapping and multi-head attention interaction mechanism. It combines the rewards of the global reward function and adaptively corrects the distribution of collaboration strength to generate a global value assessment. The collaborative optimization module is used to adopt a collaborative training framework of offline pre-training and online learning to inject controllable exploration noise into the constrained action space. Based on the empirical data collected during the training process and the global value assessment, the strategy parameters of each intelligent agent are simultaneously optimized to drive the collaborative convergence of water production, energy consumption and transmembrane pressure difference balance of the membrane pool cluster.
8. A membrane pool optimization control method and device based on multi-agent collaborative decision-making, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute the membrane pool optimization control method based on multi-agent collaborative decision-making as described in any one of claims 1-6.
9. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, the membrane pool optimization control method based on multi-agent collaborative decision-making as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Intelligent membrane pollution decision-making method based on knowledge type-2 fuzzy
CN113283481A
Water plant water intake pumping station energy-saving scheduling method based on multi-agent deep reinforcement learning
CN115544899A
Raw water system scheduling and risk early warning method based on deep reinforcement learning
CN117787631A
Control method and system of improved consistency algorithm based on multiple agents
CN119134546A
Control method and control system of multi-agent water plant process system
CN119356270A