A smart regulation system for multiphase transformation of water quality in gate-controlled river sections
By constructing an intelligent regulation and control system for multiphase transformation of water quality in gate-controlled river sections, the problem of rapid emergency response to sudden water pollution incidents in gate-controlled river sections has been solved. Real-time data assimilation and multi-gate dam collaborative regulation and control have been achieved, improving the speed of emergency response and the scientific nature of regulation and control decisions, and enhancing the level of intelligent management of river network water quality.
Patent Information
- Application Number
- CN202511696106.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-19
AI Technical Summary
Existing technologies are insufficient to respond quickly and effectively to sudden water pollution incidents in sluice-controlled river sections, especially in complex multi-sluice-dam river networks. They cannot achieve coordinated operation and real-time control of multiple sluice-dams, resulting in delayed emergency response and threatening downstream water supply security and ecological health.
A smart regulation system for multiphase transformation of water quality in sluice-controlled river sections was constructed. Through data assimilation and status update modules, proxy model calculation modules, collaborative strategy decision-making modules, and rolling optimization and command execution modules, the system can achieve real-time data assimilation, accurate prediction of multiphase transformation processes of pollutants, and automatic generation of collaborative emergency regulation schemes for multiple sluice gates and dams.
It enables precise and dynamic perception of the multiphase transformation process of water quality in gate-controlled river sections, ultra-fast prediction of pollutant concentration fields and phase evolution, generation of globally optimal gate group collaborative control schemes, improvement of emergency response speed and scientific nature of control decisions, and enhancement of the intelligent level of river network water quality management.
Smart Images

Figure CN121165498B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of water treatment, more particularly, the present application relates to a gated river section water quality multi-phase conversion intelligent regulation system. BACKGROUND
[0002] In recent years, with the rapid development of economic and social in the basin of China, the gated river section has become an important carrier of water resources allocation and flood control and safety. However, the high-risk industries such as chemical industry and metallurgy distributed along the river make the river face the serious threat of sudden water pollution events. Such events have strong unpredictability and great destructiveness. For example, after the leakage accident of toxic chemicals occurs in the upstream chemical plant, high-concentration pollutants enter the river instantaneously and are transported downstream with the water flow. In the complex multi-dam series river network, the pollution group not only undergoes convection and diffusion under the action of water power, but also undergoes a series of physical, chemical and biological processes, such as the settlement and resuspension of insoluble substances, the adsorption and desorption of pollutants between the sediment and the water phase, and the degradation and conversion under the action of microorganisms. These multi-phase conversion behaviors significantly change the migration path, existence form and ecological toxicity of pollutants. The downstream often distributes sensitive targets such as urban drinking water source protection area, important wetland or aquatic germplasm resource protection area, which have high requirements for water quality safety and ecological health. In the face of such crisis, how to quickly and effectively intercept and control pollution and reduce pollution to minimize environmental impact and economic loss has become a severe challenge for the basin management.
[0003] At present, the emergency regulation and control for sudden water pollution events in the gated river section mainly relies on single gate operation based on fixed rules or passive response mode after the event. The existing technology usually closes the downstream gate to block the pollution group from discharging or concentrates the discharge of the upstream gate to dilute the pollution. This disposal method relying on experience judgment is difficult to cope with the complex multiphase conversion of pollutants during migration, and cannot realize the coordinated operation of multiple gates in time and space. On the other hand, although the mechanism model with hydrodynamic and water quality coupling can theoretically simulate the migration and transformation process of pollutants, these models usually have high computational complexity and long time-consuming, which cannot meet the real-time decision-making demand of minutes or hours in emergency response. In addition, in the complex river network across regions, each gate usually belongs to different management subjects, and there is a lack of an intelligent decision support platform that can integrate multi-source real-time monitoring data, quickly predict the evolution trend of pollution group, and automatically generate a global optimal coordinated scheduling scheme. This leads to the difficulty in forming scientific, efficient and unified regulation and control instructions in the valuable emergency response window period, which may directly threaten the safety of downstream water supply, cause regional ecological disasters, and cause huge social and economic losses. Therefore, developing an intelligent system that can quickly assimilate real-time data, accurately predict the multiphase transformation process, and automatically generate a multi-gate coordinated emergency regulation and control scheme has become an urgent need in the field of current watershed water resources management and water environment protection. SUMMARY
[0004] The present application provides a water quality multiphase transformation intelligent regulation and control system for gated river sections to solve the problems in the background art through a data assimilation and state updating module, a proxy model calculation module, a coordinated strategy decision module and a rolling optimization and instruction execution module.
[0005] The technical solution of the present application to solve the above technical problems is as follows: specifically includes: data assimilation and state updating module, proxy model calculation module, coordinated strategy decision module and rolling optimization and instruction execution module, wherein:
[0006] The data assimilation and state updating module is used to integrate the real-time hydrological and water quality data collected by the Internet of Things sensor network arranged on the gated river section when receiving new sensor monitoring data, and dynamically correct the state variables and key parameters inside the pre-constructed mechanism model by using the ensemble Kalman filtering algorithm, and output a high-fidelity digital twin containing the assimilated river network full-element state vector.
[0007] The agent model calculation module is configured to receive a river network full-element state vector output by the high-fidelity digital twin, and take the river network full-element state vector as an input to call a pre-trained agent model based on a spatiotemporal graph neural network with an attention mechanism, and deduce spatiotemporal evolution results of a concentration field distribution of a pollutant and a multiphase conversion state of the pollutant in a preset time period in the future and under different gate operation strategies.
[0008] The collaborative strategy decision module is configured to take the concentration field distribution and the spatiotemporal evolution results of the multiphase conversion state output by the agent model calculation module as a virtual environment, and run a centralized training multi-agent reinforcement learning decision model to generate a set of collaborative gate opening adjustment action sequences for all the gates to maximize a global reward function.
[0009] The rolling optimization and instruction execution module is configured to select a gate control instruction of a first time step from the collaborative gate opening adjustment action sequences, and send the gate control instruction to a corresponding physical gate dam actuator for execution, and trigger the data assimilation and state updating module to start a new round of state updating after waiting for a sampling period, so as to form a closed-loop feedback control process.
[0010] In a preferred embodiment, the data assimilation and state updating module integrates real-time hydrological and water quality data collected by an Internet of Things sensor network arranged on a gate-controlled river section, and dynamically corrects state variables and key parameters in a pre-constructed mechanism model by using a set Kalman filtering algorithm. Specifically, the data assimilation and state updating module performs the following operations:
[0011] First, the Internet of Things sensor network continuously collects real-time hydrological and water quality data including water level, flow rate, pollutant concentration, dissolved oxygen, and conductivity at key sections of the gate-controlled river section, and performs outlier detection and smoothing filtering preprocessing on the collected real-time hydrological and water quality data to form a standardized current time observation vector.
[0012] Next, a river network full-element state vector contained in a high-fidelity digital twin output in a previous data assimilation cycle is taken as an initial state input to a pre-constructed mechanism model, and the mechanism model is a partial differential equation model coupled with hydrodynamic, water quality, and pollutant multiphase conversion processes. The mechanism model is driven to perform dynamic prediction for one time step, and the dynamic prediction process takes the river network full-element state vector at a previous time and external forcing terms at a current time as inputs to calculate a model prediction state vector at the current time by using model operators, wherein the external forcing terms include gate opening instructions and boundary condition inputs of the model.
[0013] Then, a state set of model predicted state vectors around the current time is generated, a mean and a covariance matrix of the state set are calculated; a Kalman gain matrix is calculated based on the covariance matrix, a preset observation operator matrix and an observation error covariance matrix, wherein the calculation method of the Kalman gain matrix is: multiplying the covariance matrix of the state set and the transpose of the observation operator matrix to obtain a first intermediate matrix; multiplying the observation operator matrix, the covariance matrix of the state set and the transpose of the observation operator matrix, and adding the observation error covariance matrix to obtain a second intermediate matrix; multiplying the inverse matrix of the first intermediate matrix and the second intermediate matrix to finally obtain the Kalman gain matrix;
[0014] Finally, the model predicted state vector is corrected by using the Kalman gain matrix, and the correction method is: multiplying the Kalman gain matrix and the difference between the observation vector and the predicted observation vector calculated by the observation operator matrix and the model predicted state vector to obtain a correction amount; adding the correction amount to the model predicted state vector to output the assimilated river network full element state vector.
[0015] In a preferred embodiment, the output includes a high-fidelity digital twin of the assimilated river network full element state vector, and the specific process is:
[0016] The assimilated river network full element state vector is encapsulated with the fixed structure parameters in the pre-constructed mechanism model that are not changed by the assimilation process and the timestamp information of the assimilation process to jointly constitute the high-fidelity digital twin.
[0017] In a preferred embodiment, in the proxy model calculation module, the river network full element state vector output by the high-fidelity digital twin is received and used as input.
[0018] First, the received river network full element state vector is parsed, and the hydrological and water quality data parameters at each spatial discrete node included in the state vector are respectively mapped to the corresponding nodes of a space-time graph structure according to the pre-defined river network spatial topology structure, so as to construct a graph structure instance at the current time; wherein the feature data of each node is composed of the water level value, the flow rate value, the concentration values of multiple pollutants in different phases, and the water temperature value; the edges in the space-time graph structure are used to describe the connection relationship between the nodes, and contain the length and flow direction weight information of the river section.
[0019] In a preferred embodiment, the specific process of calling the pre-trained attention mechanism based space-time graph neural network proxy model for deduction is:
[0020] The constructed current time graph structure instance is input into the pre-trained agent model, and the model performs forward propagation calculation through the internal multi-layer network structure;
[0021] The forward propagation calculation includes the following consecutive operations at each layer:
[0022] First, spatial attention aggregation is performed, that is, the interaction weight between each node in the dynamic calculation graph and its adjacent nodes is dynamically calculated, and the information from the adjacent nodes is weighted and fused according to the weight; then, the space-time state update is performed, that is, the weighted and fused spatial information is integrated with the current state information of the node itself by using a gated recurrent unit, to update the state description of the node.
[0023] After the layer-by-layer transmission and transformation of the multi-layer network, the agent model finally maps the node feature vector of the final layer to the inference result of each spatial node at a plurality of preset future time steps through a decoder network; the inference result constitutes a sequence including future time series data, and each time data in the sequence includes water level, flow rate, and concentration data of various pollutants in different phases at all spatial nodes.
[0024] In a preferred embodiment, in the collaborative strategy decision module, based on the current time river network full-element state vector, and taking the concentration field distribution and the multi-phase transformation state space evolution result output by the agent model calculation module as the specific operation of the virtual environment:
[0025] First, based on the current time river network full-element state vector obtained from the data assimilation and state update module, and the concentration field distribution and the multi-phase transformation state space evolution result of a plurality of future time steps obtained from the agent model calculation module, a simulation environment for multi-agent reinforcement learning decision is constructed; the environment can simulate the evolution of the river network system state from the current time to a plurality of future time steps; on this basis, each dam is instantiated as an independent agent, and a local observation space is defined for each agent, which includes the hydrological and water quality state data of the upstream adjacent nodes, the node where the dam is located, and the downstream adjacent nodes of the agent corresponding to the dam, and the current opening state information of the dam itself.
[0026] In a preferred embodiment, the specific operation of running the centralized trained multi-agent reinforcement learning decision model is:
[0027] The architecture adopts centralized training and decentralized execution, and a centralized critic network is arranged inside the decision model, an input of the network includes a joint observation vector spliced by local observation information of all gate intelligent agents and a joint action vector composed of actions of all intelligent agents, and an output of the network is a scalar value representing a global value function value under the joint observation and the joint action.
[0028] The centralized critic network internally includes a nonlinear value function decomposition structure, which obtains the global value function value by multiplying a local value function output of each intelligent agent with a corresponding weight coefficient and then summing and passing through a nonlinear mapping function.
[0029] The weight coefficient corresponding to each intelligent agent is calculated by an attention mechanism subunit, the attention mechanism subunit takes the joint observation vector as an input, transforms and normalizes the local observation of each intelligent agent by using a trainable parameter matrix and vector, dynamically generates the weight coefficient of each intelligent agent, and adjusts the contribution degree of the local value function of each intelligent agent to the global value function.
[0030] Meanwhile, a decentralized actor network is arranged for each gate intelligent agent, each actor network is a parameterized policy function, an input of the actor network is local observation information of the intelligent agent, and an output of the actor network is a probability distribution of executable actions of the intelligent agent; and training of the actor network adopts a proximal policy optimization algorithm.
[0031] In a preferred embodiment, the specific process of generating a set of coordinated gate opening adjustment action sequences for all gates to maximize the global reward function is as follows:
[0032] The global reward function is composed of three weighted sub-items, the first sub-item is a water quality safety reward item, which describes the water quality safety situation by calculating the sum of squares of pollutant concentration exceeding amounts at all sensitive monitoring points;
[0033] The second sub-item is an ecological protection reward item, which describes the ecological impact by using water body dissolved oxygen level and biological toxicity constraints;
[0034] The third sub-item is an operation cost reward item, which describes the control cost by using gate opening change amplitude and frequency;
[0035] The total reward value is obtained by adding the sub-items multiplied by the corresponding weight coefficients, the long-term cumulative total reward value is maximized by iterative optimization, and finally a set of coordinated gate opening adjustment action sequences of all gates in future multiple time steps is output.
[0036] In a preferred embodiment, the rolling optimization and instruction execution module selects the first time step gate control instruction from the cooperative gate opening adjustment action sequence and issues it to the corresponding physical gate actuator, which specifically includes the following steps:
[0037] First, the cooperative strategy decision module generates a cooperative gate opening adjustment action sequence, and extracts the instruction set corresponding to the current decision period, i.e., the first time step, which contains all the gate opening adjustment amounts.
[0038] Next, each opening adjustment amount in the instruction set is subjected to physical safety and executability verification, which includes amplitude verification processing and action direction verification processing. The amplitude verification processing compares each opening adjustment amount with the maximum and minimum opening limits allowed by the corresponding gate actuator. For adjustment amounts less than the minimum opening limit, it is set to the minimum opening limit. For adjustment amounts greater than the maximum opening limit, it is set to the maximum opening limit. For adjustment amounts within the limit range, the original value is maintained. After verification, the verified instruction is issued to the corresponding physical gate actuator through an industrial real-time communication protocol. At the same time, a timeout monitoring timer is started, with the waiting time set to the smaller of the system sampling period and a preset instruction execution timeout threshold. Within this waiting time, the status feedback signals from each physical gate actuator are continuously monitored and received to verify whether the instruction has been correctly executed.
[0039] In a preferred embodiment, after waiting for a sampling period, the data assimilation and state update module triggers the start of a new round of state update, with the specific operation being:
[0040] When the sampling period timer ends, the rolling optimization and instruction execution module immediately sends a trigger signal to the data assimilation and state update module. This trigger signal starts a new round of data assimilation process, using the latest sensor observation data to correct the state of the digital twin, producing an updated river network full-element state vector. This updated river network full-element state vector is set as the initial condition for the next decision period optimization problem. At the same time, the time window of the prediction time domain is rolled forward by one sampling period. In addition, the rolling optimization and instruction execution module also evaluates the control performance by calculating the weighted integral of the deviation between the actual and expected values of the water quality index over a period of time, and dynamically adjusts the weight parameters in the global reward function according to the trend of the performance evaluation results, to achieve adaptive optimization of parameters. This updated river network full-element state vector will be used as the starting point for the next decision period and sent to the agent model calculation module to start a new round of prediction, decision, and rolling optimization cycle, thus forming a complete, real-time feedback-based closed-loop intelligent control process.
[0041] The beneficial effects of the present application are: by constructing a high-fidelity digital twin and fusing real-time data assimilation technology, the precise dynamic perception of the multi-phase conversion process of the water quality of the gated river section is realized; with the help of the spatio-temporal graph neural network agent model based on the attention mechanism, the system can quickly predict the complex spatio-temporal pattern of the pollutant concentration field and phase evolution under different control strategies; on this basis, the multi-agent reinforcement learning algorithm is used to automatically generate the globally optimal coordinated control scheme of the gate group, effectively coordinating the multiple goals of water quality improvement, ecological protection and engineering operation; finally, through the rolling optimization and closed-loop execution mechanism, the theoretical strategy is reliably converted into physical action, and can be adjusted adaptively according to the environmental feedback, thereby comprehensively improving the emergency response speed, the scientificity of the control decision and the intelligent level of the whole river network water quality management of the sudden water pollution event. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 The method flowchart of the present application is shown in the figure.
[0043] Figure 2 The system structure block diagram of the present application is shown in the figure. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0045] In the description of the present application, the terms "first", "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0046] In the description of the present application, the term "for example" is used to mean "serving as an example, instance, or illustration." Any embodiment described as "for example" in this application is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is presented to enable any person skilled in the art to make and use the application. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present application. It will be apparent, however, to one skilled in the art that the present application can be practiced without using these specific details. In other instances, well-known structures and processes are not elaborated upon in order to avoid unnecessary detail, which can obscure the description of the present application. Thus, the present application is not intended to be limited by the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0047] Embodiment 1
[0048] The present embodiment provides a kind of intelligent control system for multiphase conversion of water quality in gated river section as shown in Figures 1-2 Specifically includes: data assimilation and state updating module, agent model calculation module, collaborative strategy decision module and rolling optimization and instruction execution module, wherein;
[0049] Data assimilation and state updating module, for when system starts or receives new sensor monitoring data, by integrating the real-time hydrological water quality data collected by the Internet of Things sensor network arranged on the gated river section, and adopting ensemble Kalman filter algorithm to dynamically correct the state variables and key parameters inside the mechanism model constructed in advance, output high-fidelity digital twin containing assimilated river network full-element state vector;
[0050] Agent model calculation module, for receiving the river network full-element state vector output by the high-fidelity digital twin, and taking it as input, calling the pre-trained agent model based on attention mechanism spatio-temporal graph neural network, to millisecond level speed deduce the concentration field distribution and multiphase conversion state spatio-temporal evolution result in the future preset period corresponding to different gate operation strategies;
[0051] Collaborative strategy decision module, for based on the river network full-element state vector at the current time, and taking the concentration field distribution and multiphase conversion state spatio-temporal evolution result output by the agent model calculation module as a virtual environment, running the multi-agent reinforcement learning decision model trained centrally, to generate a set of collaborative gate opening adjustment action sequence for all dams that maximizes the global reward function;
[0052] The rolling optimization and instruction execution module is used to select the first time step of the gate control instruction from the coordinated gate opening adjustment action sequence within the current decision cycle and issue it to the corresponding physical gate actuator. After waiting for one sampling cycle, it triggers the data assimilation and state update module to start a new round of state update, thereby forming a closed-loop feedback control process.
[0053] In this embodiment, it is specifically necessary to explain that in the data assimilation and state update module, real-time hydrological and water quality data collected by the Internet of Things sensor network set up on the sluice gate section are integrated, and the ensemble Kalman filter algorithm is used to dynamically correct the state variables and key parameters inside the pre-built mechanism model. The specific operation is as follows:
[0054] First, the IoT sensor network continuously collects real-time hydrological and water quality data, including water level, flow velocity, pollutant concentration, dissolved oxygen, and conductivity, at key sections of the sluice gate-controlled river section. The collected real-time hydrological and water quality data undergoes outlier detection (i.e., removing outliers that clearly exceed physical limits) and smoothing filtering preprocessing to form a standardized observation vector for the current moment. This observation vector is an mwq-dimensional column vector, where mwq represents the total number of observed variables, and each observation moment is indexed by a discrete-time index. Mark it;
[0055] Next, the state vector of all river network elements contained in the high-fidelity digital twin output from the previous data assimilation cycle is used as the initial state input to a pre-constructed mechanistic model. This mechanistic model is a partial differential equation model coupling hydrodynamics, water quality, and multiphase transformation processes of pollutants. This mechanistic model is then driven to perform dynamic forecasting for one time step. The dynamic forecasting process utilizes model operators, taking the state vector of all river network elements from the previous time step and the external forcing terms from the current time step as inputs, to calculate the model's predicted state vector for the current time step. The external forcing terms include the opening commands of each gate and the boundary conditions input of the model, as shown in the formula:
[0056] ;
[0057] in, This represents the state vector describing the entire gate-controlled river system. It is a set containing the values of all variables of interest to the model at each spatial location in the river network (water level, flow rate, concentration of pollutants in various phases, etc. in each river segment; for example, it includes the water level of segment 1, the flow velocity of segment 2, and the COD concentration at a certain cross section, etc.). The superscript indicates... and This represents a time index, used to distinguish the state at different times. Indicates the current moment. Indicates the previous time period, with the subscript... is the abbreviation of English "prior", which means the result of the mechanism model prediction, not yet combined with the actual observation data at the current time, represents the state of the river network at time, which is the result of the mechanism model running, represents the predicted value of the state at time, which is the direct result of the mechanism model running, represents the best estimate of the state of the real system at time, which is the output of the high-fidelity digital twin containing the state vector of all elements of the river network after optimization by the data assimilation algorithm (such as ensemble Kalman filter), which combines the model prediction at an earlier time and the actual observation data at represents the mechanism model itself (i.e. the model operator), which is not a simple multiplication, but a complex mathematical operation process, which encapsulates the physical and chemical laws describing the movement of water flow, the migration and diffusion of pollutants, and the phase transformation (usually derived from the discretization of partial differential equations), represents the known external influence or artificial control input (i.e. external forcing term) applied to the system within the time interval ;
[0058] Then, a state ensemble centered on the model prediction state vector at the current time is generated, each member of the state ensemble is obtained by perturbing the parameters or initial conditions of the mechanism model, and the mean and covariance matrix of the state ensemble are calculated; based on the covariance matrix, a pre-set observation operator matrix and an observation error covariance matrix, a Kalman gain matrix is calculated, wherein the calculation method of the Kalman gain matrix is: multiplying the covariance matrix of the state ensemble with the transpose of the observation operator matrix to obtain a first intermediate matrix; multiplying the observation operator matrix, the covariance matrix of the state ensemble and the transpose of the observation operator matrix, and adding the observation error covariance matrix to obtain a second intermediate matrix; multiplying the first intermediate matrix with the inverse matrix of the second intermediate matrix to finally obtain the Kalman gain matrix, and the calculation formula of the Kalman gain matrix is:
[0059] ;
[0060] wherein, represents the Kalman gain matrix, which determines the degree to which the observation data should be trusted relative to the model prediction when correcting the model prediction, P is the forecast error covariance matrix, which is calculated from the state ensemble, quantifying the uncertainty of the model forecast states and their spatial correlation, H is the observation operator matrix, a linear operator mapping the state space of the model to the observation space, essentially indicating which state variables in the model correspond to each observation and their location (e.g., an observation from a certain sensor corresponds to the state value at a specific point in the model grid), the superscript "T" denotes the matrix transpose operator, used to transpose a matrix (rows become columns and vice versa), for example, is the transpose matrix of , the superscript "-1" denotes the matrix inverse operator, used to invert a matrix, R is the observation error covariance matrix, a diagonal or approximately diagonal matrix containing information about the measurement error of each observation instrument and the representative error of the observations;
[0061] Finally, the model forecast state vector is corrected using the Kalman gain matrix. The correction is done by multiplying the Kalman gain matrix with the difference between the observation vector and the forecast observation vector calculated from the observation operator matrix and the model forecast state vector, resulting in a correction term. The correction term is then added to the model forecast state vector, outputting the assimilated river network full state vector, which is given by:
[0062] ;
[0063] where, x is the assimilated river network full state vector, which is the final output of the data assimilation process, it is the result of the fusion of model forecasts and observation data according to the optimal weight determined by the Kalman gain, and it is the optimal estimate of the true state, x is the model forecast state vector (dimension is a column vector), which is directly calculated from the mechanism model and has not been corrected by observation data, where n represents the total number of state variables of the model (such as water level, flow rate, concentration of each pollutant, etc.), K is the Kalman gain matrix, y is the observation vector (dimension is a column vector), which is the observation data actually collected and preprocessed by the sensor network at the current time , where n represents the number of observation data, H is the observation operator matrix, which functions to optimally incorporate observation information into model forecasts by minimizing analysis error covariance, correcting state variables (such as water level, concentration) and key parameters (such as diffusion coefficient, settling rate);
[0064] The high-fidelity digital twin output contains the assimilated river network full-element state vector, and the specific process is as follows:
[0065] The assimilated river network full-element state vector is encapsulated with the fixed structure parameters in the pre-constructed mechanism model that are not changed by the assimilation process and the timestamp information of the assimilation process to form a high-fidelity digital twin; wherein the assimilated river network full-element state vector contains the water level, flow velocity at spatial discrete points, and concentration data of pollutants in different phases of the gated river network, and this high-fidelity digital twin as a structured data object provides unique and authoritative input data reflecting the real-time dynamics of the river network for the subsequent agent model calculation module in the system.
[0066] In this embodiment, it is specifically required to explain that in the agent model calculation module, the river network full-element state vector output by the high-fidelity digital twin is received and used as input, and the specific operation is as follows:
[0067] First, the received river network full-element state vector is parsed, and according to the pre-defined river network spatial topology, the hydrological and water quality data parameters of each spatial discrete node contained in the state vector are respectively mapped to the corresponding nodes of a space-time graph structure, thereby constructing a graph structure instance at the current time; wherein the feature data of each node is composed of the water level value, flow velocity value, concentration values of multiple pollutants in different phases (including dissolved phase, particle adsorption phase, etc.), and water temperature value these physical parameters; the edges in the space-time graph structure are used to describe the connection relationship between the nodes, and contain the length and flow direction weight information of the river section;
[0068] The specific process of calling the pre-trained agent model based on attention mechanism space-time graph neural network for deduction is as follows:
[0069] The constructed graph structure instance at the current time is input into the pre-trained agent model, and the model performs forward propagation calculation through its internal multi-layer network structure;
[0070] The forward propagation calculation includes the following continuous operations at each layer:
[0071] First, spatial attention aggregation is performed, that is, the interaction weight between each node and its adjacent nodes in the dynamic calculation graph is dynamically calculated, and the information from the adjacent nodes is weighted and fused according to the weight, and the formula is as follows:
[0072] ;
[0073] Wherein, represents the attention weight, that is, in the lth layer network, the node When aggregating neighbor information, it is assigned to its neighbor nodes. Importance weight; the larger the value, the stronger the node. For nodes The more important the influence, the more significant the impact. 'l' represents the network layer index, used to identify the current computation occurring in the 'l'th layer of the neural network. The entire network contains L layers. This represents the index of the target node, which is the center node whose feature vector is currently being updated. This represents the neighbor node index, i.e., the target node. A directly adjacent node, This represents the exponential function, used to convert the calculated score into a positive number for normalization. This represents the activation function, a non-linear function (such as LeakyReLU) used to introduce a non-linear transformation into the attention mechanism, enhancing the model's expressive power. This represents the learnable parameter vector in the attention mechanism. It's a vector that maps the concatenated high-order features to a scalar "attention score," indicated by the superscript "". " represents the transpose operator, that is, to transpose... Transpose. Let represent the learnable weight matrix of the l-th layer, used to perform a linear transformation on the feature vectors of the nodes, projecting them into a new high-dimensional space that is more suitable for calculating the correlations between nodes. Represents the node at level l The feature vector, that is, the node when input to the l-th layer. The feature representation, for the first layer (l=1), this vector is the original input feature. Represents the node at level l The feature vectors of the nodes neighboring nodes Feature representation at layer l, This represents the vector concatenation operation, used to join two vectors end-to-end to form a longer vector. Represents a node The set of neighboring nodes, in a graph structure, is all nodes connected to the node. The set of directly connected nodes. This represents the summation index, which, in the summation terms of the denominator, represents the traversed nodes. All neighboring nodes Subsequently, a spatiotemporal state update is performed, which involves using a gated loop unit to integrate the weighted and fused spatial information with the node's own current state information to update the node's state description. The spatiotemporal state update formula is as follows:
[0074] ;
[0075] ;
[0076] where, denotes the aggregated message, i.e., the node aggregates information from all its neighbor nodes at the l-th layer, which is a weighted sum of the transformed features of the neighbor nodes, with the weights being the attention weights , denotes the weight matrix for message transformation, a learnable parameter matrix, which is specifically used to transform the features of the neighbor nodes before aggregating the information, (remark: in some architectures, and are the same matrix), denotes the gated recurrent unit, a recurrent neural network structure, which here serves as an “update function” to update the state of the node itself based on the aggregated message , thus capturing the dynamic characteristics in both space (neighbor messages) and time (history of the node’s own state), denotes the updated node feature, which is the new feature representation of the node at the (l+1)-th layer after the GRU computation, which will serve as the input to the next layer of the network, denotes the internal parameters of the GRU unit, containing all the parameter sets that need to be trained in the GRU structure, such as the update gate and the reset gate.
[0077] After the layer-by-layer transmission and transformation through the multi-layer network, the agent model finally maps the node feature vector at the final layer to the inference results of each spatial node at a preset number of future time steps through a decoder network; the inference results constitute a sequence containing future time series data, and each time point in the sequence contains water level, flow rate, and concentration data of various pollutants in different phases at all spatial nodes; the data structure and physical dimension of the inference result sequence are consistent with the spatial discrete node number and physical dimension used in the data assimilation and state update module, so that the subsequent collaborative strategy decision module can directly call it as virtual environment data.
[0078] In this embodiment, it is specifically necessary to explain that, in the collaborative strategy decision module, based on the river network full-element state vector at the current time, and taking the concentration field distribution and the spatiotemporal evolution results of the multiphase transformation state output by the agent model calculation module as the specific operation of the virtual environment:
[0079] First, based on the current state vector of all elements of the river network obtained from the data assimilation and state update module, and the concentration field distribution and spatiotemporal evolution results of multiphase transformation states at multiple future time steps obtained from the surrogate model calculation module, a simulation environment for multi-agent reinforcement learning decision-making is constructed. This environment can simulate the state evolution of the river network system at multiple future time steps starting from the current moment. On this basis, each dam in the control system is instantiated as an independent agent, and a local observation space is defined for each agent. This local observation space contains the hydrological and water quality status data of the upstream neighboring nodes of the dam, the node where the dam is located, and the downstream neighboring nodes of the dam, as well as the current opening status information of the dam itself.
[0080] The specific steps for running a centralized, multi-agent reinforcement learning decision-making model are as follows:
[0081] The architecture of centralized training and decentralized execution is adopted. Within the decision model, a centralized commentator network is set up. The input of this network includes a joint observation vector composed of the local observation information of all dam agents and a joint action vector composed of the actions of all agents. The output of the network is a scalar value representing the global value function value under joint observation and joint action.
[0082] The centralized commentator network contains a nonlinear value function decomposition structure. This structure obtains the global value function by multiplying the local value function output of each agent by a corresponding weight coefficient, summing the results, and then passing the sum through a nonlinear mapping function. The formula for the nonlinear value function decomposition structure is as follows:
[0083] ;
[0084] in, The global state-action value function is represented by a scalar, which indicates the state at time t. All intelligent agents take joint action Finally, the total expected cumulative reward value can be obtained. Its function is: it is the final output of the centralized critic network and is used to guide the agent in learning the optimal policy. This represents a nonlinear mapping function, typically a nonlinear activation function (such as ReLU) or a small neural network. Its function is to perform a nonlinear transformation on the weighted sum, enhancing the model's expressive power. The total number of intelligent agents (dams) is a constant. Index representing the agent, From 1 to , Indicates the first The attention weight of each agent, a scalar, serves to dynamically measure the attention weight of the first agent. The importance of each agent's contribution to its local Q-value in the current global state. Represents the joint observation vector, which is generated by all agents at time [time]. Local observation It is pieced together to represent the overall state of the environment. Indicates the first The local value function of an agent, a scalar, represents the value at time t. The local value of an action taken by the agent based on its own observations; its function: calculated by the subnetwork corresponding to each agent. Indicates the first An intelligent agent at time The local observation is a vector containing environmental information that the agent can perceive (such as upstream and downstream water levels, water quality, and its own opening degree). Indicates the first An intelligent agent at time Actions taken (such as changes in gate opening). Indicates the first Local value function of an agent The network parameters are updated during training;
[0085] The weight coefficients for each agent are calculated through an attention mechanism subunit. This subunit takes the joint observation vector as input and transforms and normalizes the local observations of each agent using a trainable parameter matrix and vectors, dynamically generating the weight coefficients for each agent. This adjusts the contribution of each agent's local value function to the global value function, as shown in the formula:
[0086] ;
[0087] in, This represents the normalization exponential function, which normalizes the "attention scores" calculated by all agents, ensuring that the sum of all weights is 1, and that each weight is between 0 and 1. The query vector representing the attention mechanism is a trainable column vector. Its function is to act as a general "questioner," performing a dot product with the transformed observations of each agent to calculate similarity. The superscript "" indicates this. " indicates the transpose symbol, that is, the vector Convert column vectors to row vectors to perform dot product operations. This represents the weight matrix used to transform observations in the attention mechanism. It is a trainable matrix whose function is to transform the local observations of each agent. Linear projection onto a new feature space is used to extract key information for calculating attention. Indicates the first An agent observes the local state at time , represents the hyperbolic tangent activation function, which compresses the linear transformed result to the interval (-1, 1) to prevent gradient explosion and introduces nonlinearity, represents the dimension of the key vector, i.e., the weight matrix represents the dimension of the output feature, which is also a hyperparameter, and acts as a scaling factor in the formula to scale the dot product result to prevent gradient vanishing of the softmax function;
[0088] At the same time, a decentralized actor network is set for each gate agent, and each actor network is a parameterized policy function whose input is the local observation information of the agent and whose output is the probability distribution of the executable action of the agent (i.e., the adjustment amount of the gate opening). The training of the actor network uses the proximal policy optimization algorithm, which considers both the importance sampling ratio and the advantage function estimate in the optimization objective function, and uses the clipping mechanism to constrain the amplitude of policy update to ensure the stability of the training process. The optimization objective function formula of the proximal policy optimization algorithm is:
[0089] ;
[0090] where, represents the objective function, which is the expected value that the PPO algorithm needs to maximize, represents the parameters of the policy network (actor network), and the goal of training is to adjust to make as large as possible, represents the expectation operator, which needs to average the calculation in the parentheses (i.e., the min(...) part) over all possible states, actions, and experience data to obtain a stable optimization direction, represents the importance sampling ratio, which measures the probability difference between the new policy and the old policy when taking a certain action, and is calculated as: , represents the probability of selecting action by agent according to its current (new) policy network parameters after observing state at time , represents the probability of selecting the same action by agent according to its previous (old) policy network parameters (i.e., the policy used when generating the experience data currently being evaluated) under the same observation at time This is used to correct the empirical data collected from the old strategy, making it suitable for evaluating the performance of the new strategy. At that time, the old and new strategies were similar; when When, the new strategy is more inclined to choose this action; when At other times, the opposite is true. The dominant function represents the dominance of a particular observation. Next, perform a specific action. How much better is it than performing the average action? This indicates that the action is better than average and should be encouraged. This indicates that the action is below average and should be suppressed. This is the core pruning mechanism of PPO, in which This represents the clipping function, which sets the importance sampling ratio. Limited to the range Inside, This indicates the operation of finding the minimum value. Indicates the target that has not been clipped. , The target after cropping This mechanism collectively ensures the stability of policy updates. By taking the minimum value, it ultimately optimizes a "pessimistic" estimate, thus avoiding the risk of errors in a single update. The training crashes because the difference between the new and old strategies is too large. This represents the clipping range hyperparameter, a pre-defined small positive number (such as 0.1 or 0.2), which defines the trust domain for the policy update step size;
[0091] The generalized advantage estimation formula is:
[0092] ;
[0093] in, This represents the summation over future steps, starting from the current kpu and continuing until the next future step. step( This represents the total number of time steps in a round, or the maximum forward step size (truncated step size) used to calculate advantage. It is the offset relative to the current KPU. This represents the composite discount factor, which consists of two hyperparameters: This represents the standard discount factor, with a value range of [0,1), used to measure the current value of future rewards. The closer the value is to 0, the more the agent focuses on short-term rewards; the closer it is to 1, the more farsighted it is. This represents a smoothing parameter specific to GAE, with a value range of [0,1], used to balance the bias and variance of the dominance estimate. Large time variance and small bias (closer to Monte Carlo estimation). Small time variance and large bias (closer to time series difference estimation). Indicates at time Instant rewards received This represents the state-value function, i.e., the state value function in the observation. Below is an estimate of the expected cumulative return that can be obtained by following the current strategy. The time-difference residual is considered as a single-step advantage estimate, i.e.: in At any given moment, the actual reward received is added to the estimated value of the next state, and then the estimated value of the current state is subtracted. If this value is positive, it means that the performance in the current step is better than expected.
[0094] The specific process for generating a set of coordinated gate opening adjustment action sequences that maximizes the global reward function and applies to all dams is as follows:
[0095] The global reward function consists of three weighted sub-items. The first sub-item is the water quality safety reward, which describes the water quality safety status by calculating the sum of squares of pollutant concentration exceedances at all sensitive monitoring points. The expression is:
[0096] ;
[0097] in, This indicates a sub-reward item for water quality safety. The "-" sign represents a negative value, used to make the summation result negative, because the term within the parentheses (the square of the excess) is always non-negative, and adding a negative sign means... The value is always ≤0. It is only 0 when all indicators are within the limit (no penalty); once an indicator exceeds the limit, a negative reward (i.e., penalty) is obtained. The summation symbol represents the summation of the sums of the two sides. From 1 to The summation is performed to accumulate the penalties for exceeding all monitored water quality indicators. This indicates the total number of water quality indicators monitored in the system. For example, indicators may include chemical oxygen demand (COD), ammonia nitrogen (NH3-N), and heavy metal concentrations. Indicates in At that moment, the The actual measured or predicted concentration values of the water quality indicators (this data comes from the deduction results output by the surrogate model calculation module) are the environmental feedback input for reward calculation, directly reflecting the effect of control actions on water quality. Indicates the first The preset safe concentration threshold for a water quality indicator is a constant used as a benchmark to determine whether water quality is safe. This threshold is usually determined based on national water quality standards, drinking water source protection requirements, or ecological benchmarks. represents a max function, compares 0 and , and returns the larger one, which is used to realize the logic of "punish only over-standard, and do not reward standard", if (the standard is met), then , the max function takes 0, meaning no punishment for this indicator; if (the standard is not met), then , the max function takes a positive value, meaning that this indicator will be punished, represents a square operation, which has the following effects: A1, ensures non-negativity: makes the punishment value always positive; A2, amplifies serious over-standard: when the concentration is far beyond the safety threshold, the square operation will make the punishment value grow in a square level, which can strongly guide the agent to preferentially avoid serious pollution events, rather than treating slight over-standard and serious over-standard as the same;
[0098] The second sub-item is the ecological protection reward item, which describes the ecological impact by means of water body dissolved oxygen level and biological toxicity constraints;
[0099] The third sub-item is the operation cost reward item, which describes the control cost by means of gate opening change amplitude and frequency;
[0100] After each sub-item is multiplied by the corresponding weight coefficient, the total reward value is obtained by adding them up, and the long-term cumulative total reward value is maximized through iterative optimization, and finally a set of coordinated gate opening adjustment action sequences of all gates in the future multiple time steps is output.
[0101] In this embodiment, it is specifically necessary to explain that in the rolling optimization and instruction execution module, the first time step gate control instruction is selected from the coordinated gate opening adjustment action sequence and issued to the corresponding physical gate actuator for execution, which specifically includes the following steps:
[0102] First, the coordinated gate opening adjustment action sequence generated by the coordinated strategy decision module is analyzed, and the instruction set containing the opening adjustment amount of all gates corresponding to the current decision period, i.e. the first time step, is extracted;
[0103] Then, the physical safety and executability of each opening adjustment amount in the instruction set is checked, including a limiting check process and an action direction check process. The limiting check process is to compare each opening adjustment amount with the maximum opening limit value and the minimum opening limit value allowed by the corresponding gate dam actuator, and set the adjustment amount to the minimum opening limit value if it is less than the minimum opening limit value, set the adjustment amount to the maximum opening limit value if it is greater than the maximum opening limit value, and keep the original value if it is within the limit value range. The action direction check process is to ensure that the mathematical sign of the adjustment amount is consistent with the action direction allowed by the gate mechanical characteristics. After the check passes, the checked instructions are respectively sent to the corresponding physical gate dam actuators through an industrial real-time communication protocol, which includes an OPC unified architecture protocol or a Modbus transmission control protocol. At the same time of sending the instructions, a timeout monitoring timer is started, and the waiting time of the timer is set to the smaller value between the system sampling period and a preset instruction execution timeout threshold. Within the waiting time, the state feedback signals from each physical gate dam actuator are continuously monitored and received to verify whether the instructions are correctly executed.
[0104] After waiting for a sampling period, the data assimilation and state updating module triggers the specific operation of starting a new round of state updating as follows:
[0105] When the sampling period timing ends, the rolling optimization and instruction execution module immediately sends a trigger signal to the data assimilation and state updating module. The trigger signal starts a new round of data assimilation process, and uses the latest collected sensor observation data to correct the state of the digital twin, to generate an updated river network full-element state vector. The updated river network full-element state vector is set as the initial condition of the optimization problem in the next decision-making period. At the same time, the time window of the prediction time domain is rolled forward by one sampling period. In addition, the rolling optimization and instruction execution module also evaluates the control performance by calculating the weighted integral of the deviation between the actual value and the expected value of the water quality index within a period of time, and dynamically adjusts the weight parameters in the global reward function according to the trend of the performance evaluation result, to realize the adaptive optimization of the parameters. The updated river network full-element state vector will be used as the starting point of the next decision-making period and sent to the agent model calculation module to start a new round of prediction, decision-making and rolling optimization cycle, thereby forming a complete closed-loop intelligent regulation and control process based on real-time feedback.
[0106] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0107] Those skilled in the art will appreciate that embodiments of the present application can be devised for a variety of applications. It is intended that the present application be limited only by the scope of the appended claims, and it is intended that various modifications and alterations made by those skilled in the art be considered as within the scope of the present application. The embodiments of the present application will be described with reference to the attached drawings, wherein:
[0108] The present application is described in reference to the drawings using a flowchart and / or a block diagram of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0109] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0111] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to cover all such modifications and variations as fall within the scope of the present application.
[0112] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A gated river section water quality multi-phase conversion intelligent regulation system, characterized in that, Specifically comprising: The data assimilation and state updating module, the proxy model calculation module, the collaborative strategy decision module and the rolling optimization and instruction execution module, wherein The data assimilation and state updating module is used for, when receiving new sensor monitoring data, integrating real-time hydrological and water quality data collected by the Internet of Things sensor network arranged on the gated river section, and dynamically correcting state variables and key parameters in the mechanism model constructed in advance by using the ensemble Kalman filtering algorithm, outputting a high-fidelity digital twin containing the river network full-element state vector after assimilation; The proxy model calculation module is used for receiving the river network full-element state vector output by the high-fidelity digital twin, taking it as input, calling the pre-trained proxy model based on the attention mechanism spatio-temporal graph neural network, and deducing the concentration field distribution and multi-phase conversion state spatio-temporal evolution results of the pollutants in the future preset period corresponding to different gate operation strategies; The collaborative strategy decision module is used for, based on the river network full-element state vector at the current moment, taking the concentration field distribution and multi-phase conversion state spatio-temporal evolution results output by the proxy model calculation module as a virtual environment, running the multi-agent reinforcement learning decision model after centralized training, and generating a set of collaborative gate opening adjustment action sequences for all dams to maximize the global reward function; The rolling optimization and instruction execution module is used for selecting the gate control instruction of the first time step from the collaborative gate opening adjustment action sequence and issuing it to the corresponding physical dam actuator for execution in the current decision period, and after waiting for a sampling period, triggering the data assimilation and state updating module to start a new round of state updating, thereby forming a closed-loop feedback control process.
2. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 1, characterized in that: In the data assimilation and state updating module, real-time hydrological and water quality data collected by the Internet of Things sensor network arranged on the gated river section is integrated, and the state variables and key parameters in the mechanism model constructed in advance are dynamically corrected by using the ensemble Kalman filtering algorithm. The specific operation is as follows: First, the Internet of Things sensor network continuously collects real-time hydrological and water quality data including water level, flow rate, pollutant concentration, dissolved oxygen and conductivity at key sections of the gated river section, and performs outlier detection and smoothing filtering preprocessing on the collected real-time hydrological and water quality data to form a standardized current moment observation vector; Next, the river network full-element state vector contained in the high-fidelity digital twin output in the last data assimilation period is input as an initial state to the mechanism model constructed in advance, and the mechanism model is a partial differential equation model coupled with hydrodynamic, water quality and pollutant multi-phase conversion processes; the mechanism model is driven to perform dynamic prediction for one time step, and the dynamic prediction process is performed by using model operators to calculate the model prediction state vector at the current moment by taking the river network full-element state vector at the previous moment and the external forcing term at the current moment as input, wherein the external forcing term includes the opening instruction of each gate and the boundary condition input of the model; Then, a state set around the model prediction state vector at the current moment is generated, and the mean and covariance matrix of the state set are calculated; The Kalman gain matrix is calculated based on the covariance matrix, a preset observation operator matrix, and an observation error covariance matrix, and the calculation manner of the Kalman gain matrix is as follows: the covariance matrix of the state set is multiplied by the transpose of the observation operator matrix to obtain a first intermediate matrix; the observation operator matrix, the covariance matrix of the state set, and the transpose of the observation operator matrix are multiplied together, and the observation error covariance matrix is added to obtain a second intermediate matrix; and the first intermediate matrix is multiplied by the inverse matrix of the second intermediate matrix to finally obtain the Kalman gain matrix; Finally, the model predicted state vector is corrected by using the Kalman gain matrix, and the correction manner is as follows: the Kalman gain matrix is multiplied by the difference between the observation vector and a predicted observation vector calculated from the observation operator matrix and the model predicted state vector to obtain a correction amount; and the correction amount is added to the model predicted state vector to output the assimilated river network full-element state vector.
3. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 2, characterized in that: A high-fidelity digital twin containing the assimilated river network full-element state vector is output, and the specific process is as follows: The assimilated river network full-element state vector is encapsulated with the fixed structure parameters in the pre-constructed mechanism model that are not changed by the assimilation process and the timestamp information of the assimilation process to jointly constitute the high-fidelity digital twin.
4. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 3, characterized in that: In the proxy model calculation module, the river network full-element state vector output by the high-fidelity digital twin is received and used as input, and the specific operation is as follows: First, the received river network full-element state vector is parsed, and the hydrological and water quality data parameters at each spatial discrete node contained in the state vector are respectively mapped to the corresponding nodes of a space-time graph structure according to the pre-defined river network spatial topology structure, so as to construct a graph structure instance at the current time; wherein the feature data of each node is composed of the water level value, the flow rate value, the concentration values of multiple pollutants in different phases, and the water temperature value; and the edges in the space-time graph structure are used to describe the connection relationship between the nodes, and contain the length and flow direction weight information of the river section.
5. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 4, characterized in that: The specific process of calling the pre-trained attention mechanism-based space-time graph neural network proxy model for deduction is as follows: The constructed graph structure instance at the current time is input into the pre-trained proxy model, and the model performs forward propagation calculation through its internal multi-layer network structure; The forward propagation calculation includes the following continuous operations at each layer: First, spatial attention aggregation is performed, that is, the interaction weight between each node in the dynamic calculation graph and its adjacent nodes is dynamically calculated, and the information from the adjacent nodes is weighted and fused according to the weight; Then, the space-time state is updated, that is, the weighted and fused spatial information and the current state information of the node itself are integrated by using a gated recurrent unit to update the state description of the node. After the layer-by-layer transmission and transformation through the multi-layer network, the agent model finally maps the node feature vectors of the final layer to the inference results of each spatial node at a preset continuous number of future time steps through a decoder network; the inference results constitute a sequence containing future time series data, and the data at each time in the sequence contains water level, flow rate, and concentration data of various pollutants in different phases at all spatial nodes.
6. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 5, characterized in that: In the collaborative strategy decision module, based on the river network full-element state vector at the current time, and taking the concentration field distribution and the spatiotemporal evolution results of the multiphase conversion state output by the agent model calculation module as the specific operation of the virtual environment: First, based on the current time river network full-element state vector obtained from the data assimilation and state updating module, and the concentration field distribution and the spatiotemporal evolution results of the multiphase conversion state at a plurality of future time steps obtained from the agent model calculation module, a simulation environment for multi-agent reinforcement learning decision is constructed; this environment can simulate the evolution of the river network system state from the current time to a plurality of future time steps; on this basis, each dam is instantiated as an independent agent, and the local observation space of each agent is defined, which includes the hydrological and water quality state data of the upstream adjacent nodes, the node where the dam is located, and the downstream adjacent nodes of the dam corresponding to the agent, as well as the current opening state information of the dam itself.
7. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 6, characterized in that: The specific operation of running the multi-agent reinforcement learning decision model trained centrally is: Using the architecture of centralized training and decentralized execution, a centralized critic network is set up inside the decision model, the input of the network includes the joint observation vector spliced by the local observation information of all dam agents and the joint action vector composed of the actions of all agents, and the output of the network is a scalar value representing the global value function value under the joint observation and joint action; The centralized critic network contains a nonlinear value function decomposition structure inside, which obtains the global value function value by multiplying the local value function output of each agent with a corresponding weight coefficient and then passing through a nonlinear mapping function; Wherein, the weight coefficient corresponding to each agent is calculated by an attention mechanism subunit, which takes the joint observation vector as input, uses trainable parameter matrices and vectors to transform and normalize the local observation of each agent, and dynamically generates the weight coefficient of each agent to adjust the contribution of each agent's local value function to the global value function; At the same time, a decentralized actor network is set up for each dam agent, and each actor network is a parameterized policy function whose input is the local observation information of the agent and whose output is the probability distribution of the executable action of the agent; the training of the actor network uses the proximal policy optimization algorithm.
8. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 7, characterized in that: The specific process of generating a set of collaborative gate opening adjustment action sequences for all dams that maximize the global reward function is: The global reward function is composed of three weighted sub-items, the first sub-item is a water quality safety reward item, which describes the water quality safety situation by calculating the sum of squares of the excessive amount of pollutant concentration at all sensitive monitoring points; The second sub-item is an ecological protection reward item, which describes the ecological impact by constraining the dissolved oxygen level and biological toxicity of the water body; The third sub-item is an operation cost reward item, which describes the control cost by the change amplitude and frequency of the gate opening; Each sub-item is multiplied by the corresponding weight coefficient and then added to obtain the total reward value, and the total reward value accumulated in the long term is maximized through iterative optimization, and finally a set of coordinated gate opening adjustment action sequences of all gates in multiple time steps in the future is output.
9. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 8, characterized in that: In the rolling optimization and instruction execution module, the gate control instruction of the first time step in the coordinated gate opening adjustment action sequence is selected and issued to the corresponding physical gate actuator for execution, which specifically includes the following steps: First, the coordinated gate opening adjustment action sequence generated by the coordinated strategy decision module is parsed, and the instruction set containing the opening adjustment amount of all gates corresponding to the first time step in the current decision period is extracted; Then, each opening adjustment amount in the instruction set is subjected to physical safety and executability verification, which includes amplitude limiting verification processing and action direction verification processing, wherein the amplitude limiting verification processing compares each opening adjustment amount with the maximum and minimum opening limit values allowed by the corresponding gate actuator, and sets the adjustment amount less than the minimum opening limit value to the minimum opening limit value, sets the adjustment amount greater than the maximum opening limit value to the maximum opening limit value, and keeps the adjustment amount within the limit value range unchanged; after verification, the verified instruction is issued to the corresponding physical gate actuator through an industrial real-time communication protocol; at the same time of issuing the instruction, a timeout monitoring timer is started, and the waiting time of the timer is set as the smaller one of the system sampling period and a preset instruction execution timeout threshold; within the waiting time, the state feedback signals from each physical gate actuator are continuously monitored and received to verify whether the instruction is correctly executed.
10. The intelligent control system for multiphase conversion of water quality in a controlled river section according to claim 9, characterized in that: After waiting for a sampling period, the specific operation of triggering the data assimilation and state updating module to start a new round of state updating is as follows: When the sampling period timing ends, the rolling optimization and instruction execution module immediately sends a trigger signal to the data assimilation and state updating module; the trigger signal starts a new round of data assimilation process, and uses the latest collected sensor observation data to correct the state of the digital twin, to generate an updated river network full-element state vector; the updated river network full-element state vector is set as the initial condition of the optimization problem in the next decision period; at the same time, the time window of the prediction time domain is rolled forward by one sampling period; in addition, the rolling optimization and instruction execution module also evaluates the control performance by calculating the weighted integral of the deviation between the actual value and the expected value of the water quality index within a period of time, and dynamically adjusts the weight parameters in the global reward function according to the trend of the performance evaluation result, to realize adaptive optimization of the parameters; The updated river network full-factor state vector will be sent into the agent model calculation module as the starting point of the next decision cycle to start a new round of prediction, decision and rolling optimization cycle, thus forming a complete closed-loop intelligent regulation process based on real-time feedback.
Citation Information
Patent Citations
Flood control dispatching method based on digital twin
CN113222296A
Intelligent flood forecasting method and device based on FQI-DMTSNet adaptive generative adversarial mechanism and readable storage medium thereof
CN120470540A