Sewage treatment control method and system based on multi-dimensional variable data analysis
By constructing a causal directed graph and a control-oriented feature space, rule-based and learning-based candidate strategies are generated, achieving deep coupling between data analysis and control and dynamic coordination between process units in the wastewater treatment system. This solves the problems of data analysis and control separation and conflicting process unit objectives, and improves the system's stability and operating efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIAXING UNIV
- Filing Date
- 2026-04-24
- Publication Date
- 2026-06-19
AI Technical Summary
Existing wastewater treatment control methods are characterized by a disconnect between data analysis and control execution, failure to consider the impact of control performance in feature selection, insufficient coordination of conflicting objectives between process units, and an inability to adapt to changes in temporal and spatial conditions.
Based on multidimensional variable data analysis, a causal directed graph structure is constructed to generate a control-oriented feature space. Causal paths are extracted to generate rule-based and learning-based candidate strategies. Through a two-layer collaborative control module, dynamic coordination of target conflicts between process units and spatiotemporal adaptive control within the unit are realized.
It achieves deep coupling of data analysis and control, improves the stability, response speed and robustness of the wastewater treatment system, optimizes the target coordination between process units and the spatiotemporal control within the units, and improves operating efficiency and anti-disturbance capability.
Smart Images

Figure CN122233587A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment technology, specifically to a wastewater treatment control method and system based on multidimensional variable data analysis. Background Technology
[0002] Wastewater treatment processes are characterized by nonlinearity, large time delays, and multivariate coupling, and their control effectiveness directly impacts effluent quality compliance rates and operational energy consumption. Current wastewater treatment control technologies are mainly divided into two categories: one is classical control methods based on mechanistic models, such as PID control and fuzzy control, which establish setpoints for key parameters like dissolved oxygen and nitrate nitrogen for feedback tracking; the other is data-driven optimization control methods, which collect data from multiple sensor sources, utilize algorithms such as neural networks and support vector regression to establish water quality prediction models or energy consumption models, then solve for the optimal control setpoints using optimization algorithms such as particle swarm optimization and genetic algorithms, and finally execute the control through a lower-level controller. However, existing technologies still have the following shortcomings: First, existing wastewater treatment control methods suffer from a disconnect between data analysis and control execution, and feature selection fails to consider the impact on control performance. In current wastewater treatment control methods, multidimensional variable data analysis is typically used as a preliminary step in generating optimal setpoints. The analysis results are only used to determine the optimal target values for key parameters such as dissolved oxygen and nitrate nitrogen, without directly participating in the generation of the control law, resulting in low coupling between the analysis model and the controller. Furthermore, existing methods prioritize prediction accuracy during feature extraction or dimensionality reduction. While the selected features can improve the accuracy of the water quality prediction model, they do not consider the impact of these features on the stability, response speed, and robustness of the control system. Some features with high predictive contribution may cause control oscillations due to their high noise and strong time delay. Therefore, how to achieve deep coupling between data analysis and control execution, and construct a feature selection mechanism oriented towards control performance, are urgent technical problems to be solved. Second, existing wastewater treatment control methods suffer from insufficient coordination of inter-process unit conflicting objectives and difficulty in adapting to spatiotemporal changes in homogenized control within units. In existing wastewater treatment control methods, each process unit is typically optimized and controlled as a whole. When conflicting objective functions exist between units, there is a lack of effective coordination mechanisms. Traditional multi-objective optimization methods can only convert multiple objectives into a single objective with weights, making it difficult to explicitly model the game-theoretic relationships between units. Furthermore, existing methods apply uniform control parameters to the entire unit for homogenization within units, failing to adapt to spatiotemporal changes caused by uneven influent distribution and spatial differences in microbial activity within the unit. Therefore, how to achieve dynamic coordination of inter-process unit conflicting objectives and realize spatiotemporally adaptive fine control within units is a pressing technical problem that needs to be solved. Summary of the Invention
[0003] To achieve the above objectives, the present invention provides the following technical solution: a wastewater treatment control method based on multidimensional variable data analysis, the method comprising: Collect and preprocess raw multidimensional variable data from the entire wastewater treatment process to generate standardized multidimensional variable data vectors; Based on standardized multidimensional variable data vectors, the direction of causal relationships between variables is calculated and determined, a causal relationship directed graph structure is constructed, the control performance impact metric score of each variable in the causal relationship directed graph structure is calculated, and a control-oriented feature space is generated. The causal paths in the directed graph structure of causal relationships are extracted to generate regular candidate policies. Reinforcement learning is performed on real-time variable values based on the control-oriented feature space to generate learned candidate policies. The regular candidate policies and learned candidate policies are merged to obtain a candidate policy set. Calculate the overall credibility score for each strategy in the candidate strategy set, and select the strategy with the highest overall credibility score from the strategies with credibility scores higher than the credibility threshold as the global optimal strategy. Each fixed process unit in the entire wastewater treatment process is defined as an agent. The allowable range of the operation variables of each agent is calculated. Dynamic control micro-domains are generated within each agent. With the global optimal strategy and the allowable range of the operation variables as constraints, matching strategies are searched in the candidate strategy set to generate micro-domain control instructions. Collect actual execution data after micro-domain control commands are executed, calculate the execution deviation between the actual execution data and the expected data of the global optimal strategy, and make corrections and updates based on the execution deviation.
[0004] Furthermore, this application proposes a wastewater treatment control system based on multidimensional variable data analysis to implement the aforementioned wastewater treatment control method based on multidimensional variable data analysis. The system includes: The data acquisition and processing module is used to collect and preprocess the raw multi-dimensional variable data of the entire wastewater treatment process, and generate a standardized multi-dimensional variable data vector. The causal identification and feature measurement module constructs a causal relationship directed graph structure based on standardized multidimensional variable data vectors, calculates the control performance impact measurement score of each variable in the causal relationship directed graph structure, and generates a control-oriented feature space. The candidate control strategy generation module is used to generate rule-based candidate strategies based on the causal directed graph structure, generate learning-based candidate strategies based on the control-oriented feature space, and merge the rule-based candidate strategies and learning-based candidate strategies to obtain a candidate strategy set. The strategy credibility evaluation module is used to calculate the comprehensive credibility score of each strategy in the candidate strategy set, and select the globally optimal strategy based on the comprehensive credibility score; The dual-layer collaborative control module is used to define each fixed process unit in the entire wastewater treatment process as an intelligent agent, calculate the allowable range of operation variables for each intelligent agent, divide the dynamic control micro-domain, and generate micro-domain control instructions. The feedback and evolution module is used to collect the actual execution data after the micro-domain control command is executed and calculate the execution deviation, and make corrections and updates based on the execution deviation.
[0005] This invention provides a wastewater treatment control method and system based on multidimensional variable data analysis. It has the following beneficial effects: 1. This invention achieves deep coupling between data analysis and control through causal inference and control performance impact measurement. Utilizing a constructed directed graph structure of causal relationships, the causal relationships between variables in the wastewater treatment process are quantitatively characterized. Based on this, a control performance impact measurement for each variable is calculated. This measurement comprehensively evaluates the weight of each variable's influence on system stability, response speed, and robustness. Unlike traditional methods that use analysis results only to generate optimal setpoints, this invention directly maps the causal structure to the dynamic topology of the controller, enabling the control law to adaptively adjust with changes in causal relationships. Simultaneously, the control performance impact measurement filters out beneficial feature variables for the control process, constructing a control-oriented feature space. This avoids control oscillations caused by selecting features solely for prediction accuracy. The causal inference results are not only used for strategy generation but also directly participate in strategy credibility assessment and game coordination processes, forming a fully closed-loop deep coupling mechanism from data analysis to control execution and effect feedback. This significantly improves the stability, response speed, and robustness of the control system.
[0006] 2. This invention constructs a two-layer collaborative control architecture consisting of an outer-layer multi-agent game-theoretic collaboration and an inner-layer spatiotemporal dynamic control domain partitioning. In the outer layer, each fixed process unit is defined as an agent with an independent objective function. A distributed game is used to solve for the Nash equilibrium points among the agents, explicitly modeling the competitive and cooperative relationships between units and outputting the allowable range of operational variables for each unit, thus achieving dynamic equilibrium under objective conflict. In the inner layer, each unit is partitioned in terms of time dimension and clustered in terms of space dimension to generate dynamic control micro-domains. Each micro-domain is matched with an independent control strategy based on its current state and historical response characteristics, and a coordination mechanism between micro-domains ensures that the total amount of operations does not exceed the allowable range given by the outer-layer game. This two-layer architecture achieves explicit coordination of objective conflicts between units and adaptive control of spatiotemporal changes within units, ensuring both global collaborative optimization and local fine-grained regulation, thereby improving the overall operating efficiency and anti-disturbance capability of the wastewater treatment system. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating the steps of the wastewater treatment control method based on multidimensional variable data analysis of the present invention. Figure 2 This is a data transmission flowchart of the wastewater treatment control method based on multidimensional variable data analysis according to the present invention; Figure 3 This is an architecture diagram of the wastewater treatment control system based on multidimensional variable data analysis according to the present invention. Detailed Implementation
[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0009] like Figures 1 to 2 As shown, a wastewater treatment control method based on multidimensional variable data analysis is proposed, which includes: Step S100 involves collecting and preprocessing the raw multidimensional variable data of the entire wastewater treatment process to generate a standardized multidimensional variable data vector. The raw multidimensional variable data includes four categories: influent parameters, process parameters, effluent parameters, and equipment status parameters. The standardized multidimensional variable data vector is used to eliminate numerical differences between variables of different dimensions, ensuring that subsequent causal relationship calculations and characteristic measurements are performed on a unified scale, thus avoiding interference from dimensional differences in the analysis results.
[0010] Step S200: Based on standardized multidimensional variable data vectors, calculate and determine the direction of causal relationships between variables, construct a directed causal relationship graph structure, calculate the control performance impact metric score for each variable in the directed causal relationship graph structure, and generate a control-oriented feature space. The directed causal relationship graph structure is used to quantitatively characterize the causal dependencies between variables in the wastewater treatment process, ensuring that candidate strategies are generated based on the true impact relationships between variables, reflecting the causal chain between variables in the wastewater treatment system. The control performance impact metric score is used to quantify the comprehensive impact on the system; variables with higher scores indicate a more positive effect on the stable operation, rapid response, and anti-interference capability of the control system, ensuring that the variables upon which subsequent control strategies rely are beneficial to the control process, rather than just to prediction accuracy. The control-oriented feature space is used to compress the input dimensions of subsequent strategy generation and credibility assessment, reducing computational complexity and ensuring that all subsequent processing revolves around variables that are truly important to the control process, avoiding interference from redundant and noisy variables.
[0011] Step S300: Extract causal paths from the directed graph structure of causal relationships to generate rule-based candidate strategies. Perform reinforcement learning on real-time variable values based on the control-oriented feature space to generate learned candidate strategies. Merge the rule-based and learned candidate strategies to obtain a candidate strategy set. Rule-based candidate strategies reflect stable causal relationships between variables in the wastewater treatment process. Learned candidate strategies cover complex scenarios that rule-based strategies cannot handle, reflecting implicit patterns learned by the system from historical operating data. They can adapt to changes in operating conditions, but their interpretability is relatively lower than that of rule-based strategies. The candidate strategy set reflects all available control methods the system possesses when dealing with different process states, demonstrating the comprehensiveness of strategy generation.
[0012] Step S400: Calculate the comprehensive credibility score for each strategy in the candidate strategy set. From the strategies with comprehensive credibility scores higher than the credibility threshold, select the strategy with the highest comprehensive credibility score as the globally optimal strategy. The comprehensive credibility score threshold is dynamically adjusted based on the current system stability: when the system is running smoothly, the threshold increases, prioritizing high-credibility strategies; when the system is subjected to disturbances (such as water inrush), the threshold decreases, allowing the attempt of new strategies to cope with emergencies. This reflects the system's requirement for strategy reliability under current operating conditions and demonstrates the system's adaptive decision-making capability under different risk environments. The globally optimal strategy provides top-level guidance for subsequent two-layer collaborative control. The outer-layer game uses this strategy as the initial proposal, and the inner-layer micro-domain uses this strategy as the screening criterion, ensuring consistency between local control and global objectives.
[0013] Step S500: Each fixed process unit in the entire wastewater treatment process is defined as an agent. The allowable range of the operational variables for each agent is calculated. Dynamic control micro-domains are generated within each agent. Using the global optimal strategy and the allowable range of operational variables as constraints, a matching strategy is searched in the candidate strategy set to generate micro-domain control instructions. The allowable range of operational variables provides a global constraint boundary for the inner micro-domain control, achieving a unity of global coordination and local autonomy. Through dynamic control micro-domains, refined control within the unit is achieved, enabling the control strategy to adapt to dynamic changes in space and time, avoiding the local overexposure or underexposure problems caused by traditional homogenized control. Micro-domain control instructions are specific execution instructions generated for each dynamic control micro-domain. The global optimal strategy and the allowable range of operational variables are decomposed into executable local instructions and issued to the corresponding actuators, realizing a complete link from global decision-making to local execution, ensuring that each micro-domain operates according to a control strategy adapted to its state.
[0014] Step S600: Collect actual execution data after the micro-domain control command is executed, calculate the execution deviation between the actual execution data and the expected data of the global optimal strategy, and make corrections and updates based on the execution deviation. The actual execution data is used to evaluate the difference between the actual execution effect and the expected effect of the control strategy, reflecting the execution effect of the control strategy under real operating conditions, and revealing possible prediction deviations or unforeseen disturbance effects in the strategy generation stage.
[0015] In this embodiment, the detailed implementation steps for collecting and preprocessing the original multidimensional variable data of the entire wastewater treatment process to generate a standardized multidimensional variable data vector include: Step S101: Acquire raw data of multidimensional variables for the entire wastewater treatment process. Influent parameters include COD sensor parameters, ammonia nitrogen parameters, total nitrogen parameters, total phosphorus parameters, pH parameters, flow rate parameters, and temperature parameters, collected by online monitoring instruments and flow meters installed at the influent end, used to characterize the influent load characteristics. Process parameters include multi-point dissolved oxygen, oxidation-reduction potential, mixed liquor suspended solids concentration, sludge age, return ratio, aeration rate, and chemical dosage, collected by sensors and flow meters deployed within each process unit, reflecting the current biochemical reaction status and equipment operating status. Effluent parameters include effluent COD, ammonia nitrogen, total nitrogen, and total phosphorus, collected by online monitoring instruments at the effluent end, used to assess whether the control effect meets emission standards. Equipment status parameters include blower frequency, pump frequency, valve opening degree, and meter readings, collected through feedback signals from the equipment and meter readings, used to calculate energy consumption and monitor equipment operating status. Each parameter of the above multidimensional variable data is acquired based on existing instruments and sensors used in wastewater treatment.
[0016] Step S102: Standardize the original multidimensional variable data to obtain a standardized multidimensional variable data vector. First, calculate the mean of each variable x over a historical period (e.g., the past 7 days). and standard deviation Then, the raw values collected at each time t are... The standardized values are obtained by dividing the difference between the original values and the mean by the standard deviation. Finally, the standardized values of all variables are aligned in chronological order to form a standardized multidimensional variable data vector matrix. The rows of the matrix correspond to the time step, and the columns correspond to the variable categories. This matrix serves as the basic input data for subsequent steps.
[0017] In this embodiment, based on standardized multidimensional variable data vectors, the implementation steps of calculating and determining the direction of causal relationships between variables, constructing a directed graph structure of causal relationships, calculating the control performance impact metric score for each variable in the directed graph structure of causal relationships, and generating a control-oriented feature space include: Step S201: Obtain standardized multidimensional variable data vectors, calculate the time-series correlation coefficient and partial correlation coefficient between each variable, determine the direction of the causal relationship between the variables based on the time-series correlation coefficient and partial correlation coefficient, and construct a directed graph structure of the causal relationship based on the direction of the causal relationship and the variables. The time-series correlation coefficient is used to measure the degree of linear correlation between two variables at different time lags; for variables A and B, calculate... and Pearson correlation coefficient , Where t is time; k is the lag step size, ranging from 1 to L, with L typically being 5 to 10; and These represent the means of the variables. By calculating the time-series correlation coefficient, we can initially identify potential time-series influence relationships between the variables. The partial correlation coefficient is used to measure the net correlation between two variables, after controlling for the influence of other variables. For variables A and B, and the set of variables C to be controlled, we calculate the correlation coefficient between A and B after eliminating the influence of C. , ;in , , The corresponding variable pairs are listed in order. , , The time-series correlation coefficient between the two variables is used to determine whether the correlation between the two variables is mediated by a third variable by calculating the partial correlation coefficient, thereby eliminating spurious correlations. When determining the direction of causal relationships between variables based on time-series correlation coefficients and partial correlation coefficients, firstly, for each pair of variables... Calculate the time series correlation coefficients under different lags. and Find the lag direction corresponding to the largest time-series correlation coefficient; then find the set C of all variables that are correlated with both A and B, and calculate the partial correlation coefficient under the condition of controlling for the set C of variables. If the partial correlation coefficient decreases significantly (below 50% of the original correlation coefficient), then the correlation between A and B is determined to be mediated by C, with no direct causal relationship. If the partial correlation coefficient does not decrease significantly, then a direct causal relationship exists. Finally, for variable pairs with a direct causal relationship, a comparison is made. and :like ,and Then determine There is a causal relationship in the direction; if ,and Then determine There is a causal relationship; where 0.2 is a preset difference threshold. If the difference between the two is within the preset threshold, that is... , If the direction is not marked, then the direction will not be marked for the time being.
[0018] The directed graph structure for causal relationships is constructed using variables as nodes and the determined causal relationship directions as directed edges. First, a node set is created, with each variable from the standardized multidimensional variable data vectors involved in the analysis treated as a node, and the edge set is initialized to empty. Then, all variable pairs determined to have a direct causal relationship are traversed, and directed edges from the cause variable to the effect variable are added according to the determined causal relationship direction. For each directed edge, its initial confidence value is recorded as 1. Finally, the complete directed graph structure for causal relationships is output. The causal directed graph structure contains all nodes, directed edges, and the confidence of directed edges, which are used for subsequent causal path extraction and variable measurement calculation. The causal directed graph structure contains all nodes divided into causal variable nodes and result variable nodes. The causal variable nodes contain operation variable nodes, and the result variable nodes contain control target variable nodes.
[0019] Step S202: For each variable in the directed graph structure of causal relationships, calculate the stability impact score, response speed impact score, and robustness impact score, and then sum them by weight to obtain the control performance impact measure score for each variable. The stability impact score is used to measure the degree of influence of a variable on the stability of the control system, reflecting the breadth and depth of the propagation of fluctuations of that variable through the causal path; it is obtained by statistically analyzing the current variable... Length of all causal paths starting from Calculate the average path length , , This represents the length of the j-th path in the total number of causal paths N; statistics are based on... The number N of causal paths originating from the first path; calculate the stability impact score. , The response speed impact score measures the degree of influence of a variable on the control response speed, reflecting the shortest causal distance from the variable to the control target. It extracts all paths originating from the manipulated variable node and ending at the control target variable node from the causal directed graph structure; and statistically analyzes the current variable... Shortest causal distance to any node controlling the target variable Calculate the impact of response speed on the score , Robustness impact score measures the degree of influence of a variable on the robustness of a control system, reflecting the variable's fluctuation characteristics and its transmission attenuation characteristics along the causal path; it is calculated by... Variance in historical data , Where T is the number of historical time steps; For variables The standardized value at time t; The mean is calculated from... Average fluctuation propagation coefficient of all nodes on the causal path from which it originates That is, the average attenuation rate of the fluctuation amplitude after each causal edge, used to calculate the robustness impact score. , This formula represents variables volatility variance The larger the value, the weaker the fluctuation transmission along the causal path. The smaller the value, the greater the positive impact on the system's robustness.
[0020] Among them, the average fluctuation transmission coefficient In calculations, first start with the variables Starting from the causal directed graph structure, traverse all nodes from... A directed path from which the target variable is reached or the path ends; each path is denoted as . , containing node sequences For each path Each directed edge on Calculate the wave propagation coefficient : ;in For covariance, For the causal lag step size, For variables The variance of the fluctuation, and These are all nodes corresponding to variables. This coefficient reflects the strength of the influence of the u fluctuation on w, with a value between [0,1]. The closer to 1, the stronger the fluctuation transmission. Then, the path is calculated. Comprehensive transmission coefficient Multiply the wave propagation coefficients of all edges along the path to obtain the overall propagation coefficient of the path. : Finally, the average fluctuation propagation coefficient is calculated. For all from The average of the starting paths: ;in From Total number of causal paths originating from. Average fluctuation propagation coefficient. Reflects variables The smaller the value, the faster the fluctuation attenuates during the transmission process, which is more beneficial to the system's robustness.
[0021] Control performance impact metric score It is a comprehensive evaluation variable Quantitative indicators of the impact on the stability, response speed and robustness of the control system are used to screen variables that are beneficial to the control process and construct a control-oriented feature space. The weighted fusion calculation formula is as follows: ; in, For variables The control performance affects the measurement score; For variables Stability affects the score The weighting coefficient is initially set to 0.4, representing the system's emphasis on stability; For variables Response speed affects the score The weighting coefficient is initially set to 0.3, representing the system's emphasis on response speed; For variables Robustness impact score The weighting coefficient, initially set to 0.3, represents the system's emphasis on robustness. , , It can be dynamically adjusted according to the actual operation. When the system reveals a certain performance deficiency, the weight of that aspect will be increased accordingly.
[0022] Step S203: Sort the control performance impact metric scores from high to low, and select a preset number of variables corresponding to the control performance impact metric scores to form a control-oriented feature space. The preset number is generally set to 30% to 50% of the total number of variables in the causal directed graph structure, with the specific value set based on the total number of variables and the system's computing power. When the total number of variables is large (e.g., more than 50), the preset ratio is 30% to reduce computational complexity; when the total number of variables is small (e.g., less than 20), the preset ratio is 50% to ensure the representativeness of the feature space. By setting a preset number for variable selection, key control variables are retained while the input dimensions of subsequent steps are compressed, computational efficiency is improved, and redundant and noisy variables are avoided from interfering with policy generation and credibility assessment.
[0023] When constructing the control-oriented feature space, the control performance impact metric scores of all variables in the causal directed graph structure calculated in step S201 are first obtained. Sort all variables from highest to lowest according to their control performance impact metric scores to obtain the sorted sequence. Then, based on a preset quantity K, where K is 30% to 50% of the total number of variables m, select the top K variables and record them as the variable list. ,but Finally, the causal relationships between variables in the variable list are extracted from the directed graph structure of causal relationships to form a feature causal subgraph. , variable list and feature causal subgraph Together, they form the control-oriented feature space, the output of which is used in subsequent steps. The control-oriented feature space serves as the core input for subsequent candidate policy generation, policy credibility assessment, and dynamic control micro-domain partitioning, ensuring that all subsequent processing revolves around variables that are truly important to the control process. The control-oriented feature space includes the variable names of the selected variables, standardized variable data vectors, control performance impact metric scores, and upstream and downstream relationships in a causal directed graph structure.
[0024] In this embodiment, the steps of extracting causal paths from the directed graph structure of causal relationships to generate rule-based candidate policies, performing reinforcement learning on real-time variable values based on the control-oriented feature space to generate learned candidate policies, and merging the rule-based candidate policies and learned candidate policies to obtain a candidate policy set include: Step S301, the process of generating rule-based candidate strategies includes: Traverse all nodes in the causal directed graph structure, marking the manipulated variable nodes and control target variable nodes. Manipulated variable nodes refer to the graph nodes corresponding to variables that can be directly adjusted in the wastewater treatment process, including aeration rate, return ratio, chemical dosage, sludge discharge rate, blower frequency, and pump frequency. When setting manipulated variable nodes, in the causal directed graph structure, nodes that can be directly controlled by the actuator are marked as manipulated variable nodes, serving as adjustable objects in strategy generation and the direct points of application for control commands. Control target variable nodes refer to the graph nodes corresponding to variables that need to be maintained within a certain range in the wastewater treatment process, including dissolved oxygen concentration, effluent ammonia nitrogen, effluent total phosphorus, effluent COD, and energy consumption. Nodes reflecting the system's operating effect and directly related to emission standards or operating costs are marked as control target variable nodes, serving as the optimization direction for strategy generation. The expected effect of the strategy must revolve around these variables.
[0025] For each pair of operand and target variable nodes, find all directed paths from the operand node to the target variable node. When generating directed paths, first select a pair of operand and target variable nodes, with the operand node as the starting node and the target variable node as the ending node. Then, use a depth-first traversal method, starting from the starting node and recursively visiting adjacent nodes along the directed edges, recording the sequence of nodes visited each time. When the ending node is reached, the sequence is saved as a directed path. Finally, continue traversing until all possible paths have been explored, and output all directed paths from the operand node to the target variable node. Directed paths are used to reveal the causal chain between variables, providing a basis for policy generation.
[0026] Intermediate variables are extracted from each directed path to form a causal path. Policy rules are generated based on these causal paths, and the set of policy rules for all directed paths forms a rule-based candidate policy. Intermediate variables refer to the variable nodes located between the operand variable node and the control target variable node on a directed path. ,in All variables are intermediate variables, with O representing the manipulated variable node and T representing the control target variable node. Intermediate variables reflect the specific transmission mechanism by which the manipulated variable node influences the control target variable node. By monitoring the state of intermediate variables, the control effect can be predicted in advance, thus achieving feedforward control. A causal path refers to a directed path from the manipulated variable node to the control target variable node, emphasizing the causal transmission relationship between variables. This provides a direct basis for rule-based candidate strategies, ensuring the strategy is built on a clear causal chain. When generating causal paths, firstly, all directed paths originating from the manipulated variable node and ending at the control target variable node are extracted from the causal directed graph structure. Then, for each directed path, the sequence of intermediate variables is extracted. Finally, each directed path and its sequence of intermediate variables are saved as a causal path. The set of strategy rules generated from all causal paths constitutes a rule-based candidate strategy, with each strategy rule serving as an element in the rule-based candidate strategy. Rule-based candidate strategies cover scenarios with clear causal paths and explicit causal relationships, providing highly interpretable strategy options for subsequent credibility assessment. The strategy rules are conditional action-based decision-making logic generated based on causal paths, set according to the correlation between the state of intermediate variables and the adjustment of operational variables along the causal path. Each strategy rule consists of four parts: intermediate variable threshold condition, operational variable adjustment range, expected effect, and dependent variable set. The intermediate variable threshold condition refers to the rule triggered when the real-time value of an intermediate variable exceeds or falls below a preset threshold. Intermediate variables are extracted from the causal path and determined by combining historical data distribution. Based on historical operating data of intermediate variables, the threshold is set using the percentile method (e.g., the 10th and 90th percentiles) or according to process experience (e.g., triggering increased aeration when DO is below 1.5 mg / L). The operational variable adjustment range refers to the specific amount of adjustment made to the operational variable. Typical adjustment amounts or preset discrete levels are extracted from historical operating data. Based on the process operation manual, the minimum adjustment accuracy of the actuator, and common adjustment amounts in historical operation, preset discrete levels such as ±5% and ±10% are used. The expected effect refers to the anticipated change in the target variable after implementing the strategy (e.g., an increase in DO of 0.3 mg / L). This is estimated using causal effects from historical data, specifically the average change in intermediate and target variables after adjusting the operational variables, or estimated using the transmission coefficients of the causal path. The dependency variable set refers to the list of intermediate and operational variables upon which the strategy depends, directly taken from all variable nodes along the causal path. Policy rules transform causal relationships into directly executable control logic, providing the system with highly interpretable policy options and clearly defined execution paths.
[0027] Step S302, the process of generating learning-based candidate strategies includes: The state vector is constructed by taking the current values of all variables in the control guidance feature space and the time values of the previous preset number of times. The adjustment range of the operational variables constitutes a discrete action set. The current value of all variables refers to the standardized value of each variable in the control guidance feature space at the current time t, derived from the standardized multidimensional variable data vector generated in step S100. In the process, the values of all variables in the control-guided feature space corresponding to the current time t are extracted to form the main part of the state vector, reflecting the current operating state of the system and providing input for reinforcement learning decision-making. The preset number of time values refers to the variable values at several times (such as t-1, t-2, ..., tp) before the current time t, where p is the preset number, used to introduce time series information so that the state vector can reflect the changing trend of variables and enhance the model's ability to perceive dynamic processes. The preset number p is generally taken as 3~5, based on the time constant and sampling frequency of the wastewater treatment process, usually ensuring coverage of historical information within a basic response cycle (such as 10~20 minutes). The manipulated variable is a process parameter that can be directly controlled, consistent with the manipulated variable node in step S301. The adjustment range of the manipulated variable refers to the specific value or proportion of the manipulated variable to be adjusted, such as "increase aeration by 5%" or "decrease reflux ratio by 3%"; the adjustment range is obtained by preset several discrete adjustment levels (such as -10%, -5%, 0, +5%, +10%) according to the type of manipulated variable and the current operating state. The discrete action set is the set of all possible combinations of adjustment magnitudes of the operation variables, denoted as . Each action This represents a vector containing the adjustment range of multiple operational variables, such as action a = (aeration rate + 5%, reflux ratio - 2%). The content of the discrete action set is preset according to the actual process requirements, and usually includes common adjustment combinations. It is used to discretize the continuous control space and simplify the search space of the reinforcement learning strategy.
[0028] Using a discrete set of actions as the dynamic space and in conjunction with state vectors, a reward function is constructed that measures the impact of control performance. This function is then used to train and generate learnable candidate policies. Next execution action vector The immediate reward obtained afterward, its reward function The formula is defined as follows: ; in, To perform the action The projected change in energy consumption (standardized value). To perform the action The expected degree of exceedance of the standard for the effluent quality (standardized value). For action The average CPIM score of the operational variables involved in the control-oriented feature space; This is the penalty coefficient for exceeding water quality standards, and a positive value is used. This is the CPIM incentive factor, and a positive value is taken.
[0029] The reward function guides the selection of actions in the deep Q-network model that reduce energy consumption, ensure water quality, and favor controllable variables. The training process for generating learned candidate policies begins by constructing a deep Q-network, with the state vector as input. The output is the Q-value of each action in the discrete action set. It is trained offline using historical running data, and the network parameters are updated according to the Bellman equation to make the Q-value approximate the optimal action value. After training, the deep Q-network is deployed online as a policy network. In each control cycle, the current state vector is input, and the action with the highest Q-value is output as the learning candidate policy. The deep Q-network used for training the reward function is a well-established and mature technology. Its internal architecture and layers can be directly implemented based on existing deep Q-network architectures in this embodiment. Specifically, computer software such as PyTorch can be used to set up the environment and import the deep Q-network model.
[0030] Learning-based candidate policies are policy networks trained through reinforcement learning using deep Q-networks. They can output the optimal action based on the current state vector. Learning-based candidate policies include action selection logic, key state variables (obtained through feature importance analysis), and the expected effect after executing the action (estimated through the reward function). They are used to cover complex scenarios that rule-based policies cannot handle, such as multivariate coupling and nonlinear responses, and provide flexible alternatives.
[0031] Step S303: Merge the rule-based candidate policies and the learning-based candidate policies to obtain a candidate policy set. The candidate policy set is the complete set of policies formed by merging the rule-based and learning-based candidate policies, denoted as... , Each strategy The candidate strategy set includes a strategy ID, operation instructions (specific adjustment values), expected effect, a set of dependent variables, and a strategy type tag. The strategy type tag is divided into rule-based candidate strategies and learning-based candidate strategies. This set of candidate strategies provides diverse alternatives for subsequent credibility assessment, enabling the system to select the optimal strategy based on credibility scores in different scenarios, balancing interpretability and flexibility. When generating the candidate strategy set, the rule-based candidate strategy set is first obtained from step S301. Obtain the set of learning candidate strategies from step S302. Typically, the top K actions with the highest Q-values are selected as candidates; then the two sets are merged to obtain the candidate strategy set. , Finally, a unique policy ID is generated for each policy, and the candidate policy set is stored for use in step S400.
[0032] In this embodiment, the steps of calculating the comprehensive credibility score of each strategy in the candidate strategy set and selecting the strategy with the highest comprehensive credibility score as the globally optimal strategy from the strategies with a comprehensive credibility score higher than the credibility threshold include: Step S401: Obtain the candidate strategy set, and calculate the path clarity score, variable reliability score, and uncertainty score for each strategy in the candidate strategy set. Path Clarity Score Used to measure the simplicity and interpretability of the causal path upon which the strategy depends; for rule-based candidate strategies, the length of the causal path upon which it depends is taken. Calculate path clarity score For learning-based candidate policies, trace the state vectors that the deep Q-network depends on during decision-making, and statistically determine the shortest distance from these state vectors to the control target variable within the causal directed graph structure. ,calculate This reflects the intuitiveness of the strategic decision-making logic; the shorter the path and the closer the distance, the higher the interpretability. Variable reliability score. This is used to measure the reliability of policy dependency variables. The calculation process involves selecting the variables that the candidate policy set depends on to form a dependency variable set. Calculate the set of dependent variables List of variables belonging to the control-oriented feature space proportion , Calculate the average control performance impact metric score of dependent variables. , ; Calculate the ratio to the maximum CPIM , Finally, the reliability score of the variables is calculated. The variable reliability score reflects the reliability of the variables on which the strategy depends; the more variables with high CPIM scores the strategy relies on, the higher the reliability. Uncertainty score. Used to measure the current prediction error and volatility of the variables on which the strategy depends; the calculation process is as follows: for the set of dependent variables... For each variable v in the equation, calculate its prediction error at the current time. and volatility variance ; Calculate the uncertainty contribution of this variable. , The uncertainty score is obtained by averaging all dependent variables. , Uncertainty scoring is used to quantify the uncertainty during strategy execution; the higher the uncertainty, the lower the reliability of the strategy. Specifically, it involves calculating the prediction error at the current moment. For variable v, an autoregressive model is built based on historical data to calculate the predicted value at the current time. , Where c is a constant term, It is assumed that the mean is equal to 0; , , All are autocorrelation coefficients; p is the order, usually taken as 3~5; using the actual values of the first p time steps. Predict the current value Calculate the prediction error , ;in Let v be the historical standard deviation. This represents the actual measured value of variable v at the current moment, used for normalization to make errors comparable across different variables. It is used in calculating variance. At this time, set the time window length W, such as the number of data points in the most recent hour, and obtain the value of variable v at W time points before the current time. Calculate the mean within the window , Calculate the variance of the fluctuation. ;in, This represents the actual value of variable v at time tk, where t is the current time and k is the number of steps shifted to the past.
[0033] The interpretability score for each strategy is obtained by weighting and summing the path clarity score and variable reliability score, then multiplying the sum by an uncertainty penalty coefficient. The uncertainty penalty coefficient η is a coefficient between 0 and 1 used to introduce uncertainty penalty into the interpretability score; when the strategy's uncertainty score... If the score is high, the interpretability score of the strategy is reduced. The specific calculation formula is as follows: The uncertainty penalty coefficient is used to penalize strategies with large prediction errors or drastic fluctuations in dependent variables, thus avoiding the selection of unreliable strategies. Interpretability score. It is a comprehensive quantification of the comprehensibility and reliability of a strategy, and the calculation formula is: ;in, and These are weighting coefficients, with initial values of 0.6 and 0.4, respectively. The interpretability score comprehensively assesses the causal path clarity, variable reliability, and uncertainty of the strategy, providing a foundation for subsequent credibility scoring.
[0034] The average control performance impact metric, historical success rate, and interpretability score for each strategy's dependent variables are weighted and summed to obtain a comprehensive credibility score for each strategy. Historical success rate This refers to the historical success rate when executing strategies similar to the current strategy. A similar strategy is defined as one with the same operational variables and similar adjustment ranges. The historical success rate includes the number of executions, the number of successes, and the success rate (number of successes / number of executions). The historical success rate is calculated by retrieving historical strategy records similar to the current strategy from the historical database and calculating their success rates. If a similar strategy is found, the success rate is used; otherwise, the system default value of 0.5 is used to assist in evaluating the reliability of the strategy. Overall Reliability Score It quantifies the overall credibility of the strategy, taking into account interpretability, friendliness of variable control, and historical experience. The calculation formula is: ;in, , , These are the first, second, and third weighting coefficients, with initial values of 0.4, 0.3, and 0.3 respectively, which can be dynamically adjusted based on actual performance. The overall credibility score provides a unified quantitative standard for strategy selection, ensuring the selection of the strategy with the best overall performance.
[0035] Step S402: Dynamically adjust the credibility threshold based on the current system stability, filter out strategies with a comprehensive credibility score higher than the current credibility threshold, and select the strategy with the highest comprehensive credibility score as the globally optimal strategy. Credibility threshold It is a dynamically adjusted screening threshold. Only strategies with a comprehensive credibility score higher than the credibility threshold are included in the final selection. The credibility threshold adaptively adjusts the stringency of strategy selection based on the current stability of the system, increasing the requirements during stable periods and relaxing them during disturbances. The credibility threshold is calculated by first calculating the variables in the control-oriented feature space. The average variance over the most recent preset time period (e.g., 1 hour). , Where R is the number of variables in the control-oriented feature space; This represents the variance of variable v over the most recent preset time period, calculated in the same way as the variance of variance in step S401. The same applies; after subtracting the variance-mean from 1 and normalizing the variance-mean, the current system stability is obtained. , The larger the average variance, the lower the stability of the current system. Then, obtain the preset base threshold. and minimum threshold , set the preset base threshold Multiply by one and subtract the current system stability plus the preset minimum threshold Multiply by the current system stability to obtain the confidence threshold. ,Right now Current system stability The closer the value is to 1, the higher the stability and the closer it is to the confidence threshold. ; The closer the value is to 0, the lower the stability and the closer the reliability threshold is to 0. .
[0036] The globally optimal policy is selected from candidate policies whose overall credibility score is higher than a dynamic threshold, choosing the policy with the highest credibility score as the globally optimal policy. This is denoted as... The globally optimal strategy includes operational instructions, expected effects, a set of dependent variables, and a strategy type label, providing top-level guidance for subsequent two-layer collaborative control. The outer-layer game uses the globally optimal strategy as the initial proposal, while the inner-layer micro-domain uses the globally optimal strategy as the selection criterion, ensuring consistency between local control and the global objective. When generating the globally optimal strategy, the candidate strategy set is first obtained from step S401. Overall credibility score for each strategy Obtain the current credibility threshold from step S402. Then filter out all The strategy is to form a subset of highly reliable strategies; finally, the strategy with the highest reliability score is selected from the subset of highly reliable strategies. Output the globally optimal strategy .
[0037] In this embodiment, each fixed process unit in the entire wastewater treatment process is defined as an intelligent agent, and the implementation steps for calculating the allowable range of the operational variables for each intelligent agent include: Step S501: The fixed process unit includes an aeration unit, a secondary sedimentation unit, and a deep treatment unit, which are defined as the first intelligent agent, the second intelligent agent, and the third intelligent agent, respectively. During the overall design of wastewater treatment, the wastewater treatment system is divided into several physically independent and functionally defined units according to the process flow. Each unit contains a set of operational variables (such as aeration rate, return ratio, and chemical dosage) and a set of state variables (such as DO, MLSS, and effluent quality). Each independent fixed process unit, as the basic subject of the outer multi-agent game, is defined as an intelligent agent with an independent objective function and operational variables. Cooperative optimization between units is achieved through game theory, avoiding the problem of difficult coordination of conflicting objectives in traditional overall optimization. A fixed process unit refers to a process area in the wastewater treatment system that has independent functions and is relatively fixed in space, including the aeration unit, the secondary sedimentation unit, and the deep treatment unit. The aeration unit mainly includes an aeration tank, blower, aeration pipeline and its control valves. Its main function is to provide dissolved oxygen through aeration to support the degradation of organic matter by microorganisms. The secondary sedimentation unit mainly includes a secondary sedimentation tank, sludge return pump and excess sludge discharge valve. Its main function is to achieve sludge-water separation and return activated sludge to the aeration tank. The advanced treatment unit mainly includes a coagulation tank, flocculation tank, filter, and dosing system. Its main function is to further remove suspended solids and phosphorus to ensure that the effluent meets the standards.
[0038] Step S502: Construct the objective function for each agent based on the control performance impact metric score, and extract the coupling variables between agents from the causal directed graph structure. The objective function for each agent is a mathematical quantification of the control objective for that unit, used to evaluate the merits of different operational schemes during the game process; the objective function for the aeration unit, as the first agent, is... ;in, Energy consumption of the aeration unit The average control performance impact metric score for dissolved oxygen-related variables. For weighting coefficients; the purpose of this function is to minimize energy consumption while ensuring favorable dissolved oxygen control. The objective function of the secondary sedimentation unit, which is the second agent, is: ,in, This refers to the amount of sludge discharged; The average CPIM score for sludge settling-related variables; These are the weighting coefficients; the objective function aims to minimize excess sludge discharge while maintaining sludge settling performance. The objective function of the deep processing unit, acting as the third agent, is... ,in, To increase the dosage, The average CPIM score for total phosphorus-related variables; These are the weighting coefficients. The objective function aims to minimize chemical consumption while ensuring that the total phosphorus content in the effluent meets the standard.
[0039] Each agent optimizes its own objective function locally, exchanges information through coupling variables, iteratively solves for the Nash equilibrium point, and outputs the allowable range of operational variables for each fixed process unit. Coupling variables refer to variables that transmit influence between different fixed process units, making their objectives interrelated; for example, the return ratio of the aeration unit affects the hydraulic load of the secondary sedimentation unit, and the residual sludge discharge of the secondary sedimentation unit affects the sludge concentration of the aeration unit. Coupling variables are used to explicitly characterize the interactions between units, providing a basis for information exchange in distributed game theory. When obtaining coupling variables, firstly, all cross-unit directed edges are extracted from the causal directed graph structure, i.e., directed edges where the starting node belongs to one fixed process unit and the ending node belongs to another. Then, the starting or ending node corresponding to each cross-unit directed edge is marked as a coupling variable. Finally, all coupling variables are summarized to form a coupling variable list, and the unit pairs associated with each coupling variable are recorded. Iteratively solving for the Nash equilibrium point is a distributed optimization method that allows each agent to optimize its own objective function locally while gradually adjusting its strategy by exchanging coupling variable information until all agents no longer unilaterally change their strategies, i.e., equilibrium is reached. The specific process involves initializing the operational variables of each agent and repeating the following steps until convergence: Each agent, given the operational variables of other agents, solves for the optimal solution of its own objective function and broadcasts the values of coupling variables that affect other agents to the relevant agents; when the change in the operational variables of all agents is less than a preset threshold, the iteration stops, and the combination of operational variables at this point is the Nash equilibrium point.
[0040] The permissible range of manipulated variables is the range of values for each fixed process unit's manipulated variables obtained through game theory solutions. This includes the manipulated variable name, lower limit, and upper limit. For example, the permissible range for aeration rate is 1000~1200 m³ / h. The permissible range of manipulated variables provides global constraint boundaries for inner-level micro-domain control, ensuring that the total manipulated amount does not exceed the unit-level physical limitations and game equilibrium results when each micro-domain is controlled independently. The generation process first inputs the objective function and coupling variables of each agent into a distributed optimization solver; then, iteratively solving for the Nash equilibrium point using the alternating direction multiplier method yields the optimal manipulated variable values for each agent. Finally, an allowable range is set around the optimal value, typically taking [value missing]. ,in Based on the adjustment precision of the operating variables and the allowable fluctuation range of the process (e.g., the aeration rate is set in 5% increments), the allowable range of the operating variables of each unit is output as the result of the outer game.
[0041] The distributed optimization solver is a numerical calculation method for solving the Nash equilibrium point of a multi-agent system. In this application, each fixed process unit is defined as an independent agent, each with its own objective function and operating variables. Information is exchanged through coupling variables, and the Nash equilibrium point is solved iteratively, ultimately outputting the allowable range of operating variables for each unit. When using the distributed optimization solver, initial operating variable values are first set for each agent. Set convergence threshold Let the current iteration number be k; then each agent i operates on the variables of other agents. Under the given conditions, we solve the problem of minimizing our own objective function to obtain... Each agent broadcasts the values of its coupling variables (such as reflux ratio and sludge concentration) that affect other units to relevant agents. If the changes in the manipulated variables of all agents... If the condition is met, the iteration stops; otherwise, proceed to the next round. Finally, the converged combination of operational variables is taken as the Nash equilibrium point, and an allowable interval is set around it to obtain the allowable interval of the operational variables.
[0042] In practical engineering implementations, distributed optimization solvers can utilize mature optimization tools or frameworks such as CVXPY, a Python-based convex optimization modeling tool that supports the Alternating Direction Multiplier Method (ADMM) in distributed optimization and is suitable for small-scale multi-agent problems; PyTorch-based gradient-based distributed optimization algorithms suitable for scenarios where the objective function is differentiable; the ACADO Toolkit, an optimization solver for nonlinear model predictive control (NMPC) that supports distributed control; and the MATLAB Optimization Toolbox, which supports various optimization algorithms and is used for prototype verification and small-scale deployment. The distributed optimization solver described in this application is implemented using CVXPY in practical applications because it inherently supports distributed structures and is suitable for solving game equilibrium problems in multi-unit coupled systems such as wastewater treatment. In actual use, the appropriate optimization tool or framework can be selected based on the specific circumstances.
[0043] In this embodiment, the steps for dividing and generating dynamic control micro-domains within each intelligent agent include: Step S503: Divide the internal structure of each agent into time and space dimensions. Time dimension division yields N time segments, and space dimension division yields M spatial micro-domains. Time dimension division divides the historical operating state within a fixed process unit into several time periods with similar dynamic response characteristics, enabling the control strategy to adapt to changes in operating conditions over time. Time dimension division is based on the similarity of control-oriented feature spatial variables within adjacent time windows. First, the operating data of the most recent preset time period (e.g., 24 hours) is divided into multiple time windows (e.g., each window is 10 minutes). The mean vector and variance vector of the control-oriented feature spatial variables are calculated for each window. Then, the Euclidean distance between adjacent windows is calculated using a sliding window as the similarity. When the similarity is below a preset threshold (e.g., 0.7), it is marked as a time segment boundary. Finally, the time axis is divided into N consecutive time segments based on all boundaries. By dividing the spatial dimensions, the spatial location within a fixed process unit is divided into several micro-domains with similar process states, enabling the control strategy to adapt to spatial distribution differences. This division is based on the feature similarity of multi-point sensor data within the unit. The division process first collects real-time data from all sensors within the fixed process unit to form a spatial feature matrix, with each row corresponding to a sensor location and each column corresponding to a feature variable. The feature matrix is then normalized. Next, a density-based spatial clustering method is used to calculate the characteristic Euclidean distance between sensor locations. Sensor points with a distance less than a preset radius are grouped into the same spatial micro-domain, and the boundaries of the micro-domains are marked. Finally, M spatial micro-domains and the set of sensor locations they contain are output.
[0044] Step S504 involves cross-combining time segments and spatial micro-domains to generate a dynamic control micro-domain that includes time range, spatial range, current state, and historical response characteristics. When cross-combining time segments and spatial micro-domains, the N time segments obtained by dividing the time dimension are first... M spatial micro-domains obtained by dividing with spatial dimensions Perform a Cartesian product to generate N×M spatiotemporal combinations, where M and N are both integers; then for each time segment... and each spatial micro-domain Create dynamic control microdomains, meaning each combination corresponds to one dynamic control microdomain. The system records the time and spatial ranges of the input dynamic control microdomain. The time range includes the start and end times, and the spatial range includes the set of all sensor positions within the input dynamic control microdomain. Sensor data belonging to the microdomain at the current moment is extracted from the standardized multidimensional variable data vector as the current state. Historical response data of the microdomain under similar states is retrieved from the historical database as historical response characteristics. For example, time segment T1 and spatial microdomain S1 combine to form a microdomain. This represents the control unit within the S1 spatial region during the T1 time period; finally, the output is a set of all dynamic control microdomains, achieving joint decoupling of the time and space dimensions, giving each microdomain independent spatiotemporal attributes. A dynamic control microdomain is an independent control unit formed by the cross-combination of specific time segments and specific spatial microdomains. Through dynamic control microdomains, refined control within the unit is achieved, enabling the control strategy to adapt to both temporal changes and spatial distribution differences.
[0045] The current state refers to the operating condition of the dynamic control microdomain at the current moment, including the real-time values of all sensors within the microdomain and the characteristic variables (such as rate of change and fluctuation amplitude) calculated from these data. The current sensor values belonging to the spatial range of the microdomain are extracted from the standardized multidimensional variable data vector, and combined with the time range of the microdomain to confirm that the current moment is within the validity period of the microdomain. This serves as the input basis for the matching strategy in step S505, reflecting the real-time operating condition of the microdomain. Historical response characteristics refer to the record of response results after executing different strategies when the dynamic control microdomain was in a similar operating condition to the current state in the past, including adjustments to the executed operation variables, changes in effects, and response times. This provides empirical reference for strategy matching in step S505, improving the accuracy of the matching. Historical records similar to the current state of the microdomain (feature vector Euclidean distance less than a threshold) are retrieved from the historical database, and the strategy execution effect data are extracted from them.
[0046] In this embodiment, the steps for generating micro-domain control instructions by searching for matching strategies in the candidate strategy set and generating micro-domain control instructions, constrained by the globally optimal strategy and the allowed range of the operation variables, include: Step S505: For each dynamic control microdomain, based on the current state and historical response characteristics, retrieve a matching strategy from the candidate strategy set. The retrieval of matching strategies is based on the overlap between the candidate strategy's dependency variable set and the microdomain's feature variable set, as well as the consistency between the strategy's expected effect and the microdomain's current control objective. During the retrieval, first obtain the current state of the dynamic control microdomain and extract the microdomain's feature variable set. That is, the variables corresponding to all sensors within the micro-domain, and then traversing the candidate policy set. Each strategy in Calculate its set of dependent variables With micro-domain feature variable set The degree of overlap C, The strategy with an overlap of more than a preset ratio (e.g., 70%) is selected. Finally, among the selected strategies, they are further sorted according to the matching degree between the expected effect of the strategy and the current control objective of the micro-domain (e.g., increase DO, reduce energy consumption), and the strategy with the highest matching degree is selected as the candidate matching strategy.
[0047] The candidate matching strategy is a strategy highly relevant to the current micro-domain state, retrieved from the candidate strategy set. It includes operation instructions, expected effects, a set of dependent variables, and a strategy type flag, providing a targeted control scheme for the micro-domain and ensuring the strategy adapts to the local operating conditions of the micro-domain. (Dependent variable set) refers to strategy The set of key causal variables on which the policy depends is derived from the dependency variables recorded during policy generation. Micro-domain feature variable set. This refers to the set of variables corresponding to all sensors within a dynamically controlled microdomain, determined by the spatial extent of the microdomain. The settings of both the dependent variable set and the microdomain feature variable set are based on the causal directed graph structure and sensor deployment locations. By calculating the overlap between the two sets, the degree of matching between the strategy and the microdomain is determined; a higher overlap indicates a more suitable strategy for the microdomain. The preset ratio is set based on the minimum requirement for overlap between the strategy's dependent variable set and the microdomain feature variable set, typically set to 70%. Based on engineering experience and statistical laws, this ensures that most of the key variables the strategy depends on are within the observable range of the microdomain, avoiding execution deviations caused by missing policy dependencies. It is used to filter out strategies irrelevant to the microdomain state, improving matching accuracy.
[0048] Micro-domain control instructions are generated based on candidate matching strategies. These instructions are the final execution instructions generated for each dynamic control micro-domain, including the name of the operand variable, adjustment range, execution sequence, and execution micro-domain identifier. They are used to decompose the globally optimal strategy and the allowed game interval into executable local instructions, which are then distributed to the corresponding execution mechanisms. When generating micro-domain control instructions, for each dynamic control micro-domain, a strategy base that adjusts the same operand variable and in the same direction as the globally optimal strategy is selected from the matched candidate strategies. When multiple dynamic control micro-domains need to adjust the same operand variable simultaneously, they are sorted by micro-domain priority from high to low, and the operation amount is allocated sequentially from high to low priority to ensure that the total operation amount does not exceed the upper limit of the allowed operand variable interval. If the remaining operation amount is insufficient after allocation to a high-priority micro-domain, the low-priority micro-domain automatically reduces its adjustment range or abandons the adjustment. Finally, the coordinated micro-domain control instructions for each dynamic control micro-domain are distributed to the corresponding execution mechanism.
[0049] The priority of a micro-domain is determined by the degree of influence of the dynamic control micro-domain on the global objective, and this degree of influence is extracted from the causal directed graph structure. The greater the influence, the higher the priority. The micro-domain priority setting process first extracts all variable nodes within the micro-domain from the causal directed graph structure. Starting from each variable node, it traverses the causal paths of all variable nodes, counting the path endpoints affected by these causal paths, i.e., the number of control target variables and the path length. Then, it calculates the influence degree score: influence degree score = number of affected control target variables / (average path length + 1). Finally, all dynamic control micro-domains are sorted from highest to lowest influence degree score, with higher scores indicating higher priority.
[0050] In this embodiment, the implementation steps of collecting actual execution data after the micro-domain control command is executed, calculating the execution deviation between the actual execution data and the expected data of the global optimal strategy, and correcting and updating based on the deviation include: Step S601: Obtain the actual energy consumption change, actual effluent water quality change, and actual equipment response time after the micro-domain control command is executed, constituting the actual execution data. This actual execution data is used to evaluate the true effect of the micro-domain control command execution, serving as feedback information compared with the expected effect to correct the model and update parameters. It is obtained from feedback from the actuators, online monitoring instruments, and energy consumption metering equipment, specifically including the actual frequency / opening changes reported by the actuators (such as blowers and dosing pumps), effluent indicators collected by online water quality monitors, and energy consumption data recorded by electricity meters. Among these, the actual energy consumption change... Data is obtained from electricity meters and flow meters. The cumulative power consumption of relevant equipment is recorded over a fixed time period (e.g., 5 minutes) before and after the execution of micro-domain control commands. The difference is calculated and normalized to obtain the energy consumption change, which is used to verify whether the energy-saving effect of the strategy meets expectations. Actual effluent water quality change. Data was obtained from the online monitoring instrument at the effluent outlet. Concentrations of COD, ammonia nitrogen, total nitrogen, and total phosphorus in the effluent were recorded before and after the execution of control commands. The changes in each indicator were calculated, and a weighted sum was obtained to determine the overall water quality change. This was used to verify whether the strategy's impact on effluent quality met expectations. Actual equipment response time. By obtaining feedback signals from the executing agency, and recording the time required from the issuance of the instruction to the execution agency's feedback to reach 90% of the target adjustment value, in seconds, the timeliness of the strategy execution is reflected, providing a basis for response speed assessment.
[0051] Step S602: Calculate the execution deviation between the actual execution data and the expected data of the global optimal strategy. Execution deviation refers to the difference between the actual execution data and the expected data of the global optimal strategy, and includes energy consumption deviation, water quality deviation, and response time deviation. Energy consumption deviation ;in, To represent the actual change in energy consumption; This represents the expected change in energy consumption, which is the pre-set energy consumption change value in the globally optimal strategy. It is obtained from the expected effects during the generation of the candidate strategy set in step S300. The expected energy consumption variable for rule-based candidate strategies is estimated based on historical effects, while the expected energy consumption change for learning-based strategies is determined through the reward function in step S302. Obtain.
[0052] Water quality deviation ;in, To represent the actual change in effluent water quality; This represents the expected change in effluent water quality, i.e., the pre-defined degree of water quality improvement in the globally optimal strategy. It is obtained from the expected effects during the generation of the candidate strategy set in step S300. The expected change in energy consumption for rule-based candidate strategies is obtained based on historical effects along the causal path; the expected change in energy consumption for learning-based strategies is obtained through the reward function in step S302. .
[0053] Response time deviation ;in, To indicate the actual device response time; This represents the expected device response time, which is the expected device response time in the globally optimal strategy. It is obtained from the expected effect when the candidate strategy set is generated in step S300 and is set by the average response time of similar strategies in history.
[0054] By quantifying the discrepancy between the performance of the deviation quantification strategy and the expected results, a basis for model correction can be provided.
[0055] If the execution deviation exceeds the deviation threshold, the causal path observed during the actual execution process is compared with the directed graph structure of the causal relationship. Based on the comparison results, the confidence of the causal edges in the directed graph structure is adjusted. The deviation threshold is set based on the process's allowable error range and sensor accuracy, typically taking 10% to 20% of the expected value. The specific value is adjusted according to the importance of the control objective; for example, the energy consumption deviation threshold is set to 10% of the expected value, and the water quality deviation threshold is set to 30% of the allowable fluctuation of the emission standard, used to determine whether the deviation is significant enough to trigger model correction.
[0056] When comparing the observed causal paths with the directed graph structure of causal relationships during actual execution, the actual response sequence after the change of the operational variable is first extracted from the actual execution data. For example, "after the aeration rate increases by 5%, DO increases by 0.2 mg / L within 1 minute, while the DO consumption rate decreases by 10%", forming the observed causal path. Then, the observed causal path is compared segment by segment with the corresponding causal edges in the directed graph structure of causal relationships to check whether each directed edge exists in the graph and whether the causal direction is consistent. Finally, the edges that exist in the graph but are not observed and the edges that are observed but do not exist in the graph are recorded.
[0057] When adjusting the confidence level of causal edges in a directed graph structure based on alignment results, the adjustment criteria include: For causal edges observed in the graph, increase their confidence level. The increase is positively correlated with the number of observations and the degree of consistency, for example, by 0.05 for each observation. For a causal edge that is not present in the graph but appears multiple times (e.g., 3 times consecutively), add the directed edge to the graph with an initial confidence level of 0.5. For causal edges that exist in the graph but are not observed in multiple executions (e.g., not appearing in 5 consecutive executions), reduce their confidence by 0.1 each time. When the confidence is below 0.2, mark the edge as "weakly correlated" or remove it from the graph. Finally, after adjusting all confidence levels, the graph is renormalized so that the sum of the confidence levels of all causal edges remains unchanged.
[0058] Step S603: Store the success rate of this strategy execution in the historical database and adjust the weighting coefficients in the overall credibility score. The success rate of this strategy execution refers to the degree to which the strategy execution effect meets the expected requirements. It is used to quantify the success rate of a single execution and serves as the update content for the historical database. This is achieved by obtaining the energy consumption deviation calculated in step S602. Water quality deviation and response time deviation Divide each deviation by the corresponding deviation threshold to obtain the normalized deviation, take the average to obtain the comprehensive deviation rate, and subtract the comprehensive deviation rate from 1 to obtain the success rate. ,Right now Each deviation has been normalized to the [0, 1] interval. The historical database includes a historical record of each strategy execution, specifically including strategy ID, execution time, micro-domain state (feature vector) at execution, operation instructions, actual execution effects (energy consumption changes, water quality changes, response time), success rate, and dependency variable set. The historical database provides data support for calculating historical success rates in strategy credibility assessment, provides actual response sequence references for causal graph correction, and provides a basis for retrieving historical response characteristics of dynamic control micro-domains.
[0059] When adjusting the weighting coefficients in the overall credibility score, the weight of each score component is increased based on the positive correlation between the success rate of the current strategy execution and the score components. In other words, the component that more accurately reflects the true reliability of the strategy is weighted accordingly. The correlation coefficient between the strategy execution effect and the interpretability score (based on multiple historical execution records) is calculated. If the correlation is positive and significant, the first weighting coefficient is increased. The increase is 0.01; if it is a negative correlation, it decreases. Similarly, the second weighting coefficient is adjusted based on the correlation between the performance and the average CPIM score of the dependent variables. The third weighting coefficient is adjusted based on the positive correlation between execution effectiveness and historical success rate. Finally, the three adjusted weight coefficients are normalized so that their sum is 1. The updated weight coefficients are used for the comprehensive credibility score calculation of subsequent strategies, thus realizing the self-evolution of the system.
[0060] like Figure 3 As shown, a wastewater treatment control system based on multidimensional variable data analysis is provided. The system includes: The data acquisition and processing module is used to collect and preprocess the raw multi-dimensional variable data of the entire wastewater treatment process, and generate a standardized multi-dimensional variable data vector. The causal identification and feature measurement module, based on standardized multidimensional variable data vectors, calculates and determines the direction of causal relationships between variables, constructs a directed graph structure of causal relationships, calculates the control performance impact measurement score of each variable in the directed graph structure of causal relationships, and generates a control-oriented feature space. The candidate control strategy generation module is used to extract causal paths in the causal relationship directed graph structure to generate rule-based candidate strategies, perform reinforcement learning on real-time variable values based on the control-oriented feature space to generate learned candidate strategies, and merge the rule-based candidate strategies and learned candidate strategies to obtain a candidate strategy set. The strategy credibility evaluation module is used to calculate the comprehensive credibility score of each strategy in the candidate strategy set. From the strategies with a comprehensive credibility score higher than the credibility threshold, the strategy with the highest comprehensive credibility score is selected as the global optimal strategy. The dual-layer collaborative control module is used to define each fixed process unit in the entire wastewater treatment process as an intelligent agent, calculate the allowable range of the operation variables of each intelligent agent, divide and generate dynamic control micro-domains within each intelligent agent, and search for matching strategies in the candidate strategy set with the global optimal strategy and the allowable range of operation variables as constraints to generate micro-domain control instructions. The feedback and evolution module is used to collect the actual execution data after the micro-domain control command is executed, calculate the execution deviation between the actual execution data and the expected data of the global optimal strategy, and make corrections and updates based on the execution deviation.
[0061] In this embodiment, deep coupling of data analysis and control is achieved through causal inference and control performance impact measurement. Using a constructed directed graph structure of causal relationships, the causal relationships between variables in the wastewater treatment process are quantitatively characterized. Based on this, a control performance impact measurement for each variable is calculated. This measurement comprehensively evaluates the weight of the variable's influence on system stability, response speed, and robustness. Unlike traditional methods that use analysis results only to generate optimized setpoints, this invention directly maps the causal structure to the dynamic topology of the controller, enabling the control law to adaptively adjust with changes in causal relationships. Simultaneously, the control performance impact measurement filters out characteristic variables beneficial to the control process, constructing a control-oriented feature space. This avoids control oscillations caused by selecting features solely for prediction accuracy. The causal inference results are not only used for strategy generation but also directly participate in strategy credibility assessment and game coordination processes, forming a fully closed-loop deep coupling mechanism from data analysis to control execution and effect feedback. This significantly improves the stability, response speed, and robustness of the control system.
[0062] This invention constructs a two-layer collaborative control architecture consisting of an outer-layer multi-agent game-theoretic collaboration and an inner-layer spatiotemporal dynamic control domain partitioning. In the outer layer, each fixed process unit is defined as an agent with an independent objective function. A distributed game is used to solve for the Nash equilibrium points among the agents, explicitly modeling the competitive and cooperative relationships between units and outputting the allowable range of operational variables for each unit, thus achieving dynamic equilibrium under objective conflict. In the inner layer, each unit is partitioned in terms of time and clustered in terms of space to generate dynamic control micro-domains. Each micro-domain is matched with an independent control strategy based on its current state and historical response characteristics, and a coordination mechanism between micro-domains ensures that the total amount of operations does not exceed the allowable range given by the outer-layer game. This two-layer architecture achieves explicit coordination of objective conflicts between units and adaptive control of spatiotemporal changes within units, ensuring both global collaborative optimization and local fine-grained regulation, thereby improving the overall operating efficiency and disturbance resistance of the wastewater treatment system.
[0063] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code, which, when executed by the one or more processors, can perform the wastewater treatment control method and system based on multidimensional variable data analysis as described above.
[0064] The methods and systems according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the wastewater treatment control method and system based on multidimensional variable data analysis provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs.
[0065] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0066] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A wastewater treatment control method based on multidimensional variable data analysis, characterized in that, The method includes: Collect and preprocess raw multidimensional variable data from the entire wastewater treatment process to generate standardized multidimensional variable data vectors; A causal directed graph structure is constructed based on standardized multidimensional variable data vectors. The control performance impact metric score of each variable in the causal directed graph structure is calculated, and a control-oriented feature space is generated. The causal paths in the directed graph structure of causal relationships are extracted to generate regular candidate policies. Reinforcement learning is performed on real-time variable values based on the control-oriented feature space to generate learned candidate policies. The regular candidate policies and learned candidate policies are merged to obtain a candidate policy set. Calculate the overall credibility score for each strategy in the candidate strategy set, and select the globally optimal strategy based on the overall credibility score; Each fixed process unit in the entire wastewater treatment process is defined as an intelligent agent. The allowable range of the operation variable is calculated for each intelligent agent, and a dynamic control micro-domain is divided. With the global optimal strategy and the allowable range of the operation variable as constraints, a matching strategy is searched in the candidate strategy set to generate micro-domain control instructions. Collect actual execution data after micro-domain control commands are executed, calculate the execution deviation between the actual execution data and the expected data of the global optimal strategy, and make corrections and updates based on the execution deviation.
2. The wastewater treatment control method based on multidimensional variable data analysis according to claim 1, characterized in that, The process involves constructing a causal directed graph structure based on standardized multidimensional variable data vectors, calculating the control performance impact metric score for each variable in the causal directed graph structure, and generating a control-oriented feature space, including: Obtain standardized multidimensional variable data vectors, calculate the time-series correlation coefficient and partial correlation coefficient between each variable, determine the direction of causal relationship between variables based on the time-series correlation coefficient and partial correlation coefficient, and construct a causal relationship directed graph structure based on the causal relationship direction and variables; For each variable in the causal directed graph structure, calculate the stability impact score, response speed impact score, and robustness impact score, and then sum them by weight to obtain the control performance impact measure score for each variable. The control performance impact metric scores are sorted from high to low, and a preset number of variables corresponding to the control performance impact metric scores are selected to form a control-oriented feature space.
3. The wastewater treatment control method based on multidimensional variable data analysis according to claim 1, characterized in that, The process of merging rule-based candidate strategies and learning-based candidate strategies to obtain a candidate strategy set includes: Traverse all nodes in the causal directed graph structure, mark the operation variable nodes and the control target variable nodes, and for each pair of operation variable nodes and control target variable nodes, find all directed paths from the operation variable node to the control target variable node. Intermediate variables are extracted from each directed path to form a causal path. Policy rules are generated based on the causal paths. The set of policy rules for all directed paths forms a rule-based candidate policy. The current values of all variables in the control-guided feature space and the time values of the previous preset number are used to form a state vector, and the adjustment range of the operation variables is defined to form a discrete action set. Using a set of discrete actions as the dynamic space and in conjunction with state vectors, a reward function is constructed with a score that measures the impact of control performance, and a learning-oriented candidate policy is generated through training.
4. The wastewater treatment control method based on multidimensional variable data analysis according to claim 1, characterized in that, The calculation of the comprehensive credibility score for each strategy in the candidate strategy set, and the selection of the globally optimal strategy based on the comprehensive credibility score, includes: Obtain a set of candidate strategies, and calculate the path clarity score, variable reliability score, and uncertainty score for each strategy in the set. The interpretability score for each strategy is obtained by multiplying the weighted sum of the path clarity score and the variable reliability score by the uncertainty penalty coefficient. The average control performance impact metric, historical success rate, and interpretability score of each strategy's dependent variables are obtained and weighted and summed to obtain a comprehensive credibility score for each strategy. The credibility threshold is dynamically adjusted based on the current system stability. Strategies with a comprehensive credibility score higher than the current credibility threshold are selected, and the strategy with the highest comprehensive credibility score is chosen as the globally optimal strategy.
5. The wastewater treatment control method based on multidimensional variable data analysis according to claim 4, characterized in that, The dynamic adjustment of the credibility threshold based on the current system stability includes: Calculate the average variance of each variable in the control-guided feature space over the most recent preset time period, and calculate the current system stability based on the average variance. The credibility threshold is calculated based on the preset base threshold and the current system stability.
6. The wastewater treatment control method based on multidimensional variable data analysis according to claim 1, characterized in that, The process involves defining each fixed process unit in the entire wastewater treatment process as an intelligent agent, calculating the allowable range of operational variables for each intelligent agent, and dividing the dynamic control micro-domain, including: The fixed process unit includes an aeration unit, a secondary sedimentation unit, and a deep treatment unit, which are respectively defined as a first intelligent agent, a second intelligent agent, and a third intelligent agent. The objective function for each agent is constructed based on the control performance impact metric score, and the coupling variables between agents are extracted from the causal directed graph structure. Each agent optimizes its own objective function locally, exchanges information through coupling variables, iteratively solves for the Nash equilibrium point, and outputs the allowable range of the operating variables for each fixed process unit.
7. The wastewater treatment control method based on multidimensional variable data analysis according to claim 6, characterized in that, The process of dividing the dynamic control microdomain includes: Each agent is divided into time and space dimensions. The time dimension division yields N time segments, and the space dimension division yields M spatial micro-domains. By cross-combining time segments with spatial micro-domains, a dynamic control micro-domain is generated that includes time range, spatial range, current state, and historical response characteristics.
8. The wastewater treatment control method based on multidimensional variable data analysis according to claim 7, characterized in that, The process of retrieving a matching strategy from the candidate strategy set and generating micro-domain control instructions, constrained by the globally optimal strategy and the allowed range of the operation variables, includes: For each dynamic control microdomain, a matching strategy is retrieved from the candidate strategy set based on the current state and historical response characteristics; Among the matched strategies, select the one that adjusts the same operational variable and in the same direction as the globally optimal strategy. When multiple dynamic control microdomains need to adjust the same operational variable at the same time, the operational quantity is allocated in descending order of microdomain priority to ensure that the total operational quantity does not exceed the upper limit of the allowable range of the operational variable. The coordinated micro-domain control commands for each dynamic control micro-domain are then sent to the corresponding actuators.
9. The wastewater treatment control method based on multidimensional variable data analysis according to claim 1, characterized in that, The process involves collecting actual execution data after the micro-domain control command is executed, calculating the execution deviation between the actual execution data and the expected data of the global optimal strategy, and correcting and updating based on the deviation, including: The actual execution data consists of the actual energy consumption change, the actual effluent water quality change, and the actual equipment response time after the micro-domain control command is executed. Calculate the execution deviation between the actual execution data and the expected data of the globally optimal strategy; If the execution deviation exceeds the deviation threshold, the causal path observed during the actual execution process is compared with the causal relationship directed graph structure, and the confidence of the causal edges in the causal relationship directed graph structure is adjusted according to the comparison results. The success rate of this strategy will be stored in the historical database, and the weighting coefficients in the overall credibility score will be adjusted accordingly.
10. A wastewater treatment control system based on multidimensional variable data analysis, characterized in that, The system includes: The data acquisition and processing module is used to collect and preprocess the raw multi-dimensional variable data of the entire wastewater treatment process and generate standardized multi-dimensional variable data vectors. The causal identification and feature measurement module constructs a causal relationship directed graph structure based on standardized multidimensional variable data vectors, calculates the control performance impact measurement score of each variable in the causal relationship directed graph structure, and generates a control-oriented feature space. The candidate control strategy generation module is used to generate rule-based candidate strategies based on the causal directed graph structure, generate learning-based candidate strategies based on the control-oriented feature space, and merge the rule-based candidate strategies and learning-based candidate strategies to obtain a candidate strategy set. The strategy credibility evaluation module is used to calculate the comprehensive credibility score of each strategy in the candidate strategy set, and select the globally optimal strategy based on the comprehensive credibility score; The dual-layer collaborative control module is used to define each fixed process unit in the entire wastewater treatment process as an intelligent agent, calculate the allowable range of operation variables for each intelligent agent, divide the dynamic control micro-domain, and generate micro-domain control instructions. The feedback and evolution module is used to collect the actual execution data after the micro-domain control command is executed and calculate the execution deviation, and make corrections and updates based on the execution deviation.