A sewage treatment plant hybrid intelligent optimization decision method and system
Patent Information
- Application Number
- CN202610507265.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-04-17
AI Technical Summary
[0006]本发明的目的是提供一种污水处理厂混合智能优化决策方法及系统,通过双层控制架构实现全厂级资源协同分配与处理单元级快速响应控制的有机结合,解决现有技术中优化决策滞后、理论最优参数难以落地执行以及系统稳定性和鲁棒性不足等技术问题
[0071] 1. This invention adopts a two-layer control architecture. The centralized optimization layer is responsible for setting plant-wide goals and allocating resource quotas, while the distributed autonomous layer is responsible for optimal autonomous control at the unit level, thus realizing an organic combination of global resource optimization and local rapid response.
Smart Images

Figure CN122043972B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment technology, and in particular to a hybrid intelligent optimization decision-making method and system for wastewater treatment plants, applicable to the intelligent operation control and optimization management of urban wastewater treatment plants. Background Technology
[0002] Wastewater treatment plants, as a crucial component of urban infrastructure, play a vital role in ensuring environmental safety and resource conservation through efficient, stable, and low-cost operation. With the development of wastewater treatment technology and the increase in automation, intelligent control and optimization decision-making systems are playing an increasingly important role in the operation and management of wastewater treatment plants.
[0003] Currently, the automated control systems of wastewater treatment plants are mainly divided into two typical modes: one is a centralized control system based on a central control room, which realizes unified monitoring and adjustment of parameters of the entire plant through a SCADA (Supervisory and Data Acquisition) system; the other is a decentralized system in which each treatment unit is independently controlled, such as a PLC-based dissolved oxygen control system for biological treatment tanks and an automatic operation system for membrane systems.
[0004] Existing advanced intelligent optimization systems for wastewater treatment plants typically employ a single central optimization decision-making model. This model collects operational data from all treatment units across the plant and performs overall simulation and optimization calculations using complex mathematical models (such as the activated sludge model ASM and membrane filtration models) to generate optimal control parameter setpoints, which are then distributed to each execution unit. While this method can theoretically achieve global optimization, it faces several key technical challenges: First, the plant-wide optimization model has high computational complexity, requiring consideration of the interactions between numerous variables, resulting in long computation times. Second, due to the time delay between the central decision-making and execution units, it is difficult to respond quickly to sudden situations such as fluctuations in water quality and quantity. Furthermore, a single model is insufficient to adapt to the specific needs of each treatment unit, lacking versatility and adaptability.
[0005] These technical deficiencies often lead to problems in actual operation, such as optimization decisions lagging behind changes in operating conditions, difficulty in implementing theoretically optimal parameters, and insufficient system stability and robustness. They cannot simultaneously meet the needs of plant-wide resource collaborative optimization and unit-level rapid autonomous control, thus restricting the improvement of the overall operating efficiency and economic benefits of wastewater treatment plants. Summary of the Invention
[0006] The purpose of this invention is to provide a hybrid intelligent optimization decision-making method and system for wastewater treatment plants. Through a two-layer control architecture, it achieves an organic combination of plant-wide resource collaborative allocation and treatment unit-level rapid response control, solving technical problems in the prior art such as lagging optimization decisions, difficulty in implementing theoretically optimal parameters, and insufficient system stability and robustness.
[0007] To achieve the above objectives, the present invention provides a hybrid intelligent optimization decision-making method for wastewater treatment plants, comprising:
[0008] Obtain process flow data and control requirement data of the wastewater treatment plant, and design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture.
[0009] Based on the hierarchical control architecture, historical operating data and real-time status data of each processing unit are collected and preprocessed and feature extracted to construct a global optimization model, which reflects the interaction relationship between each processing unit.
[0010] Based on the global optimization model, the plant-wide resource optimization allocation scheme and control parameter constraints are calculated. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain the global optimization decision.
[0011] Edge controllers are deployed in each processing unit. Based on the global optimization decision, local sensor data is collected, optimization algorithms are run within the constraints of the control parameters, unit-level real-time control commands are obtained, and the operation of each processing unit is controlled.
[0012] The system collects operational status information after execution and reports it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer then makes optimization decisions and adjustments for the next cycle, thereby achieving collaborative control between the centralized optimization layer and the distributed autonomous layer.
[0013] Furthermore, the process flow data and control requirement data of the wastewater treatment plant are acquired, and a two-layer control architecture is designed based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture, including:
[0014] Acquire process flow data and control requirement data of the wastewater treatment plant. The process flow data includes the connection relationship of each treatment unit, the material flow path and energy consumption characteristics. The control requirement data includes effluent water quality standards, treatment capacity requirements and safe operation constraints, to obtain the basic data of the wastewater treatment plant.
[0015] Based on the basic data of the wastewater treatment plant, the plant-wide resource coordination needs and unit-level real-time response needs are analyzed to determine the system's functional requirements. These functional requirements include global resource collaborative allocation needs, local rapid response control needs, and abnormal situation handling needs, resulting in a list of system functional requirements.
[0016] Based on the system functional requirements list, a centralized optimization layer is designed to be responsible for plant-wide target setting and resource allocation, and a distributed autonomous layer is designed to be responsible for unit-level optimal autonomous control, resulting in a two-layer control architecture and responsibility boundary definition.
[0017] Based on the aforementioned two-layer control architecture and responsibility boundary definition, the hardware configuration and installation locations of the central server, edge computing devices, and communication network are planned to obtain the hierarchical control architecture.
[0018] Furthermore, the collection of historical operating data and real-time status data of each processing unit, and the subsequent data preprocessing and feature extraction, include:
[0019] Historical operating data and real-time status data are collected from the sensors and control systems of each treatment unit in the wastewater treatment plant. The historical operating data and real-time status data include water quality parameters, equipment operating parameters and energy consumption data to obtain the raw dataset.
[0020] The original dataset is subjected to missing value imputation, outlier detection, and data format standardization to obtain a cleaned dataset.
[0021] The cleaned dataset is then subjected to sensitivity classification identification, and divided into high-sensitivity data, medium-sensitivity data, and low-sensitivity data to obtain the classified dataset.
[0022] Based on the graded dataset, time features, frequency domain features, and statistical features are extracted to obtain preprocessed plant-wide data.
[0023] Furthermore, the construction of a global optimization model, which reflects the interaction relationships between processing units, includes:
[0024] Based on the preprocessed plant-wide data, a process network diagram is constructed to represent the material flow, energy flow, and information flow relationships between each processing unit. The process network diagram includes nodes and edges, where nodes represent processing units and edges represent the connection relationships between units, thus obtaining a wastewater treatment process network diagram.
[0025] Based on the wastewater treatment process network diagram, a simplified activated sludge model and a membrane filtration model are used to establish a local dynamic model for each treatment unit, resulting in a simplified model for each unit.
[0026] Based on the simplified models of each unit and the coupling relationships between units, model parameters are identified through data-driven machine learning methods, including neural networks or support vector machines, to obtain the global optimization model.
[0027] Furthermore, the calculation of the plant-wide resource optimization allocation scheme and control parameter constraints, wherein the plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain a global optimization decision, includes:
[0028] Based on the global optimization model, the influent water quality and influent water volume for future periods are predicted to obtain influent prediction data.
[0029] Based on the influent prediction data and effluent water quality requirements, the energy consumption budget and chemical consumption budget of each treatment unit are calculated by an optimization algorithm. The optimization algorithm aims to minimize the total cost and obtains a resource budget allocation scheme for each unit.
[0030] Based on the resource budget allocation scheme and process safety constraints of each unit, the reasonable setting range of each key control parameter is calculated. The key control parameters include dissolved oxygen concentration, reverse osmosis system recovery rate and reagent dosage, and the global optimization decision is obtained.
[0031] Further, obtaining the global optimization model includes:
[0032] Based on the wastewater treatment process network diagram, tree width analysis is performed on the wastewater treatment process network diagram to identify the key coupling points and bottlenecks of the system and obtain the tree width parameters of the process network.
[0033] Based on the tree width parameter of the process network, the global optimization model is decomposed into several sub-models. Each sub-model corresponds to a cluster of tree decomposition. Tree decomposition ensures that the connection between sub-models follows a tree structure, resulting in a set of sub-models after tree decomposition.
[0034] Based on the set of sub-models after tree decomposition, a hierarchical optimization model is constructed. The top-level node of the hierarchical optimization model corresponds to the macro-decision variables of the entire plant, the intermediate-level nodes correspond to the collaborative optimization variables of the processing unit group, and the leaf nodes correspond to the micro-control variables of a single processing unit. This results in a hierarchical optimization model based on a tree structure, which serves as the global optimization model.
[0035] Furthermore, after obtaining the global optimization model, the process further includes:
[0036] Based on the set of sub-models after tree decomposition and the wastewater treatment process network diagram, tree coverage metrics are calculated, including node coverage, edge coverage, and critical path retention rate, to obtain tree coverage evaluation results.
[0037] Based on the tree coverage assessment results and computational resource constraints, the tree width parameter of the process network is dynamically adjusted. When the optimization period is shorter than the preset response time threshold, the tree width parameter value is reduced to reduce computational complexity. When the optimization period is longer than the preset response time threshold, the tree width parameter value is increased to improve model accuracy, thus obtaining the adaptively adjusted tree width parameter.
[0038] Based on the adaptively adjusted tree width parameter, the step of decomposing the global optimization model into several sub-models is re-executed to obtain an updated hierarchical optimization model based on tree structure.
[0039] The updated tree-based hierarchical optimization model is mapped to a centralized optimization layer and a distributed autonomous layer. The optimization of the top-level nodes and intermediate nodes of the updated tree-based hierarchical optimization model is implemented in the centralized optimization layer, and the optimization of the leaf nodes is implemented in the distributed autonomous layer, resulting in the mapped tree-based hierarchical optimization model, which serves as the global optimization model.
[0040] Furthermore, the reporting to the centralized optimization layer via a two-way communication mechanism includes:
[0041] Based on the running status information after execution and the hierarchical dataset, differential privacy sensitivity is calculated for each type of sensitive data. The differential privacy sensitivity represents the maximum impact of a change in a single record on the query results, thus obtaining the sensitivity parameters for each type of data.
[0042] Based on the sensitivity parameters and privacy budget parameters of the various types of data, the standard deviation of Gaussian noise is calculated, and a time-adaptive noise calibration mechanism is designed to obtain the calibrated Gaussian noise parameters.
[0043] Based on the running status information after execution and the calibrated Gaussian noise parameters, Gaussian noise is applied before the edge controller reports sensitive data. Different levels of privacy protection are applied to data with different sensitivity levels to obtain privacy-protected status reporting data.
[0044] Based on the privacy-protected status reporting data, it is transmitted to the centralized optimization layer through the communication network to realize privacy-protected status reporting.
[0045] Furthermore, after applying Gaussian noise before reporting sensitive data to the edge controller, applying different levels of privacy protection to data with different sensitivity levels, and obtaining privacy-protected status reporting data, the process further includes:
[0046] Based on the privacy-protected state reporting data, a noise-aware optimization objective function is run in the centralized optimization layer to incorporate data uncertainty into the decision-making process, thereby obtaining an optimization objective that takes noise into account.
[0047] Based on the optimization objective that takes noise into account, a probabilistic inference method is used to estimate and filter out noise, calculate the decision confidence interval, and obtain a near-optimal decision result.
[0048] Based on the near-optimal decision results, the decision quality loss index is calculated by comparing the difference in the optimization objective function values before and after privacy protection, and the decision error evaluation results are obtained.
[0049] Based on the decision error assessment results, a quantitative relationship model between privacy protection strength and decision quality is established. The calibrated Gaussian noise parameters are dynamically adjusted according to the comparison between the decision quality loss index and the preset quality loss threshold to obtain an optimized decision that balances privacy protection and decision quality, which serves as the global optimized decision for the next cycle.
[0050] Furthermore, the step of running an optimization algorithm within the constraints of the control parameters to obtain real-time control commands at the unit level includes:
[0051] Based on the local sensor data and the control parameter constraints, the core treatment units affecting the effluent quality are identified. The core treatment units include a biological reactor, a membrane filtration unit, and a disinfection unit, thus obtaining a set of key treatment units.
[0052] Based on the control problem of the set of key processing units, the discretized control parameter values within the set interval of the control parameter constraints are modeled as a matroid structure. The basic element set is defined as the set of control parameter values discretized within the set interval of the control parameter constraints according to a preset step size. The independent set family is defined as the set of parameter combinations that meet the process requirements, thus obtaining the matroid model of the wastewater treatment process.
[0053] Based on the aforementioned wastewater treatment process matroid model, matroids corresponding to different optimization objectives are constructed, including energy efficiency matroid, treatment efficiency matroid, and stability matroid. Through matroid cross-operation, parameter combinations that simultaneously satisfy multiple constraints are found, resulting in a multi-objective matroid structure.
[0054] Based on the multi-objective matroid structure, a parallel basis search algorithm is run to find a set of mutually independent basis, each basis corresponding to an optimization dimension, resulting in multiple parallel basis;
[0055] Based on the multiple parallel bases, each parallel base is assigned to different computing threads for parallel computation. The optimization results of each parallel base are combined through a weighted voting-based fusion algorithm to obtain the real-time control instructions at the unit level.
[0056] Furthermore, the process of assigning each parallel basis to different computing threads for parallel computation, and then combining the optimization results of each parallel basis through a weighted voting-based fusion algorithm to obtain the unit-level real-time control instructions, includes:
[0057] Based on the computing resources of the multiple parallel bases and the edge controller, the computing resource allocation of each parallel base is dynamically adjusted according to the process status and control requirements, so as to allocate more computing resources to the bases corresponding to key process parameters, thereby obtaining an adaptive resource allocation scheme.
[0058] According to the adaptive resource allocation scheme, an information sharing protocol and a co-evolution mechanism are realized among parallel bases, enabling each base to learn from and improve each other during the iteration process, and obtain the co-optimized parallel base results;
[0059] Based on the parallel basis results after collaborative optimization, an anomaly detection mechanism is used to identify and exclude basis results that deviate too much. The contribution of each parallel basis is evaluated using a Shapley value-based method, the weight of each parallel basis in the weighted voting is determined, and the real-time control command at the unit level is obtained.
[0060] Furthermore, it also includes fault tolerance and emergency response procedures, including:
[0061] The communication status between the centralized optimization layer and the distributed autonomous layer is detected. When a communication interruption is detected, the autonomous operation mode of the distributed autonomous layer is started. Based on the historical operation data stored in the edge controller and the control parameter constraints received last time, the autonomous control command under the communication interruption is obtained.
[0062] Based on the prediction deviation between the prediction results of the global optimization model and the actual measured values, the system detects whether the global optimization model has failed. When the prediction deviation exceeds a preset threshold, the system switches to a rule-based control mode to obtain a safety control command for model failure.
[0063] The system monitors the operating status of hardware components, including a central server, edge controllers, sensors, and actuators. When a failure is detected in any of the hardware components, a backup system switching mechanism is activated, and detailed system operation logs are recorded to achieve fault-tolerant operation of the system.
[0064] The present invention also provides a hybrid intelligent optimization decision-making system for wastewater treatment plants, comprising:
[0065] The architecture design module is used to acquire process flow data and control requirement data of the wastewater treatment plant, and to design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture.
[0066] The model building module is used to collect historical operating data and real-time status data of each processing unit based on the hierarchical control architecture, perform data preprocessing and feature extraction, and build a global optimization model that reflects the interaction relationship between each processing unit.
[0067] The centralized optimization module is used to calculate the plant-wide resource optimization allocation scheme and control parameter constraints based on the global optimization model. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain the global optimization decision.
[0068] The edge control module is used to deploy edge controllers in each processing unit, collect local sensor data based on the global optimization decision, run optimization algorithms within the range of the control parameter constraints, obtain real-time control commands at the unit level, and control the operation of the devices in each processing unit.
[0069] The communication and collaboration module is used to collect the running status information after execution and report it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer makes optimization decisions and adjustments for the next cycle, realizing collaborative control between the centralized optimization layer and the distributed autonomous layer.
[0070] The beneficial effects of this invention are:
[0071] 1. This invention adopts a two-layer control architecture. The centralized optimization layer is responsible for setting plant-wide goals and allocating resource quotas, while the distributed autonomous layer is responsible for optimal autonomous control at the unit level, thus realizing an organic combination of global resource optimization and local rapid response.
[0072] 2. This invention reduces computational complexity and improves decision-making speed by using a lightweight global optimization model, thus solving the problem of long computation time in traditional global optimization models;
[0073] 3. This invention deploys edge controllers in each processing unit, enabling local fast response control and effectively solving the problem of time delay between the central decision-making and execution units;
[0074] 4. This invention achieves collaborative control between the centralized optimization layer and the distributed autonomous layer through a two-way communication mechanism, thereby improving the system's adaptability and robustness;
[0075] 5. This invention designs a fault-tolerant and emergency handling mechanism, which can maintain stable system operation in abnormal situations such as communication interruption, model failure or hardware failure, thereby improving the reliability of the system. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1This is a flowchart illustrating the hybrid intelligent optimization decision-making method for wastewater treatment plants provided in an embodiment of the present invention;
[0078] Figure 2 This is a flowchart illustrating the design of the two-layer control architecture in an embodiment of the present invention;
[0079] Figure 3 This is a schematic diagram of the structure of the hybrid intelligent optimization decision-making system for wastewater treatment plants provided in an embodiment of the present invention. Detailed Implementation
[0080] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0081] Example 1
[0082] like Figure 1 As shown, this embodiment provides a hybrid intelligent optimization decision-making method for wastewater treatment plants, including:
[0083] Step S1: Obtain the process flow data and control requirement data of the wastewater treatment plant, and design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture.
[0084] Specifically, this step begins with collecting and organizing the wastewater treatment plant's process flow data, including static data such as process flow diagrams, equipment layout diagrams, pipeline connections, treatment unit specifications, and design capacity. Multiple methods are employed to collect this data: directly extracting parameters and diagrams from engineering design documents; conducting on-site surveys of actual equipment layouts and pipeline connections; and interviewing process engineers to obtain undocumented process knowledge and experience. This data is then standardized to establish a unified process flow data model for subsequent control system design.
[0085] Simultaneously, control requirement data is collected, including treatment volume requirements, effluent quality standards, energy consumption indicators, operating cost targets, and special treatment requirements (such as seasonal requirements and emergency response capabilities). Control requirement data is collected primarily through the following methods: analyzing environmental regulations and emission standards; surveying plant operational goals and cost control requirements; consulting operators to understand key control points and challenges in actual operation; and analyzing historical operating data to identify key control requirements. This requirement data is then prioritized and quantified to form a set of actionable control objectives.
[0086] Based on collected process flow data and control requirements data, a two-tier control architecture was designed. First, the modular structure of the process flow was analyzed, identifying relatively independent processing units and key interaction points between units. Then, the control complexity, response time requirements, and collaborative control needs of each processing unit were evaluated to determine the appropriate allocation of control functions. The two-tier control architecture design adopted a combined top-down and bottom-up approach: the top-down approach analyzed plant-level optimization goals and resource allocation requirements, defining the functions and responsibilities of the centralized optimization layer; the bottom-up approach analyzed unit-level control requirements and local optimization space, defining the functional scope and autonomous permissions of the distributed autonomous layer.
[0087] The centralized optimization layer in the control architecture is responsible for plant-wide decision-making, including high-level decisions such as load allocation, resource optimization, determination of key parameter ranges, and operating condition switching strategies. The distributed autonomous layer is deployed in each processing unit, responsible for local real-time control, achieving optimal unit-level operation within allocated resource budgets and parameter constraints. The two layers exchange information through a bidirectional communication mechanism: the centralized layer sends optimization strategies and constraints to the distributed layer; the distributed layer reports operating status and performance data to the centralized layer.
[0088] Ultimately, a complete hierarchical control architecture design is formed, including control level definitions, function allocation schemes, communication protocol specifications, and interface definitions. This two-layer control architecture integrates the global optimality of centralized decision-making with the real-time response capability of distributed control, making it suitable for intelligent optimization control of complex wastewater treatment systems.
[0089] Step S2: Based on the hierarchical control architecture, collect historical operating data and real-time status data of each processing unit, perform data preprocessing and feature extraction, and construct a global optimization model. The global optimization model reflects the interaction relationship between each processing unit.
[0090] Specifically, this step begins with establishing a comprehensive data collection plan, determining the types of data to be collected, the sampling frequency, and the data sources. Historical operational data collected includes operating records from the past 6-12 months, covering various operating conditions and seasonal variations. Data types include influent and effluent water quality parameters of the treatment unit (such as COD, BOD, ammonia nitrogen, total phosphorus, etc.), process control parameters (such as dissolved oxygen concentration, pH value, MLSS concentration, etc.), equipment operating parameters (such as blower frequency, pump flow rate, valve opening, etc.), and energy and chemical consumption data. Real-time status data is collected through an online sensor network to gather current operating information, including data from online water quality analyzers, flow meters, pressure sensors, and equipment status signals.
[0091] The collected raw data undergoes a rigorous preprocessing process to ensure data quality. Preprocessing steps include: data cleaning (detecting and handling missing values, outliers, and duplicates); data standardization (converting data of different dimensions to a uniform scale); time series alignment (ensuring consistency of timestamps from different data sources); and noise filtering (using techniques such as moving averages and median filtering to reduce data noise). For missing data of specific parameters, the system employs different processing strategies based on the missing data pattern: short-term missing data is filled using interpolation methods (such as linear interpolation and spline interpolation); medium-term missing data is replaced using similar day patterns; and long-term missing data is filled using machine learning models (such as random forests and neural networks) for prediction.
[0092] After data preprocessing, feature extraction is performed to transform the raw data into a more informative feature representation. Feature extraction methods include: time-domain features (such as mean, standard deviation, peak value, rate of change, etc.); frequency-domain features (extracting periodic patterns through Fourier transform); statistical features (such as distribution parameters, correlation coefficients, etc.); and derived features (such as process performance indicators such as load rate, removal efficiency, and energy efficiency). Specific features, such as the F / M ratio (food-to-microbe ratio) and SVI (sludge volume index), are also extracted to address the characteristics of wastewater treatment. Furthermore, dimensionality reduction techniques (such as principal component analysis and t-SNE) are used to reduce the dimensionality of the feature space, improving the efficiency of subsequent modeling.
[0093] Based on the preprocessed data and extracted features, the system constructs a global optimization model. This model employs a hybrid modeling approach, combining the advantages of physical mechanism models and data-driven models. The physical mechanism part, based on the fundamental principles of wastewater treatment, establishes mass balance equations, reaction kinetic equations, and energy balance equations to describe the intrinsic mechanisms of the main treatment processes. The data-driven part utilizes machine learning algorithms to capture complex nonlinear relationships and process characteristics that are difficult to model explicitly. The two models are integrated through a gray-box model framework: the physical model provides the structural skeleton and constraints; the data model supplements unknown parameters and complex relationships.
[0094] The global optimization model pays particular attention to the interactions between treatment units, such as load transfer between upstream and downstream units, the effects of recirculation and circulation, and competition for shared resources. The model structure adopts a modular design, with each treatment unit corresponding to a sub-model, and units connected through interface variables. Each sub-model employs a modeling method suitable for its specific characteristics: a variant of the ASM (Activated Sludge Model) is used for biochemical treatment units; kinetics and equilibrium equations are used for physicochemical units; and resistance models and fouling kinetics are used for membrane treatment units. Model parameters are determined through a multi-stage calibration process: first, initial parameter ranges are set using professional knowledge; then, coarse calibration is performed using historical data; and finally, fine calibration is performed using experimental data under specific operating conditions.
[0095] The constructed global optimization model can not only accurately predict the performance and mutual influence of each processing unit, but also support multi-objective optimization and scenario analysis, providing a scientific basis for subsequent optimization decisions.
[0096] Step S3: Based on the global optimization model, calculate the plant-wide resource optimization allocation scheme and control parameter constraints. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain the global optimization decision.
[0097] Specifically, this step begins with inflow prediction based on a global optimization model. Utilizing historical data and current operating conditions, combined with external information such as weather forecasts, industrial wastewater plans, and urban water usage patterns, the inflow situation for the next 24 to 72 hours is predicted. The prediction employs a multi-model fusion approach, including time series analysis models (such as ARIMA and SARIMA) to capture regular changes, machine learning models (such as random forests and LSTM neural networks) to handle complex nonlinear relationships, and rule-based models based on industrial knowledge to handle special cases. The prediction process first calculates the prediction results of each model separately, then dynamically adjusts the weights based on the performance of each model in recent predictions, generating a weighted fusion final prediction. Inflow prediction covers multiple indicators, including water volume (hourly average flow and peak flow), conventional water quality indicators (such as COD, BOD, ammonia nitrogen, total phosphorus, and SS), and concentrations of key pollutants. The prediction results simultaneously provide expected values and 95% confidence intervals to address uncertainties. A prediction accuracy evaluation mechanism is also established, comparing the deviation between predicted and actual values in real time. When the deviation exceeds a preset threshold, automatic model adjustment or manual intervention is triggered.
[0098] Next, based on the influent forecast data and effluent quality requirements, a resource budget allocation scheme is calculated using an optimization algorithm. The optimization objective is to minimize the total operating cost while meeting the effluent quality requirements. The total operating cost includes energy costs (electricity consumption of various equipment), chemical costs (such as coagulants, pH adjusters, disinfectants, etc.), and maintenance costs (costs related to equipment lifespan and maintenance cycles). The optimization process first establishes an objective function, converting various costs into unified economic indicators based on actual prices; then, constraints are set, including effluent quality compliance constraints (such as COD, ammonia nitrogen, total phosphorus, etc., below discharge standards), treatment capacity constraints (hydraulic load not exceeding design values), and process stability constraints (process parameter variations not exceeding safe ranges). The algorithm employs a phased optimization strategy: the first phase performs coarse optimization to determine the allocation ratio of major resources; the second phase performs fine optimization, further adjusting the specific parameters of each unit based on the coarse scheme. Considering the characteristics of multi-objective optimization, the algorithm adopts the Pareto optimality principle, generating a set of solutions balancing different objectives, and a decision-making mechanism selects the solution most suitable for the current operating conditions. Finally, a detailed resource budget allocation plan for each treatment unit was obtained, including energy consumption budget (such as the electricity quota for the aeration system and the operating schedule of the water pump) and chemical consumption budget (such as the dosage plan for various chemicals).
[0099] Then, based on the resource budget allocation scheme and process safety constraints, reasonable setting ranges for key control parameters are calculated. For each key control parameter, its physical limits are first determined (e.g., the maximum and minimum allowable values for the equipment), then the range is narrowed down according to process safety constraints (e.g., the safe range to prevent inhibition by the biochemical system), and finally, the optimization range is further refined based on the resource budget and optimization objectives. Taking dissolved oxygen concentration as an example, its physical limit may be 0-10 mg / L, process safety constraints may narrow the range to 0.5-6 mg / L, and the target range based on energy consumption optimization may be further narrowed to 1.2-2.5 mg / L. For the recovery rate of the reverse osmosis system, membrane fouling risk, energy consumption level, and permeate demand are considered to determine an optimization range such as 75%-85%. For the dosage of chemicals, a dosage range such as 18-25 mg / L is determined based on influent water quality prediction, coagulation test results, and empirical models. These parameter ranges are not isolated; the mutual influence and constraint relationships between parameters are also considered, such as the correspondence between dissolved oxygen and fan frequency, and the response curves of chemical dosage and coagulation effect.
[0100] Global optimization decision-making encompasses not only static parameter ranges but also dynamic adjustment strategies. Based on the temporal distribution characteristics of influent forecasts, time-segmented parameter adjustment plans are developed, such as increasing treatment intensity during peak influent load periods and reducing energy consumption during off-peak periods. Simultaneously, robust control strategies to address forecast uncertainties are designed, evaluating system response under different influent conditions through scenario analysis and developing a decision tree including alternative solutions. Furthermore, global optimization decision-making considers smooth transitions during equipment start-up and shutdown and operational mode switching to avoid system instability caused by sudden parameter changes. The final global optimization decision is output in structured data format, including resource quotas for each unit, control parameter ranges, time execution plans, and emergency response schemes, providing a decision-making framework and constraint boundaries for the distributed autonomous layer.
[0101] Step S4: Deploy edge controllers in each processing unit, collect local sensor data based on the global optimization decision, run optimization algorithms within the range of the control parameter constraints to obtain real-time control commands at the unit level, and control the operation of the devices in each processing unit.
[0102] Specifically, this step begins by deploying an edge control system, consisting of hardware and software, in each processing unit. Hardware deployment includes edge computing devices (industrial-grade computers or dedicated controllers), communication modules (supporting wired and wireless communication), power modules, and protective measures (such as waterproof, dustproof, and corrosion-resistant structures). Equipment selection considers the characteristics and control requirements of the processing unit; for example, high-performance controllers are chosen for units with high real-time requirements (such as aeration control), while ruggedized equipment is selected for areas with harsh environmental conditions (such as disinfection units). Software deployment includes a base operating system (typically a real-time operating system or a simplified Linux distribution), control algorithm libraries, a communication protocol stack, and security components. A modular software architecture is adopted to facilitate subsequent functional expansion and maintenance.
[0103] The edge controller establishes connections with field sensors and actuators to form a complete control loop. Sensor inputs include analog signal inputs (such as standard signals like 4-20mA and 0-10V), digital signal inputs (switching signals, counting pulses, etc.), and intelligent device inputs (supporting industrial bus protocols such as RS-485, Modbus, and PROFINET). Real-time calibration and validity verification of the input sensor data ensures accuracy and reliability. Actuator connections include inverter control interfaces, electric valve control interfaces, and metering pump control interfaces, ensuring that control commands are accurately transmitted to the terminal equipment.
[0104] Once deployed, the edge controller begins collecting local sensor data in real time. The data acquisition process follows a predetermined sampling strategy, with different parameters using different sampling frequencies: critical process parameters (such as dissolved oxygen and pH) are sampled at high frequencies (e.g., on a second or minute scale); slowly changing parameters (such as MLSS concentration) are sampled at low frequencies (e.g., on an hourly scale); and equipment status signals (such as pump operating status) are sampled based on event triggers. The collected data undergoes local processing, including noise filtering, anomaly detection, trend analysis, and status estimation, generating a higher-quality information stream.
[0105] Based on locally processed sensor data and global optimization decisions (including resource budgets and control parameter constraints) received from the centralized layer, the edge controller runs a local optimization algorithm to generate real-time control commands at the unit level. The local optimization algorithm is designed with computational resource constraints and real-time requirements in mind, employing efficient solution methods. The algorithm selects appropriate control strategies based on the characteristics of the processing unit: for units with well-defined dynamic characteristics (such as pH adjustment), classical control algorithms (such as PID control and feedforward control) are used; for units with strong nonlinearity and multivariate coupling (such as biochemical reactors), advanced control algorithms (such as model predictive control and adaptive control) are used; and for units with complex uncertainties (such as membrane filtration), intelligent control algorithms (such as fuzzy control and neural network control) are used. These control algorithms operate within the control parameter constraints determined by the global optimization decisions, ensuring both the flexibility and response speed of local control while ensuring that the control behavior conforms to the global optimization objective.
[0106] The local optimization process also considers various practical constraints, such as limits on the number of equipment start-ups and shutdowns, limits on parameter change rates, and safety interlock requirements. A preventative control strategy is implemented, predicting future state changes and taking control actions in advance to prevent process parameters from exceeding safe ranges. Furthermore, self-diagnostic and anomaly handling functions are implemented, capable of detecting control anomalies and initiating safety measures to ensure the safe and stable operation of the system.
[0107] After undergoing a rationality check, the control commands generated by the optimization algorithm are converted into specific equipment control signals and sent to each executing device through the execution interface. During the execution of the control commands, the execution effect is continuously monitored, and closed-loop control is implemented to ensure the achievement of the control objectives. For complex equipment (such as centrifuges and membrane systems), control includes not only simple start-stop and speed adjustment, but also operation sequence control and operating condition management, achieving more refined equipment control.
[0108] In this way, the edge controller achieves unit-level intelligent control at the distributed autonomous layer, which not only follows the guidance of global optimization decisions, but also makes real-time optimization adjustments based on local operating conditions, improving the system's response speed and adaptability, while maintaining the overall coordination of the plant's operation.
[0109] Step S5: Collect the running status information after execution and report it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer will make optimization decisions and adjustments for the next cycle to achieve collaborative control between the centralized optimization layer and the distributed autonomous layer.
[0110] Specifically, in this step, the edge controller first collects comprehensive operational status information after control execution. The collected data includes four main categories: process performance data, such as influent and effluent water quality indicators, treatment efficiency, and hydraulic load for each treatment unit; equipment operation data, such as equipment operating status, operating parameters, energy consumption data, and fault information; resource consumption data, such as energy usage, reagent consumption, and other operating cost data; and environmental data, such as water temperature, ambient temperature, and weather conditions. Data acquisition employs a multi-level caching mechanism to ensure no data loss in the event of communication interruptions, while data compression and priority management are implemented to improve data transmission efficiency.
[0111] The collected data undergoes local processing and analysis before being reported, generating more valuable information. Processing steps include: data validation, checking the completeness and validity of the data; data aggregation, aggregating high-frequency raw data into more meaningful statistics, such as averages, peak values, and fluctuation ranges; performance calculation, calculating key performance indicators based on the raw data, such as removal rate and energy efficiency ratio; and anomaly labeling, identifying and labeling abnormal operating conditions and events to facilitate centralized analysis. Data processing also includes privacy protection measures, de-identifying or encrypting sensitive data to ensure data security.
[0112] The processed data is reported to the centralized optimization layer via a two-way communication mechanism. The communication mechanism is designed with reliability, security, and efficiency in mind, employing a multi-layered communication protocol stack: the physical layer supports various communication media, such as wired Ethernet, wireless WiFi, and 4G / 5G mobile networks, and enables automatic switching; the transport layer uses the TCP / IP protocol to ensure reliable data transmission; and the application layer uses standard industry protocols (such as OPC UA, MQTT, etc.) or custom protocols to achieve structured data exchange. To ensure communication security, data encryption, authentication mechanisms, and access control are implemented to prevent unauthorized access and data leakage. The communication strategy combines event-driven and periodic approaches: critical data changes and abnormal events trigger immediate reporting; routine status information is reported in batches according to set periods. The communication mechanism not only supports uplink data transmission but also downlink transmission of parameters and commands, forming a true two-way communication closed loop.
[0113] The operational status information received by the centralized optimization layer first enters the data management system, where it is integrated, cleaned, and stored to form structured data assets. The data analysis system then performs multi-dimensional analysis on this data: performance evaluation, comparing actual operational results with expected targets; efficiency analysis, assessing resource utilization efficiency and processing costs; trend analysis, identifying changing trends and potential problems in key parameters; and correlation analysis, exploring relationships and potential causal chains between parameters. The analysis results are presented to operators and managers through a visual interface, supporting interactive decision-making.
[0114] Based on data analysis results and the latest operational requirements, the centralized optimization layer adjusts the optimization decisions for the next cycle. The adjustment process first evaluates the effectiveness of the current global optimization decisions, identifying areas for improvement, such as uneven resource allocation, overly strict or lenient parameter constraints, and substandard performance of certain units. Then, it updates key parameters of the global optimization model, such as reaction kinetic coefficients and equipment efficiency curves, to more accurately reflect the current operating conditions. Based on the updated model and the latest influent forecast, the plant-wide resource optimization allocation scheme and control parameter constraints are recalculated, generating the adjusted global optimization decisions.
[0115] The decision-making process integrates both data-driven and knowledge-driven approaches: on the one hand, it learns optimization patterns and association rules from historical data and applies machine learning techniques to predict decision outcomes; on the other hand, it applies domain expert knowledge and best practice guidelines to ensure that decisions comply with process specifications and safety requirements. For major adjustments, the system provides decision support functions, presenting multiple alternatives for human decision-makers to choose from, achieving human-machine collaborative decision-making.
[0116] The finalized optimization decision is distributed to each edge controller via the communication network, completing the guidance update from the centralized optimization layer to the distributed autonomous layer. Based on the updated global optimization decision, the edge controllers adjust their local control strategies and parameters, achieving a new round of unit-level optimization control. Through this periodic bidirectional information flow and decision update, close collaboration between the centralized optimization layer and the distributed autonomous layer is achieved, forming an adaptive closed-loop control system. This collaborative control mechanism combines the global optimality of central decision-making with the real-time response capability of distributed control, effectively addressing the complex changes and uncertainties in the wastewater treatment process and achieving stable and efficient wastewater treatment.
[0117] Example 2
[0118] like Figure 2 As shown, in this embodiment, the process flow data and control requirement data of the wastewater treatment plant are acquired, and a two-layer control architecture is designed based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture, including:
[0119] Step S11: Obtain the process flow data and control requirement data of the wastewater treatment plant. The process flow data includes the connection relationship of each treatment unit, the material flow path and energy consumption characteristics. The control requirement data includes the effluent water quality standards, treatment capacity requirements and safe operation constraints, thus obtaining the basic data of the wastewater treatment plant.
[0120] Step S12: Based on the basic data of the wastewater treatment plant, analyze the plant-wide resource coordination requirements and unit-level real-time response requirements to determine the system's functional requirements. The functional requirements include global resource collaborative allocation requirements, local rapid response control requirements, and abnormal situation handling requirements, resulting in a list of system functional requirements.
[0121] Step S13: Based on the system functional requirements list, design a centralized optimization layer responsible for setting plant-wide goals and allocating resource quotas, and design a distributed autonomous layer responsible for optimal autonomous control at the unit level, thus obtaining a two-layer control architecture and responsibility boundary definition;
[0122] Step S14: Based on the two-layer control architecture and responsibility boundary definition, plan the hardware configuration and installation location of the central server, edge computing devices and communication network to obtain the hierarchical control architecture.
[0123] Specifically, the first step is to acquire process flow data and control requirements data for the wastewater treatment plant. Process flow data acquisition is primarily accomplished through on-site surveys, process flow diagram analysis, and compilation of historical operation records. This includes the connectivity, material flow paths, and energy consumption characteristics of each treatment unit (such as pretreatment units, biological reaction tanks, sedimentation tanks, membrane filtration units, and disinfection units). Connectivity refers to the physical connection methods and sequence between different treatment units; material flow paths refer to the flow direction and quantity relationships of water, sludge, and gas between units; and energy consumption characteristics include the rated power, actual operating power, and energy consumption distribution characteristics of each piece of equipment. Control requirements data is obtained by analyzing the wastewater treatment plant's operational objectives and regulatory requirements. This mainly includes effluent quality standards (such as emission limits for COD, ammonia nitrogen, total phosphorus, and total nitrogen), treatment capacity requirements (design treatment capacity, peak treatment capacity, etc.), and safety constraints (such as equipment operating temperature range and pressure limits). These data collectively constitute the basic dataset for the wastewater treatment plant, providing data support for subsequent architecture design.
[0124] Next, based on the acquired basic data of the wastewater treatment plant, the system systematically analyzes the plant-wide resource coordination requirements and unit-level real-time response requirements to determine the system's functional requirements. Plant-wide resource coordination requirements refer to the resource allocation issues that need to be considered holistically from the perspective of the entire wastewater treatment plant, including energy allocation (electricity quotas for each unit), chemical allocation (the usage and timing of various chemical agents), and treatment load allocation (the allocation of treated water volume among parallel units). Unit-level real-time response requirements focus on the ability of individual treatment units to quickly adjust control parameters when faced with fluctuations in influent water quality and quantity, changes in equipment status, such as real-time adjustment of dissolved oxygen concentration in the biological treatment tank and dynamic optimization of membrane system flux. The system functional requirements comprehensively consider the global resource coordination and allocation requirements, local rapid response control requirements, and abnormal situation handling requirements, forming a system functional requirement list that provides clear goals and constraints for the architecture design.
[0125] Then, based on the system functional requirements list, a two-layer control architecture was designed. The centralized optimization layer is responsible for plant-wide target setting and resource allocation. It runs on a central server, possessing strong computing power and a global perspective, but its update frequency is low (typically hourly or shift-level). The centralized optimization layer calculates the optimal resource allocation scheme by analyzing historical data and real-time status across the entire plant, and sets constraint ranges for control parameters for each processing unit, but does not directly participate in the real-time control of the units. The distributed autonomous layer is responsible for unit-level optimal autonomous control. Deployed on the edge controllers of each processing unit, it has rapid response capabilities (typically on the order of seconds or minutes) and can autonomously make decisions and execute optimal control strategies based on local sensor data, within the constraints set by the centralized layer. By clearly defining the functional boundaries and interaction mechanisms of the two layers, the two-layer control architecture and responsibility boundary definitions were obtained, achieving an organic combination of global optimization and local autonomy.
[0126] Finally, based on the two-layer control architecture and defined responsibility boundaries, a specific hardware deployment scheme was planned. The central server, an industrial-grade server with sufficient computing power and storage capacity, was selected and installed in the control room of the wastewater treatment plant to handle the computational tasks of the centralized optimization layer. Edge computing devices, such as industrial PCs or high-performance PLCs, suitable for the industrial environment, were selected and installed in the field control cabinets of each key processing unit, responsible for the real-time control of the distributed autonomous layer. The communication network adopted a layered design, with the plant backbone network using industrial Ethernet for high-speed data transmission, and industrial buses (such as PROFIBUS, Modbus, etc.) used at the field device level for reliable data acquisition and control command issuance. Through detailed planning of hardware configuration and installation locations, a complete hierarchical control architecture was formed, laying the foundation for the actual deployment and operation of the system.
[0127] The two-layer control architecture refers to a hierarchical control system composed of a centralized optimization layer and a distributed autonomous layer. The centralized optimization layer is like the "brain" of the wastewater treatment plant, responsible for global decision-making and resource scheduling; the distributed autonomous layer is like the "nerve nodes" of various organs, responsible for local perception and rapid response. This architecture draws inspiration from the working mode of the human nervous system. Through clear hierarchical division and division of responsibilities, it solves the problems of slow response and poor adaptability of traditional centralized control systems, as well as the lack of global coordination in fully decentralized control systems. This allows the system to maintain global optimality while achieving rapid response, adapting to the complexity and dynamism of wastewater treatment processes.
[0128] A hierarchical control architecture is the physical implementation of a two-tier control architecture. It includes not only logical functional layering but also physical hardware configuration, network topology, and communication protocols. The hierarchical control architecture adopts a "pyramid" structure: the top layer is the central server and data center; the middle layer consists of area controllers and communication gateways; and the bottom layer comprises field sensors and actuators. Devices at different levels have different functional roles and performance requirements, achieving seamless integration and collaborative operation through standardized interfaces and communication protocols. This hierarchical design gives the system excellent scalability and flexibility, enabling it to adapt to wastewater treatment plants of different sizes and process characteristics, while also facilitating phased implementation and upgrades.
[0129] Example 3
[0130] In this embodiment, the collection of historical operating data and real-time status data of each processing unit, and the subsequent data preprocessing and feature extraction, include:
[0131] Historical operating data and real-time status data are collected from the sensors and control systems of each treatment unit in the wastewater treatment plant. The historical operating data and real-time status data include water quality parameters, equipment operating parameters and energy consumption data to obtain the raw dataset.
[0132] The original dataset is subjected to missing value imputation, outlier detection, and data format standardization to obtain a cleaned dataset.
[0133] The cleaned dataset is then subjected to sensitivity classification identification, and divided into high-sensitivity data, medium-sensitivity data, and low-sensitivity data to obtain the classified dataset.
[0134] Based on the graded dataset, time features, frequency domain features, and statistical features are extracted to obtain preprocessed plant-wide data.
[0135] Specifically, firstly, data is collected from sensors and control systems in each treatment unit of the wastewater treatment plant. Historical operational data is typically stored in the wastewater treatment plant's SCADA system, database, or historical record files, covering a relatively long period (e.g., 3 months to 1 year) to reflect the system's long-term operational characteristics and seasonal variation patterns. Real-time status data is collected in real time through online sensors and control systems to reflect current operating conditions and short-term trends. The collected data mainly includes three categories: water quality parameters (e.g., pH, dissolved oxygen, COD, ammonia nitrogen, SS), equipment operating parameters (e.g., fan speed, pump flow rate, valve opening, membrane pressure difference), and energy consumption data (e.g., electricity consumption, chemical consumption). This data is collected through industrial communication networks (e.g., Modbus, Profibus) or database interfaces and stored at preset time intervals (e.g., minutes, hours) to form a raw dataset containing timestamps, measurement point IDs, numerical values, and quality labels.
[0136] Next, the collected raw dataset underwent data cleaning and preprocessing. First, missing value imputation was performed. For data missing due to sensor malfunctions, communication interruptions, etc., different imputation methods were used based on the data characteristics: short-term missing data (less than a preset threshold, such as 30 minutes) used linear interpolation or spline interpolation; medium-term missing data (a few hours) used historical data under similar operating conditions; long-term missing data (more than one day) was marked as invalid and excluded from subsequent analysis. Then, outlier detection was performed. Outliers in the data were identified using statistical methods (such as the 3σ principle and box plots) or machine learning methods (such as isolated forests and local anomaly algorithms), and obviously abnormal data points were corrected or removed. Finally, data format standardization was performed, unifying data from different systems and sensors into a standard format, including unit conversion, time alignment, and data structure normalization, for subsequent analysis and processing. After these processes, a cleaned dataset was obtained, improving data quality and usability.
[0137] Then, the cleaned dataset was subjected to sensitivity classification. Based on the sensitivity of the data and protection requirements, the data was divided into three sensitivity levels using the following specific criteria:
[0138] Highly sensitive data includes data that meets any of the following conditions: (1) special process parameters protected by patents or constituting trade secrets, such as key control parameters of proprietary biochemical reactors and parameters of independently developed membrane fouling control algorithms; (2) precise formulations and dosing strategies of special agents, such as independently developed flocculant combination formulations and proprietary defoamer formulations; (3) key efficiency indicators and optimization algorithms that can significantly affect the competitive advantage of treatment plants, such as core algorithm parameters for energy consumption optimization and precise control strategies for agent dosage; (4) shared technical data with third parties under confidentiality agreements; (5) system configuration information that may involve the protection of critical infrastructure, such as key control points and safety thresholds of large municipal wastewater treatment plants. The leakage of highly sensitive data may lead to the loss of core technological advantages, significant economic losses, or violations of laws and regulations.
[0139] Medium-sensitivity data includes data that meets any of the following conditions: (1) energy consumption distribution data of the processing unit, including the power consumption and energy efficiency of each device; (2) processing efficiency indicators, such as COD removal rate and nitrogen and phosphorus removal efficiency; (3) operating cost composition data, such as reagent costs, energy costs, and maintenance costs; (4) statistical characteristics of process parameters, such as the fluctuation range and average value of key parameters; (5) equipment maintenance records and performance degradation data; (6) non-core but still somewhat proprietary control strategy parameters; (7) operational data such as supplier and procurement information. The leakage of medium-sensitivity data may affect the normal operation of the enterprise or cause certain competitive disadvantages, but will not lead to the leakage of core technologies.
[0140] Low-sensitivity data includes data that meets any of the following conditions: (1) monitoring results of routine water quality indicators, such as publicly reported indicators of COD, BOD, ammonia nitrogen, total phosphorus, and total nitrogen in influent and effluent; (2) basic operating status of equipment, such as start-up and shutdown status and operating time; (3) environmental parameters, such as water temperature, ambient temperature, and atmospheric pressure; (4) publicly published or industry-standard process parameter ranges; (5) compliant data that meets emission standards; (6) operation records without specific sensitivity; and (7) environmental monitoring data that is required to be disclosed by regulations. Low-sensitivity data usually does not involve trade secrets, has a low risk of leakage, and some data may even need to be disclosed to the public or regulatory authorities in accordance with regulations.
[0141] The implementation process of data sensitivity grading includes: first, an assessment team composed of data security experts and process experts conducts an initial classification of data items according to the aforementioned standards; then, risk quantification is performed using a risk assessment matrix (combining the likelihood and severity of data breaches); finally, a sensitive data grading catalog is created, clearly defining the protection requirements, access permissions, and usage guidelines for each type of data. This grading method ensures that data protection measures match actual risks, avoiding the problems of insufficient or excessive protection. After grading is completed, a graded dataset containing sensitivity labels is obtained, providing a foundation for subsequent data processing and protection.
[0142] Finally, feature extraction and data analysis are performed based on the graded dataset. Temporal feature extraction focuses on the time-series characteristics of the data, including trend analysis, seasonality identification, periodic pattern extraction, and key time point detection, using methods such as sliding window, autocorrelation analysis, and time series decomposition. Frequency domain feature extraction transforms the time-domain data to the frequency domain using Fourier transform or wavelet transform, analyzing the spectral characteristics of the data and identifying the contributions of different frequency components and the characteristics of periodic fluctuations. Statistical features include calculations of various statistics (such as mean, variance, skewness, kurtosis, etc.) and distribution characteristic analysis, reflecting the central tendency and dispersion of the data. These extracted features are combined to form the preprocessed plant-wide data, providing rich and effective input information for subsequent model building and optimization decisions.
[0143] Sensitivity classification refers to the process of categorizing data into different security levels based on their sensitivity and protection requirements. Sensitivity assessment considers several factors: potential business losses from data breaches (such as leaks of proprietary technology, loss of competitive advantage, etc.), legal compliance risks (such as violations of data protection regulations and contractual obligations), operational security risks (such as system failures caused by malicious tampering of critical parameters), and privacy protection needs (such as personal or corporate privacy information). Sensitivity classification is not merely a one-time data classification activity, but an ongoing process requiring regular review and updates to adapt to technological advancements, regulatory changes, and evolving business needs. In the actual operation of wastewater treatment systems, sensitivity classification results directly impact data storage location (local or cloud), encryption strength, access control policies, transmission protection measures, and the scope of data sharing, forming the foundation of the data security management system. Reasonable sensitivity classification can improve system operational efficiency and data utilization value while ensuring data security, providing strong support for the intelligent and digital transformation of wastewater treatment plants.
[0144] Example 4
[0145] In this embodiment, the construction of a global optimization model, which reflects the interaction relationships between processing units, includes:
[0146] Based on the preprocessed plant-wide data, a process network diagram is constructed to represent the material flow, energy flow, and information flow relationships between each processing unit. The process network diagram includes nodes and edges, where nodes represent processing units and edges represent the connection relationships between units, thus obtaining a wastewater treatment process network diagram.
[0147] Based on the wastewater treatment process network diagram, a simplified activated sludge model and a membrane filtration model are used to establish a local dynamic model for each treatment unit, resulting in a simplified model for each unit.
[0148] Based on the simplified models of each unit and the coupling relationships between units, model parameters are identified through data-driven machine learning methods, including neural networks or support vector machines, to obtain the global optimization model.
[0149] Specifically, firstly, a wastewater treatment process network diagram is constructed based on pre-processed plant-wide data. A process network diagram is a directed graph structure used to represent the material flow, energy flow, and information flow relationships between various treatment units within a wastewater treatment plant. Nodes represent treatment units (such as screens, grit chambers, biological treatment tanks, secondary sedimentation tanks, membrane units, disinfection units, etc.), and each node contains attribute information such as unit type, treatment capacity, key parameters, and control objectives. Edges represent the connections between units, including material flow connections (such as water flow, sludge flow, gas flow, etc.), energy flow connections (such as energy supply, heat transfer, etc.), and information flow connections (such as control signals, status feedback, etc.). The weight information attached to the edges indicates the connection strength or degree of influence, determined by process flow analysis and data correlation analysis. The process network diagram construction process first determines the basic structure based on the process flow diagram, then identifies hidden connections and degree of influence using data-driven methods (such as Granger causality tests, mutual information analysis, etc.), and finally is reviewed and confirmed by domain experts to form a complete wastewater treatment process network diagram, providing a structural framework for subsequent model construction.
[0150] Next, based on the wastewater treatment process network diagram, simplified process models were used to establish local kinetic models for each treatment unit. For the biological treatment unit, a simplified activated sludge model (SASM) was adopted. This model retains the core mechanisms of traditional ASM models (such as ASM1, ASM2d, etc.) but reduces state variables and reaction processes, focusing primarily on the removal kinetics of key indicators such as COD, ammonia nitrogen, and total nitrogen. Simplification methods included merging similar processes, ignoring minor reactions, and linearizing nonlinear terms, significantly reducing computational complexity while maintaining the model's main predictive capabilities. For the membrane filtration unit, a simplified membrane filtration model (SMFM) was used. This model mainly describes the relationship between membrane flux, transmembrane pressure difference (TMP), and membrane fouling. By combining a classical membrane resistance model and data-driven correction terms, it accurately predicts membrane system performance while reducing computational burden. Other treatment units (such as sedimentation and disinfection) also adopted corresponding simplified models. These local kinetic models together constitute the simplified model set for each unit, describing the internal kinetic behavior of individual treatment units.
[0151] Then, based on the simplified models of each unit and the coupling relationships between units, data-driven machine learning methods are used to identify model parameters. For parameters that are difficult to determine directly in the model (such as reaction rate constant, half-saturation coefficient, yield coefficient, etc.), actual operating data is used for parameter identification. Parameter identification adopts a two-stage strategy: first, independent parameter identification is performed on each unit model, using the unit's input and output data and optimization algorithms (such as genetic algorithms, particle swarm optimization, etc.) to find the parameter set that minimizes the model's prediction error; then, global parameter fine-tuning is performed, considering the coupling effects between units, using plant-wide input and output data to fine-tune key coupling parameters, thereby improving the overall model accuracy. During the parameter identification process, machine learning methods such as neural networks or support vector machines are used to assist in modeling, especially for complex processes that are difficult to accurately describe with mechanistic models (such as membrane fouling dynamics, microbial community changes, etc.). Machine learning models can capture hidden patterns in the data and improve prediction accuracy. By combining mechanistic models and data-driven methods, a hybrid model with physical interpretability and high prediction accuracy is constructed, which is the final global optimization model.
[0152] The global optimization model is a comprehensive mathematical model that reflects the interactions between various units in wastewater treatment. It can simulate and predict the dynamic response of the system under different control strategies, providing support for optimization decisions. The model adopts a modular design, with each treatment unit corresponding to a sub-model. These sub-models are coupled through boundary conditions to form the overall model. The model considers both the internal dynamic processes of the treatment units (such as biochemical reactions, solid-liquid separation, and membrane filtration) and the inter-unit influences (such as the impact of upstream water quality changes on downstream treatment effects and the impact of returned sludge on the biological system). The model's time scale covers short-term dynamic responses (minute-level) to long-term operating trends (day-level or month-level), supporting optimization decisions at different time scales. The mathematical form of the model includes ordinary differential equations (describing continuous process dynamics), discrete event equations (describing intermittent operations), and algebraic constraint equations (describing physicochemical equilibrium relationships), and the system state is calculated in real time using numerical solution methods.
[0153] A process network diagram (CBD) is a graphical model used to represent the structure and connections of a wastewater treatment system. Unlike traditional process flow diagrams (PFDs) or pipe and instrumentation diagrams (P&IDs), CBDs emphasize the functional relationships between units rather than physical connections, and can quantify the degree of influence and dependence. In a CBD, connection strength is typically represented by edge weights, which can be derived based on physical flow, energy transfer, or data correlation analysis. For example, the greater the material flow between two units, the higher their connection weight; the greater the impact of a parameter change in one unit on another, the higher its information flow connection weight. CBDs can identify key nodes (such as treatment bottlenecks and control hubs) and important paths in the system through graph theory analysis, providing a structured perspective for system optimization. Furthermore, CBDs serve as the foundation for subsequent tree decomposition; a reasonable graph structure can reduce the computational complexity of the system and improve the model's adaptability.
[0154] Simplified activated sludge (SASM) models are simplified mathematical models of wastewater biological treatment based on classic activated sludge models (such as ASM1, ASM2d, and ASM3). Traditional ASM models contain a large number of state variables (typically 10-20) and process equations (typically 15-25), resulting in high computational complexity and unsuitability for online optimization applications. SASM simplifies these models through the following methods: 1) merging similar components, such as combining different forms of organic matter into a few categories; 2) simplifying biological processes, retaining only the main degradation, transformation, and growth processes; 3) adopting a quasi-steady-state assumption, simplifying some rapid equilibrium processes into algebraic equations; and 4) introducing empirical correction terms, replacing complex mechanisms with simpler functions. Through these simplifications, SASM models typically retain only 3-5 key state variables and 5-8 main processes, significantly reducing the computational burden while maintaining the predictive power for the removal of major pollutants (COD, ammonia nitrogen, total nitrogen, etc.), making real-time optimization possible. Typical inputs to the SASM model include control parameters such as influent water quality indicators, aeration rate, and return ratio; outputs include effluent water quality, sludge production, and energy consumption indicators.
[0155] Example 5
[0156] In this embodiment, the calculation of the plant-wide resource optimization allocation scheme and control parameter constraints, wherein the plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain a global optimization decision, includes:
[0157] Based on the global optimization model, the influent water quality and influent water volume for future periods are predicted to obtain influent prediction data.
[0158] Based on the influent prediction data and effluent water quality requirements, the energy consumption budget and chemical consumption budget of each treatment unit are calculated by an optimization algorithm. The optimization algorithm aims to minimize the total cost and obtains a resource budget allocation scheme for each unit.
[0159] Based on the resource budget allocation scheme and process safety constraints of each unit, the reasonable setting range of each key control parameter is calculated. The key control parameters include dissolved oxygen concentration, reverse osmosis system recovery rate and reagent dosage, and the global optimization decision is obtained.
[0160] Specifically, firstly, inflow prediction is performed based on a global optimization model. Utilizing historical data and current operating conditions, combined with external information such as weather forecasts, industrial wastewater plans, and urban water usage patterns, the inflow situation for the next 24 to 72 hours is predicted. The prediction employs a multi-model fusion approach, including time series analysis models (such as ARIMA and SARIMA) to capture regular changes, machine learning models (such as random forests and LSTM neural networks) to handle complex nonlinear relationships, and rule-based models based on industrial knowledge to handle special cases. The prediction process first calculates the prediction results of each model separately, then dynamically adjusts the weights based on the performance of each model in recent predictions, generating a weighted fusion final prediction. Inflow prediction covers multiple indicators, including water volume (hourly average flow and peak flow), conventional water quality indicators (such as COD, BOD, ammonia nitrogen, total phosphorus, SS, etc.), and concentrations of key pollutants. The prediction results simultaneously provide expected values and 95% confidence intervals to address uncertainties. The system also establishes a prediction accuracy evaluation mechanism, comparing the deviation between predicted and actual values in real time. When the deviation exceeds a preset threshold, automatic model adjustment or manual intervention is triggered. This step allows the system to obtain relatively accurate water inflow prediction data for a future period, providing a foundation for subsequent optimization.
[0161] Next, based on the influent forecast data and effluent quality requirements, a resource budget allocation scheme is calculated using an optimization algorithm. The optimization objective is to minimize the total operating cost while meeting the effluent quality requirements. The total operating cost includes energy costs (electricity consumption of various equipment), chemical costs (such as coagulants, pH adjusters, disinfectants, etc.), and maintenance costs (costs related to equipment lifespan and maintenance cycles). The optimization process first establishes an objective function, converting various costs into unified economic indicators based on actual prices; then, constraints are set, including effluent quality compliance constraints (such as COD, ammonia nitrogen, total phosphorus, etc., below discharge standards), treatment capacity constraints (hydraulic load not exceeding design values), and process stability constraints (process parameter variations not exceeding safe ranges). The algorithm employs a phased optimization strategy: the first phase performs coarse optimization to determine the allocation ratio of major resources; the second phase performs fine optimization, further adjusting the specific parameters of each unit based on the coarse scheme. Considering the characteristics of multi-objective optimization, the algorithm adopts the Pareto optimality principle, generating a set of solutions balancing different objectives, and a decision-making mechanism selects the solution most suitable for the current operating conditions. Finally, a detailed resource budget allocation plan for each treatment unit was obtained, including energy consumption budget (such as the electricity quota for the aeration system and the operating schedule of the water pump) and chemical consumption budget (such as the dosage plan for various chemicals).
[0162] Then, based on the resource budget allocation scheme and process safety constraints, reasonable setting ranges for key control parameters are calculated. For each key control parameter, its physical limits are first determined (e.g., the maximum and minimum allowable values for the equipment), then the range is narrowed down according to process safety constraints (e.g., the safe range to prevent inhibition by the biochemical system), and finally, the optimization range is further refined based on the resource budget and optimization objectives. Taking dissolved oxygen concentration as an example, its physical limit may be 0-10 mg / L, process safety constraints may narrow the range to 0.5-6 mg / L, and the target range based on energy consumption optimization may be further narrowed to 1.2-2.5 mg / L. For the recovery rate of the reverse osmosis system, membrane fouling risk, energy consumption level, and permeate demand are considered to determine an optimization range such as 75%-85%. For the dosage of chemicals, a dosage range such as 18-25 mg / L is determined based on influent water quality prediction, coagulation test results, and empirical models. These parameter ranges are not isolated; the mutual influence and constraint relationships between parameters are also considered, such as the correspondence between dissolved oxygen and fan frequency, and the response curves of chemical dosage and coagulation effect. Ultimately, the system integrates various constraints and optimization objectives to determine a reasonable set range for each key control parameter, thus forming a global optimization decision.
[0163] Global optimization decision-making encompasses not only static parameter ranges but also dynamic adjustment strategies. Based on the temporal distribution characteristics of influent forecasts, time-segmented parameter adjustment plans are developed, such as increasing treatment intensity during peak influent load periods and reducing energy consumption during off-peak periods. Simultaneously, robust control strategies to address forecast uncertainties are designed, evaluating system response under different influent conditions through scenario analysis and developing a decision tree including alternative solutions. Furthermore, global optimization decision-making considers smooth transitions during equipment start-up and shutdown and operational mode switching to avoid system instability caused by sudden parameter changes. The final global optimization decision is output in structured data format, including resource quotas for each unit, control parameter ranges, time execution plans, and emergency response schemes, providing a decision-making framework and constraint boundaries for the distributed autonomous layer.
[0164] The resource budget allocation scheme refers to the plan for allocating limited resources to each treatment unit of a wastewater treatment plant while meeting treatment requirements. The resource budget mainly includes two categories: energy budget and chemical budget. The energy budget details the electricity quotas for each energy-consuming equipment (such as blowers, pumps, and agitators), down to the hourly electricity consumption limit or equipment operating schedule. The chemical budget specifies the usage quotas for various chemical agents (such as coagulants, flocculants, disinfectants, and alkaline agents). Resource budget allocation follows the principles of "total control, differentiated allocation, and dynamic adjustment": total control ensures that the total resource usage does not exceed the economically reasonable level; differentiated allocation assigns different resource quotas based on the treatment load and efficiency of each unit; and dynamic adjustment allows for flexible adjustment of quotas between units according to actual needs within the total control framework. The resource budget allocation scheme is usually presented in tabular or time-series format, visually displaying the resource quotas of each unit and their temporal distribution, facilitating implementation and monitoring.
[0165] Control parameter constraints refer to a set of conditions that limit the range of variation of key control parameters. They consider both physical and technological feasibility constraints, as well as economic and optimization objective guiding constraints. Control parameter constraints are typically expressed as upper and lower limits, such as the control range of dissolved oxygen concentration, the allowable range of reverse osmosis system recovery rate, and the adjustment range of pH value. These constraints are set based on three levels of consideration: first, a safety layer, ensuring that parameters do not exceed the limits that would lead to system failure or process collapse; second, a stability layer, ensuring that the process operates in a stable state and avoiding oscillations or loss of control; and finally, an optimization layer, guiding parameters towards optimal operating conditions. Control parameter constraints are not static but dynamically updated as operating conditions change, optimization objectives are adjusted, and equipment status changes. In actual implementation, these constraints are passed to the distributed autonomous layer as the decision boundary of the local control algorithm, ensuring that local optimization does not deviate from the global objective.
[0166] Example 6
[0167] In this embodiment, obtaining the global optimization model includes:
[0168] Based on the wastewater treatment process network diagram, tree width analysis is performed on the wastewater treatment process network diagram to identify the key coupling points and bottlenecks of the system and obtain the tree width parameters of the process network.
[0169] Based on the tree width parameter of the process network, the global optimization model is decomposed into several sub-models. Each sub-model corresponds to a cluster of tree decomposition. Tree decomposition ensures that the connection between sub-models follows a tree structure, resulting in a set of sub-models after tree decomposition.
[0170] Based on the set of sub-models after tree decomposition, a hierarchical optimization model is constructed. The top-level node of the hierarchical optimization model corresponds to the macro-decision variables of the entire plant, the intermediate-level nodes correspond to the collaborative optimization variables of the processing unit group, and the leaf nodes correspond to the micro-control variables of a single processing unit. This results in a hierarchical optimization model based on a tree structure, which serves as the global optimization model.
[0171] Specifically, firstly, tree width analysis is performed based on the wastewater treatment process network diagram. Tree width analysis is a graph theory method used to assess the complexity and computational difficulty of a graph structure. The process of performing tree width analysis on the wastewater treatment process network diagram begins by treating the process network diagram as an undirected graph, where nodes represent treatment units and edges represent material, energy, or information flow connections between units. Then, tree decomposition algorithms (such as minimum filling algorithms, greedy elimination algorithms, etc.) are used to calculate the optimal or near-optimal tree decomposition of the graph, and the corresponding tree width values are obtained. Tree width is an important indicator of graph complexity; a smaller tree width means a relatively simple graph structure, allowing for more efficient model solving. During the tree width analysis, key coupling points and system bottlenecks are also identified. Key coupling points are nodes connecting multiple treatment units, whose state changes significantly affect multiple downstream units; system bottlenecks are nodes with limited treatment capacity or high control difficulty, often becoming key factors affecting overall performance. For example, the secondary sedimentation tank may be a key coupling point connecting biological treatment and subsequent treatments, while the biological reactor may become a system bottleneck under high load conditions. The importance and influence of each node are quantitatively assessed by calculating indicators such as the degree (number of edges connected to it), betweenness centrality (number of shortest paths through the node), and flow centrality (material flow through the node). Finally, the system obtains the tree width parameters of the process network, including the tree width value, tree decomposition structure, and a list of key nodes, laying the foundation for subsequent model decomposition.
[0172] Next, based on the tree width parameter of the process network, the global optimization model is decomposed into several sub-models. Tree decomposition is a method that maps complex graph structures to tree structures. It groups the nodes in the original graph into a series of overlapping clusters (called "bags"), and the connections between these clusters form a tree. Nodes within each cluster may have complex internal connections, but the connections between clusters follow the properties of a tree, i.e., any two clusters containing the same node form a connected path in the tree. Based on the tree decomposition structure, the global optimization model is divided into several sub-models, each corresponding to a cluster in the tree decomposition. The sub-model contains the state variables, control variables, and constraints of all nodes (processing units) within the cluster, as well as the relationship equations between these variables. Sub-models are coupled through shared variables (i.e., node variables that appear in multiple clusters). For example, if the secondary sedimentation tank belongs to both the biological treatment cluster and the deep treatment cluster, then the state variables of the secondary sedimentation tank (such as effluent SS, reflux ratio, etc.) become shared variables connecting the two sub-models. During the model decomposition process, special attention is paid to maintaining physical and technological consistency to ensure that the decomposed set of sub-models can equivalently represent the original global model. For each sub-model, further simplification and normalization are performed, removing redundant variables and constraints, standardizing the equation form, and optimizing solution efficiency. Ultimately, the resulting set of tree-decomposed sub-models provides the foundational components for constructing hierarchical optimization models.
[0173] Then, based on the set of sub-models after tree decomposition, a hierarchical optimization model is constructed. The hierarchical optimization model is a hierarchical decision structure that decomposes complex optimization problems into sub-problems at different levels and achieves overall optimization through inter-layer coordination. In this system, the hierarchical optimization model contains three levels: top layer, intermediate layer, and leaf layer. Top-level nodes correspond to macro-level decision variables for the entire plant, such as the total energy consumption allocation ratio, the treatment load allocation of each treatment stage, and the total control of key resources (such as chemicals and energy). These variables affect overall performance but do not directly control specific equipment. Intermediate-level nodes correspond to collaborative optimization variables for treatment unit groups, such as the carbon-to-nitrogen ratio control within biochemical treatment unit groups and the flux allocation within membrane treatment unit groups. These variables need to consider the synergistic effects and balance relationships within the unit groups. Leaf nodes correspond to micro-level control variables for individual treatment units, such as the dissolved oxygen setpoint of a single aeration tank and the dosage control of a single dosing pump. These variables directly affect the equipment operating status and treatment effect. The inter-layer relationships follow the characteristics of a tree structure. The decision results of upper-level nodes become constraints for lower-level nodes, while the execution feedback of lower-level nodes becomes the optimization basis for upper-level nodes. During the construction process, the system determines the composition and connection relationships of nodes at each level based on the tree decomposition structure, designs inter-layer communication and coordination mechanisms, and optimizes the solution algorithms and execution cycles of each level. Ultimately, the system obtains a hierarchical optimization model based on a tree structure, which serves as the implementation form of the global optimization model.
[0174] The hierarchical optimization model employs an alternating direction method and an information transmission mechanism. The solution process begins at the top level, calculating the optimal values of macroscopic decision variables based on global objectives and constraints, and transmitting the results as constraints to intermediate levels. Intermediate levels receive constraints from upper levels, combine them with their own optimization objectives, calculate the optimal values of collaborative optimization variables, and transmit the results to leaf nodes. Leaf nodes, within their received constraints, optimize microscopic control variables and feed back the results and status information to upper-level nodes. This top-down decision transmission and bottom-up information feedback form a closed-loop optimization mechanism, ensuring global optimality while considering local response speed. A multi-timescale coordination mechanism is also designed, allowing optimization at different levels to run at different cycles: the top-level optimization cycle is longer (e.g., several hours), emphasizing long-term stability; the intermediate-level optimization cycle is moderate (e.g., tens of minutes), balancing stability and responsiveness; and the leaf node optimization cycle is short (e.g., several minutes or seconds), achieving rapid response. Through this hierarchical and time-sharing coordination mechanism, the system can efficiently handle the multi-scale dynamic characteristics of complex wastewater treatment processes.
[0175] Tree width is an important concept in graph theory, used to measure the complexity of a graph structure. Specifically, in wastewater treatment process network graphs, tree width reflects the complexity of the most complex local structures within the system. The calculation of tree width is based on the tree decomposition of the graph: first, the graph is decomposed into a series of connected clusters of nodes (called "bags"), and the connections between these clusters form a tree; then, the size of each cluster (the number of nodes it contains) is calculated; the tree width is equal to the largest cluster size minus one. A smaller tree width indicates a graph structure closer to a tree, resulting in lower computational complexity; a larger tree width indicates highly coupled local structures within the graph, leading to higher computational complexity. In wastewater treatment systems, tree width analysis can help identify complex coupling regions and guide the design of model decomposition and solution strategies. For example, if a wastewater treatment plant has a small tree width (e.g., 2-3), a simpler decomposition method can be used; if the tree width is large (e.g., 5-6 or higher), a more complex decomposition and coordination strategy is required. Tree width analysis is an important tool for system structure simplification and computational optimization, significantly improving the solution efficiency of complex systems.
[0176] Tree decomposition is a technique for mapping a complex graph structure to a tree structure. In wastewater treatment systems, tree decomposition transforms tightly coupled networks of treatment units into hierarchical tree structures, facilitating distributed optimization and parallel computing. The core of tree decomposition is organizing the nodes in the graph into a series of overlapping clusters that satisfy three conditions: 1) Coverage condition: each node in the original graph belongs to at least one cluster; 2) Edge coverage condition: both endpoints of each edge in the original graph appear in at least one cluster simultaneously; 3) Connectivity condition: all clusters containing any given node form a connected subtree in the tree. This decomposition method transforms the originally complex graph structure into a more manageable tree structure while preserving key structural information and node relationships. In the optimization of wastewater treatment systems, a good tree decomposition can group highly coupled units (such as biochemical reactors and secondary sedimentation tanks) into the same cluster, while placing weakly coupled units (such as pretreatment and disinfection) into different clusters, thereby reducing computational complexity while ensuring optimization quality. The quality of tree decomposition directly affects the efficiency and accuracy of subsequent hierarchical optimization and is a crucial step in constructing the global optimization model.
[0177] Hierarchical optimization models are structured solution methods that decompose complex optimization problems into multi-level sub-problems. In wastewater treatment systems, hierarchical optimization models divide the system into different levels according to the scope and time scale of decisions: the top level handles plant-wide decisions, such as total resource allocation and treatment load allocation, which have a long optimization cycle; the middle level handles unit-level collaborative control, such as the internal balance of biochemical units and load allocation of multiple parallel units, which have a moderate optimization cycle; the leaf level handles specific control parameters at the unit level, such as the setting of operating parameters for individual equipment, which have a short optimization cycle. The layers coordinate with each other through information transmission and constraint relationships: the decision results of the upper level form the constraints of the lower level, and the execution feedback of the lower level provides the upper level with state information and optimization basis. This hierarchical structure enables the system to simultaneously consider global optimality and local response speed, effectively handling the multi-timescale dynamic characteristics and multi-objective optimization needs in the wastewater treatment process. Compared with traditional centralized optimization, hierarchical optimization models have higher computational efficiency, stronger adaptability, and are particularly suitable for real-time optimization control of complex large-scale systems.
[0178] Example 7
[0179] In this embodiment, after obtaining the global optimization model, the process further includes:
[0180] Based on the set of sub-models after tree decomposition and the wastewater treatment process network diagram, tree coverage metrics are calculated, including node coverage, edge coverage, and critical path retention rate, to obtain tree coverage evaluation results.
[0181] Based on the tree coverage assessment results and computational resource constraints, the tree width parameter of the process network is dynamically adjusted. When the optimization period is shorter than the preset response time threshold, the tree width parameter value is reduced to reduce computational complexity. When the optimization period is longer than the preset response time threshold, the tree width parameter value is increased to improve model accuracy, thus obtaining the adaptively adjusted tree width parameter.
[0182] Based on the adaptively adjusted tree width parameter, the step of decomposing the global optimization model into several sub-models is re-executed to obtain an updated hierarchical optimization model based on tree structure.
[0183] The updated tree-based hierarchical optimization model is mapped to a centralized optimization layer and a distributed autonomous layer. The optimization of the top-level nodes and intermediate nodes of the updated tree-based hierarchical optimization model is implemented in the centralized optimization layer, and the optimization of the leaf nodes is implemented in the distributed autonomous layer, resulting in the mapped tree-based hierarchical optimization model, which serves as the global optimization model.
[0184] Specifically, firstly, based on the set of sub-models after tree decomposition and the wastewater treatment process network diagram, a tree coverage metric is calculated. The tree coverage metric is a set of indicators used to evaluate the quality of tree decomposition, measuring the degree to which the tree decomposition results retain the original process network structure. The calculation process first compares the original process network diagram and the structure after tree decomposition, and then performs quantitative evaluation from three dimensions: nodes, edges, and critical paths. Node coverage measures the completeness of the representation of all nodes in the original network in the tree decomposition. It is calculated by statistically analyzing the proportion of original nodes included in the tree decomposition; ideally, it should reach 100%, indicating that all treatment units are included in the model. Edge coverage measures the degree to which the connections in the original network are retained in the tree decomposition. It is calculated by checking whether the two endpoints of each edge (representing the material flow, energy flow, or information flow between units) in the original diagram appear simultaneously in at least one sub-model; if so, the edge is covered. Critical path retention rate assesses the complete retention of important process paths in the system. It is calculated by identifying critical process paths (such as the main water flow path, critical return path, etc.) in the original diagram, and then checking whether these paths are completely retained or appropriately segmented in the tree decomposition. The system calculates a weighted average of these three indicators to obtain a comprehensive tree coverage index, with the weights dynamically adjusted based on the degree of influence of each indicator on the optimization quality. In addition, the balance of tree decomposition (the uniformity of the size of each sub-model) and the number of boundary nodes (the number of nodes appearing in multiple sub-models) are also evaluated. These indicators together constitute a complete tree coverage evaluation result, providing a quantitative basis for subsequent adjustments.
[0185] Next, based on the tree coverage assessment results and computational resource constraints, the tree width parameter of the process network is dynamically adjusted. First, a model relating the tree width parameter to computational complexity is established, and the correspondence between the optimization cycle and the response time threshold is defined. The optimization cycle refers to the time required to complete one global optimization, including the entire process of data processing, model calculation, and result generation. The response time threshold is the maximum allowable optimization time set according to process requirements and control objectives, typically determined based on the dynamic characteristics and control requirements of the process. For example, rapidly changing processes may require a response time in minutes, while stable operating conditions can accept a response time in hours. By monitoring the execution time of optimization calculations in real time and comparing it with the preset response time threshold, the tree width parameter is dynamically adjusted. When the optimization cycle is shorter than the preset response time threshold, it indicates that current computational resources are abundant, and the tree width parameter value is moderately increased to improve model accuracy and optimization quality. Specifically, the strategy involves selecting the boundary node with the most frequent information flow in the current tree decomposition, merging its surrounding sub-models, and increasing model complexity to obtain more accurate optimization results. When the optimization cycle is longer than the preset response time threshold, it indicates that the computational burden is too heavy, and the tree width parameter value is decreased to reduce computational complexity and accelerate the optimization speed. The reduction strategy involves identifying the largest or most computationally complex sub-model in the current tree decomposition and employing a more aggressive simplification method or further decomposing it into smaller sub-models. The adjustment process considers the tree coverage assessment results, prioritizing the preservation of structures with high coverage while also taking into account the integrity of key nodes and paths. Through this dynamic balancing act, an adaptively adjusted tree width parameter is obtained, achieving an optimal balance between model accuracy and computational efficiency.
[0186] Then, based on the adaptively adjusted tree width parameters, the global optimization model decomposition step is re-executed. First, the parameter configuration of the tree decomposition algorithm is updated, incorporating the adaptively adjusted tree width parameters to control the maximum size and complexity of the sub-models during the decomposition process. The tree decomposition algorithm is then re-executed to obtain an updated set of sub-models. Compared to the original decomposition, the updated decomposition is more adaptable to the current computing resources and optimization needs, either by increasing the size of the sub-models to improve modeling accuracy or by decreasing their size to accelerate computation. During the update of the sub-model set, the selection and allocation of boundary nodes are also optimized, reducing the coupling complexity between sub-models and improving the efficiency of distributed solution. For each updated sub-model, parameter reestimation and equation reconstruction are performed to ensure that the model accurately reflects the dynamic characteristics of the original system. Based on the updated set of sub-models, a hierarchical decision structure is reconstructed, determining the composition, connection relationships, and optimization objectives of each layer of nodes, forming an updated tree-based hierarchical optimization model. While maintaining the original hierarchical structure, the new model reallocates optimization tasks according to the adaptively adjusted tree width parameters, better adapting to the current computing environment and control requirements.
[0187] Finally, the updated tree-based hierarchical optimization model is mapped onto the actual control system architecture. Based on the physical deployment and computing power distribution of the control architecture, different levels of the optimization model are allocated to corresponding hardware platforms. The optimization calculations for top-level and intermediate-level nodes are relatively complex, requiring global information and strong computing power; therefore, they are mapped to a centralized optimization layer and deployed on a central server or computing center. The optimization calculations for leaf nodes are relatively simple, mainly focusing on local control and rapid response; therefore, they are mapped to a distributed autonomous layer and deployed on edge controllers or field devices. The mapping process not only considers matching computing power but also optimizing communication efficiency, minimizing cross-layer data transmission, and allocating frequently interacting nodes to the same hardware platform. Furthermore, the optimization algorithm is customized according to the characteristics of the hardware platform; for example, simplified algorithms with lower computational complexity are used on edge devices with limited computing power, while more accurate but computationally intensive advanced algorithms are used on the central server. Through this refined mapping and adjustment, the mapped tree-based hierarchical optimization model is obtained. As the final global optimization model, it fully leverages the advantages of the two-layer control architecture, achieving an organic combination of global optimization and local autonomy.
[0188] Tree coverage metrics are a set of indicators used to evaluate the quality of tree decomposition. They quantify the ability of the decomposition to represent the original graph structure from multiple dimensions. These metrics not only reflect the mathematical integrity of the model but also focus on the degree to which the actual connectivity relationships of the wastewater treatment process are preserved. Tree coverage metrics include three core indicators: node coverage, edge coverage, and critical path retention. Node coverage assesses the completeness of the representation of treatment units (such as aeration tanks, sedimentation tanks, membrane units, etc.) in the model; ideally, it should be 100%, indicating that all units are included in the model. Edge coverage assesses the degree to which the connectivity relationships between units (such as water flow paths, sludge return, etc.) are preserved; a high edge coverage means that the model can accurately reflect the transfer relationships of matter and energy. Critical path retention focuses specifically on the core process chains (such as main water flow paths, critical control loops, etc.), which have a decisive impact on system performance and must be accurately represented in the model. Tree coverage metrics are typically calculated using graph theory analysis methods, comparing the original graph with the set of subgraphs after tree decomposition to calculate the degree of retention of structural features. In practical applications, these indicators are weighted according to the characteristics of specific wastewater treatment processes. For example, for processes that are mainly biological treatment, more attention may be paid to the coverage of nodes and edges related to the biochemical unit.
[0189] The adaptively adjusted tree width parameter refers to the tree decomposition control parameter dynamically optimized based on system operating conditions and computing resources. Unlike a fixed tree width value, the adaptive tree width parameter can automatically adjust according to real-time needs, seeking the optimal balance between model accuracy and computational efficiency. The adaptive adjustment mechanism is based on the closed-loop feedback control principle. By monitoring the difference between the actual execution time and the target response time of the optimization calculation, it dynamically adjusts the tree width value: when the execution time is too long, it decreases the tree width to simplify the model and speed up computation; when the execution time is sufficient, it increases the tree width to improve model accuracy. The adjustment strategy adopts a gradual approach to avoid drastic changes that could lead to system instability; typically, each adjustment is limited to ±20% of the current tree width. Furthermore, the adjustment process also considers historical data and predictive factors, such as load trends and changes in data complexity, making the adjustment forward-looking and smooth. The adaptive tree width parameter is a key mechanism for achieving efficient utilization of computing resources, maintaining optimal performance under different operating conditions and computational loads, and is particularly suitable for wastewater treatment systems with limited computing resources or frequently changing operating conditions.
[0190] Example 8
[0191] In this embodiment, the reporting to the centralized optimization layer via a two-way communication mechanism includes:
[0192] Based on the running status information after execution and the hierarchical dataset, differential privacy sensitivity is calculated for each type of sensitive data. The differential privacy sensitivity represents the maximum impact of a change in a single record on the query results, thus obtaining the sensitivity parameters for each type of data.
[0193] Based on the sensitivity parameters and privacy budget parameters of the various types of data, the standard deviation of Gaussian noise is calculated, and a time-adaptive noise calibration mechanism is designed to obtain the calibrated Gaussian noise parameters.
[0194] Based on the running status information after execution and the calibrated Gaussian noise parameters, Gaussian noise is applied before the edge controller reports sensitive data. Different levels of privacy protection are applied to data with different sensitivity levels to obtain privacy-protected status reporting data.
[0195] Based on the privacy-protected status reporting data, it is transmitted to the centralized optimization layer through the communication network to realize privacy-protected status reporting.
[0196] Specifically, firstly, based on the post-execution operational status information and the tiered dataset, differential privacy sensitivity is calculated for each type of sensitive data. Differential privacy is a data protection technique that prevents the inference of individual data from the results by adding carefully calibrated random noise to the query results. Before applying differential privacy, it is necessary to first determine the sensitivity of each type of data, that is, the maximum possible impact of a change in a single data record on the statistical results. The calculation process first matches the post-execution operational status information with the tiered dataset to determine the sensitivity level (high, medium, low) of each data point. Then, sensitivity analysis is performed on each type of sensitive data separately. For high-sensitivity data (such as proprietary process parameters, special formulas, etc.), a strict sensitivity calculation method is adopted: first, the value range and distribution characteristics of this type of data are determined; then, the impact on the query results under the worst-case scenario (i.e., when the data undergoes the maximum possible change) is calculated; finally, the sensitivity parameter is determined based on the degree of impact. For example, for proprietary process parameters, a high sensitivity value (such as 5.0) may be calculated, indicating that the data requires strong protection. For moderately sensitive data (such as energy consumption distribution and processing efficiency), a balanced sensitivity calculation method is used, comprehensively considering the actual magnitude of data changes and commercial value, typically yielding a moderate sensitivity value (e.g., 2.0). For low-sensitivity data (such as routine water quality indicators), a more lenient sensitivity calculation is used, resulting in a lower sensitivity value (e.g., 0.5). This tiered calculation yields a sensitivity parameter table for various data types, providing fundamental parameters for subsequent privacy protection.
[0197] Next, the standard deviation of Gaussian noise is calculated based on the sensitivity parameters and privacy budget parameters of various types of data. The privacy budget is a key parameter in the differential privacy mechanism, representing the total amount of privacy loss the system is willing to accept. A smaller budget value means stronger privacy protection but may reduce data utility. First, a global privacy budget value is set according to business needs and regulatory requirements, and then it is allocated to different types of data query and reporting tasks. The allocation principle is to give more budget to important queries and less budget to sensitive data queries, balancing data protection and usage needs. Based on the allocated privacy budget and the previously calculated data sensitivity, the standard deviation of the added noise is calculated: the standard deviation is directly proportional to the sensitivity and inversely proportional to the privacy budget; the higher the sensitivity or the smaller the budget, the greater the noise needs to be added. To cope with data changes over long periods of operation, the system also designs a time-adaptive noise calibration mechanism. This mechanism monitors the changing trends of data distribution and query frequency, dynamically adjusting the noise parameters: when the data distribution is stable, noise is appropriately reduced to improve data utility; when data changes drastically or queries are frequent, noise is increased to strengthen protection; when potential privacy attack patterns (such as high-frequency targeted queries) are detected, protection is temporarily enhanced. Furthermore, the privacy budget is dynamically managed based on time and historical query data to prevent premature depletion of the budget and subsequent data unprotection. Through these mechanisms, calibrated Gaussian noise parameters are obtained, ensuring both the effectiveness of privacy protection and maintaining data availability.
[0198] Then, based on the post-execution operational status information and calibrated Gaussian noise parameters, Gaussian noise is applied before sensitive data is reported by the edge controller. The application process follows a differentiated protection principle: different levels of privacy protection are applied to data with different sensitivity levels. For highly sensitive data, Gaussian noise with a larger standard deviation is applied to ensure the original data value is effectively masked; for moderately sensitive data, noise with a moderate standard deviation is applied to maintain some data utility while protecting the data; for low-sensitivity data, only the minimum necessary noise is applied, or in some cases (such as publicly available data), no noise is added at all. The noise addition process is completed before the data leaves the edge controller, ensuring that the original sensitive data is always retained locally and not transmitted over the network. Furthermore, a context-aware noise adjustment mechanism is employed, considering the correlation and physical constraints between data items to avoid violating physical laws (such as generating negative concentration values) or disrupting the logical relationships between data after noise is added. For critical control parameters that require accuracy, functional mechanisms are used instead of direct noise addition, such as range fuzzification and numerical rounding, to maintain functional effectiveness while protecting accurate values. These processes generate privacy-protected status reporting data, which protects sensitive information while maintaining the functionality and usability of the data.
[0199] Finally, the privacy-protected status reporting data is transmitted to the centralized optimization layer via the communication network. The transmission process employs secure communication protocols, including data encryption, integrity verification, and authentication, further enhancing data protection. TLS / SSL encryption ensures confidentiality during data transmission, preventing eavesdropping and man-in-the-middle attacks; digital signatures and checksum mechanisms ensure data integrity, preventing tampering during transmission; and a two-way authentication mechanism ensures the authenticity of both communicating parties, preventing impersonation attacks. Data compression and batch transmission strategies are also employed during transmission to improve communication efficiency and reduce network load. Upon reaching the centralized optimization layer, the data is verified and decrypted before entering the data management system. The centralized optimization layer's data management system implements multi-level access control on the received data, controlling the visibility and usage of the data based on its sensitivity level and the user's permission level. Furthermore, a complete data access log is maintained, recording all queries and usage behaviors for subsequent auditing and anomaly detection. Through these mechanisms, the system achieves privacy-protected status reporting, ensuring data security while guaranteeing the centralized optimization layer obtains necessary operational status information to support global optimization decisions.
[0200] Differential privacy sensitivity is a quantitative indicator that measures the impact of data changes on query results. Within the differential privacy framework, sensitivity is defined as the maximum possible change in the query function result when a record in the dataset changes (added, deleted, or modified). Formally, if two datasets D1 and D2 differ by only one record, the sensitivity Δf of the query function f is represented as the maximum possible difference between the results of f(D1) and f(D2). In wastewater treatment systems, differential privacy sensitivity calculation needs to consider the physical meaning and relevance of the data: for example, for water quality parameter queries, sensitivity might be the normal fluctuation range of the parameter; for reagent ratio queries, sensitivity might be the key difference in the formulation patent. Sensitivity calculation typically employs a combination of theoretical analysis and historical data statistics: first, theoretical analysis is conducted based on the physical meaning and value range of the data to determine the maximum possible change; then, historical data analysis is used to verify and calibrate this theoretical value to obtain a more accurate sensitivity estimate. Accurate sensitivity calculation is the foundation for effective differential privacy protection; it directly determines the scale of added noise, thus affecting the strength of data protection and the utility of data use.
[0201] Time-adaptive noise calibration is a technique for dynamically adjusting the strength of privacy protection to address privacy budget depletion and data distribution changes in long-running systems. Unlike static noise parameters, the adaptive mechanism adjusts noise parameters in real time based on system state and environmental changes, maximizing data utility while protecting privacy. This mechanism calibrates noise based on several key factors: time (the sensitivity of early data may decrease over time, thus weakening protection); query patterns (monitoring query frequency and patterns to enhance protection against potential attacks); data state (noise can be reduced when the process is stable and increased when the state changes drastically); and privacy budget depletion (dynamically allocating and managing the privacy budget to prevent premature exhaustion). The adaptive calibration process consists of three phases: a monitoring phase that collects system state and query behavior data; an analysis phase that assesses privacy risks and data needs; and an adjustment phase that calculates new noise parameters and applies them to subsequent queries. This mechanism is particularly suitable for environments like wastewater treatment systems that require long-term continuous operation and where data state changes over time, providing effective privacy protection throughout the system's lifecycle while maintaining data availability and accuracy.
[0202] Gaussian noise is a commonly used noise addition method in differential privacy protection. Its characteristic is that the noise values follow a normal distribution with a mean of 0 and a preset standard deviation. Compared to Laplace noise, Gaussian noise performs better in protecting continuous and complex queries, making it particularly suitable for time-series data protection in wastewater treatment systems. The core of the Gaussian mechanism is to calculate an appropriate noise standard deviation based on query sensitivity and privacy parameters, and then sample noise values from this distribution to add to the original query results. In practical implementation, the application of Gaussian noise considers the physical constraints and logical relationships of the data: for example, for parameters that cannot be negative (such as concentration and flow rate), truncated Gaussian distributions or post-processing techniques are used to ensure that the results after adding noise still meet the physical constraints; for data items that need to maintain relationships (such as influent and effluent water balance), relevant noise generation techniques are used to ensure that the necessary data relationships are maintained after adding noise. The advantage of Gaussian noise lies in its good mathematical properties and composability, making it particularly suitable for wastewater treatment optimization systems that require complex analysis and multiple queries, providing privacy protection while maintaining the statistical characteristics and analytical value of the data.
[0203] Example 9
[0204] In this embodiment, after applying Gaussian noise before the edge controller reports sensitive data, applying different levels of privacy protection to data with different sensitivity levels, and obtaining privacy-protected status reporting data, the method further includes:
[0205] Based on the privacy-protected state reporting data, a noise-aware optimization objective function is run in the centralized optimization layer to incorporate data uncertainty into the decision-making process, thereby obtaining an optimization objective that takes noise into account.
[0206] Based on the optimization objective that takes noise into account, a probabilistic inference method is used to estimate and filter out noise, calculate the decision confidence interval, and obtain a near-optimal decision result.
[0207] Based on the near-optimal decision results, the decision quality loss index is calculated by comparing the difference in the optimization objective function values before and after privacy protection, and the decision error evaluation results are obtained.
[0208] Based on the decision error assessment results, a quantitative relationship model between privacy protection strength and decision quality is established. The calibrated Gaussian noise parameters are dynamically adjusted according to the comparison between the decision quality loss index and the preset quality loss threshold to obtain an optimized decision that balances privacy protection and decision quality, which serves as the global optimized decision for the next cycle.
[0209] Specifically, firstly, based on privacy-preserving state reporting data, a noise-aware optimization objective function is run at the centralized optimization layer. Traditional optimization models assume the input data is accurate, but in a differential privacy-preserving environment, the data contains artificially added random noise. Directly using this data for optimization can lead to suboptimal or erroneous decisions. The noise-aware optimization objective function is a specially designed objective function that incorporates the uncertainty in the data into the decision-making process, improving the robustness of decisions in noisy environments. The construction process first identifies key decision variables and their corresponding state data, and assesses the noise level in each data set. For highly sensitive data, due to the added noise, the weight of dependence on these data is reduced, and more reliance is placed on historical patterns and stable characteristics. For moderately sensitive data, appropriate noise estimation and compensation are performed, such as using smoothing techniques to reduce the impact of noise. For low-sensitivity data, due to the low noise level, it can be used with higher confidence. A probabilistic model of data uncertainty is also constructed, explicitly describing the uncertainty distribution of each parameter in the state reporting data. These distributions are related to the added Gaussian noise parameters, but also consider the inherent randomness and measurement errors of the system. Based on these uncertainty analyses, the optimization objective function was reconstructed, shifting from deterministic optimization to a robust or stochastic optimization framework. The reconstructed objective function no longer simply pursues the optimal value, but rather balances expected performance and risk, aiming to achieve a solution with good performance under various possible data implementations. For example, for energy consumption optimization, a traditional function might directly minimize expected energy consumption, while a noise-aware function might minimize "expected energy consumption + risk penalty term," where the risk penalty term is positively correlated with data uncertainty. In this way, an optimization objective considering the impact of noise is obtained, enabling more robust decisions in noisy environments.
[0210] Next, based on the optimization objective considering the impact of noise, a probabilistic inference method is used to estimate and filter noise. The system recognizes that while direct access to the original noise-free data is impossible (which is also for privacy protection purposes), the true characteristics of the signal can be partially recovered through statistical methods and domain knowledge. The noise estimation and filtering process first constructs prior knowledge based on historical data and process models, including the reasonable range of parameters, typical distribution, rate of change, and correlation between parameters. Then, a Bayesian inference framework is used to combine the prior knowledge with the currently observed noise data to estimate the most likely true state. In specific implementation, a variety of techniques are used, including Kalman filtering (suitable for state estimation under linear systems and Gaussian noise assumptions), particle filtering (suitable for state estimation under nonlinear systems and arbitrary noise distributions), and a deep learning-based denoising autoencoder (using the high-dimensional features of the data to learn a denoising map). For different types and sensitivity levels of data, the system applies different filtering strategies: for data with strong temporal continuity (such as continuous monitoring of water quality parameters), the system makes more use of temporal continuity for filtering; for data with strong spatial correlation (such as correlation readings from distributed measuring points), spatial correlation is used for cross-validation and filtering. Based on the estimated true state, the system calculates confidence intervals for the decision variables, quantifying the level of uncertainty in the decision. The calculation of confidence intervals considers multiple sources of uncertainty, including data noise, model error, and system randomness, providing a reliable measure of uncertainty for the decision. Finally, by integrating the optimization objective and uncertainty analysis, near-optimal decision results are obtained, which maintain good performance under various possible true states.
[0211] Then, based on the near-optimal decision results, the decision quality loss caused by privacy protection was evaluated. The evaluation process was based on simulation comparative analysis, constructing two parallel optimization scenarios: one using original data (or simulated noise-free data) to represent the optimal decision under ideal conditions; the other using privacy-protected data to represent the decision in the actual system. By comparing the optimization objective function values under these two scenarios, the system calculated a decision quality loss index. This index can be expressed in several forms, such as relative error percentage (the proportion of the difference between the objective function value of the privacy-protected decision and the optimal decision to the optimal value), absolute performance gap (the difference in actual performance indicators under the two decisions, such as the increase in energy consumption and the decrease in processing efficiency), and the degree of risk increase (the additional operational risks that privacy-protected decisions may cause, such as an increased probability of failure). Sensitivity analysis was also conducted to evaluate the changes in decision quality under different privacy protection intensities (noise levels), and a response curve between privacy intensity and decision quality was established. In addition, the differentiated impact of privacy protection on decision quality for different types of data was analyzed, and the key data types with the greatest impact on decision quality were identified. These analyses formed a comprehensive decision error evaluation result, providing a basis for subsequent optimization of privacy protection strategies. The evaluation results are presented in the form of a multidimensional set of indicators, including overall quality loss measurement, categorical data impact analysis, sensitive parameter impact ranking, and risk assessment results, comprehensively reflecting the impact of privacy protection on decision quality.
[0212] Finally, based on the decision error assessment results, a quantitative relationship model between privacy protection strength and decision quality was established, achieving a balanced optimization of privacy protection and decision quality. This model describes the decision quality loss under different noise parameter settings and is constructed based on extensive historical data analysis and simulation experiments. The model adopts a piecewise function form, reflecting the nonlinear relationship between privacy protection strength and decision quality: typically, in the low-noise region, the decision quality loss increases slowly; after the critical region, the loss increases rapidly. Based on this model, a dynamic noise parameter adjustment strategy was implemented. First, a preset quality loss threshold was defined, which is the maximum decision quality loss the system is willing to accept, usually determined based on business needs and regulatory requirements. The actual decision quality loss is periodically compared with this threshold: when the actual loss is below the threshold, privacy protection may be enhanced (increasing noise) to improve data security; when the actual loss is close to or exceeds the threshold, privacy protection will be moderately weakened (reducing noise) to ensure decision quality. The adjustment process is gradual to avoid drastic changes that could lead to system instability. Furthermore, the adjustments are differentiated, employing different adjustment strategies for data with different sensitivity levels and different impacts. For example, data that is highly sensitive but has little impact on decision quality can be protected more strongly; data that is moderately sensitive but critical to decision-making can be protected less. This dynamic balancing mechanism yields an optimized decision that balances privacy protection and decision quality, serving as the global optimization decision for the next cycle. This continuous optimization closed-loop mechanism ensures that the system maintains high-quality operational decisions while protecting data privacy, achieving the dual goals of security and efficiency.
[0213] Among them, the noise-aware optimization objective function is a specially designed optimization model capable of making high-quality decisions even when the input data contains random noise (especially noise added artificially for privacy protection). Unlike traditional optimization functions, the noise-aware function does not assume that the data is accurate, but explicitly considers the uncertainty in the data and its impact on the decision. Its core principle is to transform deterministic optimization into a stochastic or robust optimization problem, describing the data uncertainty through a probabilistic model and integrating this uncertainty into the decision-making process. Specific implementation methods include several main modes: the expected optimization mode, which aims to maximize / minimize the average performance under all possible real data states; the risk-averse mode, which aims to maximize / minimize the performance guarantee in the worst case; and the chance-constrained mode, which aims to satisfy the constraints under a certain probability guarantee. In wastewater treatment systems, the noise-aware function may be formalized as "minimize expected energy consumption + λ × energy consumption variance", where λ is the risk aversion coefficient, and the energy consumption variance reflects the instability of the decision under noisy data. This function design enables the system to maintain high-quality optimization decisions while protecting sensitive data, and is a key technology for achieving a balance between privacy protection and system performance.
[0214] Decision quality loss metrics are a set of quantitative measures to assess the impact of privacy protection measures on optimization decisions. These metrics measure the degree to which decisions deviate from the ideal optimal solution due to added privacy protection noise from multiple dimensions. Decision quality loss metrics typically include three key dimensions: performance loss, which measures degradation in the objective function value, such as the rate of increase in energy consumption or the percentage increase in processing costs, usually expressed as relative values; reliability loss, which measures the increased risk of decisions failing to meet constraints, such as an increased probability of exceeding water discharge standards or reduced system stability; and robustness loss, which measures the decreased adaptability of decisions to changes in operating conditions, such as a smaller adjustment range for control parameters or an increased system response delay. These metrics are usually obtained through comparative experiments, i.e., optimization is performed using noisy and noiseless (or low-noise) data respectively, and then the differences in results are compared. In practical applications, these metrics are weighted and combined according to the system's operational priorities to form a comprehensive decision quality loss metric, serving as the basis for dynamically adjusting the strength of privacy protection. By continuously monitoring these metrics, the system can find the optimal balance between protecting data privacy and maintaining optimization effectiveness, achieving the goal of "near-optimal control under privacy protection."
[0215] Example 10
[0216] In this embodiment, the step of running an optimization algorithm within the constraints of the control parameters to obtain real-time control commands at the unit level includes:
[0217] Based on the local sensor data and the control parameter constraints, the core treatment units affecting the effluent quality are identified. The core treatment units include a biological reactor, a membrane filtration unit, and a disinfection unit, thus obtaining a set of key treatment units.
[0218] Based on the control problem of the set of key processing units, the discretized control parameter values within the set interval of the control parameter constraints are modeled as a matroid structure. The basic element set is defined as the set of control parameter values discretized within the set interval of the control parameter constraints according to a preset step size. The independent set family is defined as the set of parameter combinations that meet the process requirements, thus obtaining the matroid model of the wastewater treatment process.
[0219] Based on the aforementioned wastewater treatment process matroid model, matroids corresponding to different optimization objectives are constructed, including energy efficiency matroid, treatment efficiency matroid, and stability matroid. Through matroid cross-operation, parameter combinations that simultaneously satisfy multiple constraints are found, resulting in a multi-objective matroid structure.
[0220] Based on the multi-objective matroid structure, a parallel basis search algorithm is run to find a set of mutually independent basis, each basis corresponding to an optimization dimension, resulting in multiple parallel basis;
[0221] Based on the multiple parallel bases, each parallel base is assigned to different computing threads for parallel computation. The optimization results of each parallel base are combined through a weighted voting-based fusion algorithm to obtain the real-time control instructions at the unit level.
[0222] Specifically, firstly, based on local sensor data and control parameter constraints, the core treatment units affecting effluent quality are identified. A multi-stage screening method is used to identify key treatment units: The first stage is a static importance analysis based on the process flow, which, based on process design and expert knowledge, preliminarily identifies key links in the treatment chain, such as the biological reactor (responsible for the removal of organic matter and nitrogen), the membrane filtration unit (responsible for solid-liquid separation and partial removal of dissolved substances), and the disinfection unit (responsible for pathogen inactivation). The second stage is a dynamic sensitivity analysis based on historical data, which quantifies the impact of each unit by analyzing the statistical correlation and causal relationship between changes in parameters of each treatment unit and changes in effluent quality indicators. Various data analysis methods are used, including correlation analysis (calculating the Pearson correlation coefficient or Spearman rank correlation coefficient between the control parameters of each unit and the effluent indicators), principal component analysis (identifying the treatment unit characteristics that best explain the effluent quality variations), and gradient-based sensitivity analysis (calculating the partial derivatives of the effluent indicators with respect to the parameters of each unit to assess the magnitude of the impact of parameter changes). The third stage is a real-time importance assessment based on the current operating conditions, which dynamically adjusts the importance weight of each unit according to the current influent characteristics, equipment status, and treatment objectives. For example, under high ammonia nitrogen load conditions, the nitrification function in the biological reactor receives a higher importance score; when treating reclaimed water, the integrity and performance of the membrane filtration unit are of greater concern. Based on the analysis results from these three stages, a ranking of the treatment units' importance was established, and the top-ranked units were selected to form a set of critical treatment units. This set is not static but dynamically updated as operating conditions change and treatment objectives are adjusted, ensuring that optimized resources are always focused on the aspects that have the greatest impact on current water quality.
[0223] Next, based on the control problem of the key processing unit set, the discretized control parameter values within the set interval of the control parameter constraints are modeled as a matroid structure. A matroid is an algebraic structure that can abstractly represent a set of elements satisfying specific independence conditions, making it particularly suitable for handling discrete optimization problems with complex constraints. The modeling process first defines the basic element set, i.e., all possible values in the control parameter space. The continuous control parameter interval is then discretized, dividing the set interval of each parameter into finite discrete value points according to a preset step size. The choice of discretization step size balances optimization accuracy and computational complexity; smaller step sizes are used for important parameters to improve accuracy, while larger step sizes are used for less important parameters to reduce complexity. For example, the dissolved oxygen setpoint of the bioreactor might range from 1.0 mg / L to 3.0 mg / L, discretized in increments of 0.1 mg / L; the flux setpoint of the membrane filtration unit might range from 15 LMH to 25 LMH, discretized in increments of 1 LMH; and the reagent dosage of the disinfection unit might range from 1.5 mg / L to 3.0 mg / L, discretized in increments of 0.3 mg / L. After discretization, a candidate set containing all possible parameter combinations is obtained. Then, a family of independent sets is defined, which is the set of parameter combinations that satisfy the process requirements. The definition of independent sets is based on several constraints: physical feasibility constraints (parameter values must be within the physical limitations of the equipment), process safety constraints (parameter combinations cannot lead to process instability or safety risks), treatment effect constraints (parameter combinations must achieve the expected treatment effect), and energy efficiency constraints (energy consumption of parameter combinations cannot exceed budget limits). These constraints are formalized as the independence axiom, that is, any subset of parameter combinations that satisfies the constraints is considered an independent set. Based on these definitions and constraints, a complete matroid model of the wastewater treatment process was constructed. The basic elements are discretized control parameter values, and the family of independent sets represents the set of feasible parameter combinations that satisfy various constraints. This matroid model provides a formalized problem statement for subsequent optimization algorithms, enabling the solution of complex wastewater treatment control problems using mature combinatorial optimization theory and algorithms.
[0224] Then, based on the wastewater treatment process matroid model, specific matroids corresponding to different optimization objectives were constructed. Since wastewater treatment systems typically need to balance multiple objectives such as energy efficiency, treatment efficiency, and stability, a dedicated matroid structure was built for each objective. The energy efficiency matroid focuses on minimizing energy consumption. Its basic elements are the same as those of the wastewater treatment process matroid, but the definition of independent sets is more stringent, requiring parameter combinations to minimize energy consumption while meeting basic treatment requirements. Based on historical data and energy consumption models, the expected energy consumption of each parameter combination is evaluated, and combinations with energy consumption below a certain threshold are defined as independent sets of energy efficiency matroids. The treatment efficiency matroid focuses on maximizing pollutant removal, and its independent sets include parameter combinations that can achieve treatment effects higher than expected. Based on water quality models and historical performance data, the system predicts pollutant removal rates under different parameter combinations and includes high-efficiency combinations in its independent sets. The stability matroid focuses on the robustness and anti-interference ability of the system operation, and its independent sets include parameter combinations that can maintain stable operation under various disturbance conditions. By simulating the system response under different operating condition fluctuations, the stability performance of parameter combinations is evaluated, and robust combinations are included in the independent sets. These three matroids represent different optimization directions, but they overlap and conflict; for example, a highly efficient parameter combination may have high energy consumption, while a combination with good stability may have low efficiency. To find a solution that balances multiple objectives, the intersection and the largest common independent subset of the three matroids are calculated through matroid cross-operation, which represents the set of parameter combinations that simultaneously satisfy the requirements of energy efficiency, efficiency, and stability. This multi-objective matroid structure enables the system to systematically handle complex multi-objective optimization problems and find the optimal balance point among the objectives.
[0225] Next, based on the multi-objective matroid structure, a parallel basis search algorithm is run to find the optimal parameter combination. In matroid theory, a basis is a maximal independent set that cannot be added without maintaining its independence. In wastewater treatment optimization, a basis represents a set of coordinated control parameter settings that can achieve a specific optimization objective. Parallel basis search is a special type of matroid optimization algorithm that simultaneously searches for bases in multiple optimization directions, each base corresponding to an optimization dimension. The algorithm first initializes multiple empty sets, each corresponding to an optimization dimension (such as energy efficiency, treatment efficiency, stability, etc.). Then, following a greedy strategy, the algorithm gradually adds elements (control parameter values) to each set until a maximal independent set (i.e., a basis) is formed. During the addition process, the algorithm uses a specific evaluation function to assess the contribution of candidate elements to the corresponding optimization dimension and selects the element with the largest contribution to add to the set. To ensure balance between different optimization dimensions, the algorithm checks the impact of each added element on other dimensions, preventing extreme optimization in one dimension from causing severe deterioration in other dimensions. Furthermore, the algorithm implements a diversity preservation strategy to ensure that the multiple bases searched have sufficient diversity, providing a variety of decision options. Through multiple rounds of iterative search, several parallel bases were obtained, each representing a feasible combination of control parameters, with different focuses on optimization directions. For example, the energy efficiency base prioritizes minimizing energy consumption, the treatment efficiency base prioritizes maximizing pollutant removal, and the stability base prioritizes maximizing operational stability. These parallel bases together constitute the decision space, providing a rich pool of candidate solutions for generating the final control command.
[0226] Finally, based on multiple parallel bases, a parallel computing and integrated decision-making strategy is employed to generate the final control instructions. To fully utilize computing resources and accelerate the decision-making process, the system assigns each parallel base to different computing threads for parallel computation. Each thread is responsible for evaluating the specific performance of the control scheme corresponding to a base under the current operating conditions, including expected energy consumption, processing effect, stability indicators, and other key performance parameters. The evaluation process utilizes a performance prediction model trained on historical data, enabling rapid and accurate estimation of the effects of different control schemes. The computation results of each thread are aggregated to the central decision-making unit via a data bus for comprehensive evaluation and final decision-making. The decision-making process employs a weighted voting-based fusion algorithm to intelligently integrate the optimization results of each parallel base. The voting weights are dynamically adjusted according to the current operating objectives and conditions. For example, during peak energy cost periods, the weight of the energy efficiency base is increased; when the influent load fluctuates significantly, the weight of the stability base is increased; and when the effluent indicators are close to the limit, the weight of the processing efficiency base is increased. Furthermore, the continuity of historical decisions is considered to avoid drastic fluctuations in control parameters and maintain the stability of system operation. The fusion algorithm integrates various factors to calculate the final setpoint of each control parameter, forming a complete set of control instructions. These instructions include specific operating parameters for each key piece of equipment, such as blower speed, pump flow rate, valve opening, and chemical dosage. Ultimately, these instructions are converted into standardized control signal formats and distributed to each field device via a control network to achieve real-time control. Through this parallel optimization and integrated decision-making mechanism based on matrix theory, efficient, stable, and balanced control strategies can be rapidly generated under complex and ever-changing wastewater treatment conditions, achieving refined real-time control at the unit level.
[0227] Among them, matroid is an algebraic structure used to abstractly represent a set system that satisfies specific independence conditions. Unlike matrices in traditional linear algebra, matroid focuses more on the combinatorial relationships between elements rather than numerical computation. Formally, a matroid is a binary tuple (E, I), where E is a finite set of elements and I is a family of subsets of E, satisfying the following three axioms: (1) Non-emptiness: the empty set belongs to I; (2) Heredity: if X belongs to I, then any subset of X also belongs to I; (3) Commutativity: for any X, Y belongs to I, if |X| < |Y|, then there exists an element e belonging to Y\X such that X∪{e} belongs to I. In wastewater treatment control, matroid models are particularly suitable for representing parameter selection problems under complex constraints: the basic element set E can be all possible control parameter values; the family of independent sets I can be a set of parameter combinations that satisfy process safety and performance requirements. The core advantage of matroid lies in its ability to elegantly handle combinatorial optimization problems, especially in multi-constraint, multi-objective environments. Compared with traditional linear or nonlinear programming methods, matroid-based optimization typically has better computational efficiency and robustness, and is particularly suitable for handling real-time control problems with discrete decision spaces and complex constraints, such as multi-parameter collaborative optimization in wastewater treatment systems.
[0228] Parallel basis search (PBS) is a multi-objective optimization method based on matroid theory, specifically designed to simultaneously search across multiple dimensions to find the optimal set of solutions that satisfy different optimization objectives. In matroid theory, a basis is the largest independent set that cannot be added to maintain its independence, representing an extreme point in the solution space. Traditional matroid optimization typically searches for only one basis, while PBS maintains and optimizes multiple bases simultaneously, each corresponding to a different optimization direction or objective function. The core idea of the algorithm is to construct multiple bases in parallel along multiple different optimization directions using a greedy strategy. The specific process includes: an initialization phase, creating an empty set as a prototype basis for each optimization dimension; an expansion phase, evaluating candidate elements for each prototype basis according to the corresponding objective function, selecting the optimal element to add to the basis, until a largest independent set is formed; a coordination phase, assessing the conflicts and complementarities between bases, and possibly making local adjustments to improve the quality of the overall solution; and diversity maintenance, ensuring sufficient differences between different bases to avoid solution centralization. Compared to traditional multi-objective optimization methods, parallel basis search has the advantage of systematically exploring a multi-dimensional solution space to find multiple high-quality solutions representing different trade-offs, while ensuring the independence and feasibility of these solutions. This characteristic makes it particularly suitable for complex control problems such as wastewater treatment systems that require balancing multiple objectives, including energy efficiency, treatment quality, and stability.
[0229] Weighted voting-based fusion algorithms are integrated decision-making methods used to synthesize multiple optimization results or decision schemes to generate final control commands. Based on the idea of democratic decision-making, this algorithm treats each parallel basis as a "voter," and its recommended control scheme as a "vote." The final decision is formed by weighted aggregation of these votes. The core steps of the algorithm include: weight determination, assigning voting weights to each basis based on the current system state, operating objectives, and the reliability of each basis; standardization transformation, unifying the parameter values recommended by different bases to the same computational scale; weighted aggregation, calculating the weighted average or median of all basis recommendations for each control parameter; consistency check, verifying whether the fusion result meets all necessary constraints and making adjustments if necessary; and smoothing, considering historical control values to avoid drastic parameter fluctuations. Compared with simple averages or majority voting, weighted voting fusion considers the differences in reliability and importance of different decision sources, producing higher quality and more balanced decision results. This algorithm is particularly valuable in wastewater treatment control because it can flexibly adapt to changes in control preferences under different operating conditions while maintaining the stability and consistency of decisions, ensuring the reliable operation of the control system.
[0230] Example 11
[0231] In this embodiment, the process of assigning each parallel basis to different computing threads for parallel computation, and then combining the optimization results of each parallel basis using a weighted voting-based fusion algorithm to obtain the unit-level real-time control instructions includes:
[0232] Based on the computing resources of the multiple parallel bases and the edge controller, the computing resource allocation of each parallel base is dynamically adjusted according to the process status and control requirements, so as to allocate more computing resources to the bases corresponding to key process parameters, thereby obtaining an adaptive resource allocation scheme.
[0233] According to the adaptive resource allocation scheme, an information sharing protocol and a co-evolution mechanism are realized among parallel bases, enabling each base to learn from and improve each other during the iteration process, and obtain the co-optimized parallel base results;
[0234] Based on the parallel basis results after collaborative optimization, an anomaly detection mechanism is used to identify and exclude basis results that deviate too much. The contribution of each parallel basis is evaluated using a Shapley value-based method, the weight of each parallel basis in the weighted voting is determined, and the real-time control command at the unit level is obtained.
[0235] Specifically, firstly, a dynamic resource allocation strategy is implemented based on the computing resources of multiple parallel bases and edge controllers. Edge controllers typically have limited computing resources, including processor cores, memory, and power. In multi-base parallel computing, these resources need to be intelligently allocated according to current process requirements and priorities. The system first performs a real-time assessment of the current process state, identifying process units and parameters in critical operating states. The assessment criteria include multiple indicators: deviation degree (the degree of deviation of parameters from target values), rate of change (the trend and speed of parameter changes), process importance (the degree of impact of parameters on the overall processing effect), and anomaly risk (potential process problems or safety hazards). The assessment process uses a multi-indicator comprehensive scoring method to calculate a criticality score for each process unit and control parameter. Based on the criticality score, a priority allocation strategy is determined: bases corresponding to critical process parameters receive more computing resources, while bases corresponding to non-critical parameters receive fewer resources. Resource allocation adopts a dynamic adjustment mechanism, updating in real time according to changes in the process state during system operation. Specific allocation methods include processor core allocation (allocating more CPU cores or higher execution priority to high-priority bases), computational precision adjustment (using more accurate algorithms and smaller step sizes for high-priority bases, and simplified algorithms and larger step sizes for low-priority bases), and iteration count control (allowing more iterations for high-priority bases to obtain more accurate results, while limiting iterations for low-priority bases to save resources). Furthermore, resource allocation strategies are continuously optimized based on historical optimization results and resource utilization efficiency, such as reducing computational resource investment that does not significantly improve optimization results and increasing resource allocation for sensitive operating conditions. Through this adaptive resource allocation scheme, maximum utilization of computational efficiency is achieved with limited computational resources, ensuring sufficient optimization quality for key process parameters while maintaining the responsiveness and real-time performance of the overall control system.
[0236] Next, based on the adaptive resource allocation scheme, an information-sharing protocol and a co-evolutionary mechanism among parallel bases are implemented. Traditional parallel base search algorithms typically assume that each base is independent, but in actual wastewater treatment optimization, there are complex mutual influences and constraints between different optimization objectives corresponding to each base. To fully utilize these correlations, a structured information-sharing protocol is designed, enabling each parallel base to exchange and utilize each other's intermediate results and optimization experience during the computation process. The information-sharing protocol defines the exchange format and mechanism for four types of key information: search space information (feature descriptions of explored and promising regions), performance evaluation information (evaluation results and performance indicators of different parameter combinations), constraint conflict information (difficulties and solutions encountered by each base in satisfying constraints), and anomaly information (discovered extreme cases and special operating condition handling strategies). Efficient data structures and compression coding schemes are used to ensure low overhead and real-time performance of information exchange. Based on the shared information, a co-evolutionary mechanism for parallel bases is implemented, enabling each base to learn from and improve each other during the iteration process. Co-evolution comprises four core components: cross-learning (one basis learns effective strategies from the successful experiences of other bases), complementary collaboration (each base focuses on different optimization directions, jointly constructing a complete solution space), conflict coordination (when the optimization directions of different bases conflict, a balance is found through negotiation), and diversity maintenance (ensuring sufficient diversity among bases to avoid premature convergence to local optima). In its implementation, techniques such as adaptive weight adjustment, policy cross-learning, and solution space partitioning are employed to promote efficient collaboration among parallel bases. Through multiple rounds of iteration and collaborative optimization, each parallel base continuously improves and adjusts its search strategy, ultimately generating optimization results that focus on its own optimization direction while maintaining overall balance. This collaborative optimization mechanism significantly improves the overall optimization effect, overcomes the limitations of single-objective optimization, and provides high-quality, diverse candidate solutions for subsequent decision fusion.
[0237] Then, based on the parallel basis results after collaborative optimization, anomaly detection and contribution evaluation are performed to generate the final control command. In multi-basis parallel optimization, individual bases may produce abnormal or extreme optimization results due to algorithm convergence problems, local process model errors, or special operating conditions. If these results directly participate in decision fusion, they may lead to unstable or suboptimal control commands. To address this issue, the system designs a multi-layered anomaly detection mechanism: first, statistical anomaly detection is performed, using methods such as box plots or Z-scores to identify basis results that deviate significantly from the statistical distribution; then, physical consistency checks are performed to verify whether the basis results conform to process physical constraints and energy balance relationships; finally, historical trend comparisons are performed to check whether the results are consistent with historical successful optimization patterns. For detected abnormal results, a hierarchical processing strategy is adopted: minor anomalies are corrected and included (adjusted to a reasonable range through interpolation or constraint projection); moderate anomalies have their weights reduced (significantly reducing their influence in the final fusion); and severe anomalies are completely excluded (not participating in decision fusion calculations). After excluding anomalies, the system needs to evaluate the contribution of each parallel basis to the final decision as the basis for weighting during fusion. The evaluation employs a Shapley value-based approach, a fair allocation mechanism derived from cooperative game theory that quantifies the marginal contribution of each participant (in this case, each parallel basis) to the collective outcome. The calculation process considers all possible basis combinations, evaluating the gain of each basis when added to different combinations, ultimately yielding a Shapley value reflecting the true contribution. Based on the calculated Shapley values, corresponding decision weights are assigned to each parallel basis: bases with higher contributions receive greater voting power, while bases with lower contributions receive smaller weights. Furthermore, the specialization and applicability of the basis results are considered, with additional weights given to bases particularly suitable for specific operating conditions in relevant decision dimensions. Finally, through a weighted voting mechanism, the optimization results of each parallel basis are synthesized to generate unit-level real-time control instructions that balance multiple objectives. These instructions are sent to the execution layer in a standardized format, enabling precise control of each processing unit. This contribution-based fusion mechanism continuously generates high-quality, stable, and reliable control strategies under complex and variable operating conditions, ensuring the efficient operation of the wastewater treatment system.
[0238] The adaptive resource allocation scheme is a strategy system that dynamically adjusts computing resources, intelligently allocating limited computing power based on the current process status and control requirements. Unlike static resource allocation, the adaptive scheme can respond to system changes in real time, concentrating computing power on the most critical process steps requiring optimization. The core of this scheme is the "process criticality assessment model," which dynamically calculates the priority scores of each process unit and control parameter using multi-dimensional indicators (such as parameter deviation, response time requirements, and risk levels). Resource allocation follows a variation of the "80 / 20 rule," where approximately 20% of critical parameters receive approximately 80% of the computing resources, ensuring sufficient optimization quality for critical steps. The allocation mechanism employs a multi-layered strategy: tiered by computational precision (high-precision algorithms are used for critical parameters, while simplified algorithms are used for non-critical parameters), tiered by execution frequency (critical parameters are optimized more frequently, while the optimization cycle for non-critical parameters is extended), and tiered by iteration depth (critical parameters are allowed more iterations, while non-critical parameters have earlier stopping conditions). A "resource benefit evaluation mechanism" is also designed to continuously monitor the optimization effect improvement brought about by additional computing resource investment, automatically adjusting the allocation strategy when diminishing returns become apparent. This adaptive resource allocation scheme is particularly suitable for edge computing environments. It can maximize control performance under conditions of limited computing resources, while maintaining system response speed and scalability. It is a key supporting technology for intelligent edge control.
[0239] The Shapley value-based contribution assessment method is a fair allocation mechanism derived from cooperative game theory, used to quantify the marginal contribution of each parallel basis in collective decision-making. The core idea of the Shapley value is to consider the marginal value increment brought about by participants (parallel bases) joining various possible alliances (combinations of bases), and then take a weighted average of these increments to obtain the fair contribution value of each participant. In wastewater treatment optimization, this method is particularly suitable for evaluating the contribution of different optimization objectives (such as energy efficiency, treatment efficiency, stability, etc.) to the final control effect. The calculation process considers all possible base combinations: first, a value function is defined to measure the collective performance of any base combination (such as the comprehensive control effect score); then, for each base, the value difference before and after its addition in all possible base combinations is calculated; finally, these differences are averaged according to specific weights to obtain the Shapley value of that base. Compared with simple performance scoring, the Shapley value considers the synergistic effects and complementary relationships between bases, and can more fairly and comprehensively reflect the true contribution of each base. In practical applications, since the computational complexity of the complete Shapley value increases exponentially with the number of basis cells, approximate calculation methods such as Monte Carlo sampling are usually used to reduce the computational burden while maintaining the fairness of the evaluation. The contribution determined by the Shapley value is directly converted into weight allocation in decision fusion, ensuring that the final control command fully utilizes the beneficial characteristics of each optimization objective, and achieving multi-objective balanced optimal control.
[0240] Example 12
[0241] In this embodiment, the method further includes fault tolerance and emergency handling steps, including:
[0242] The communication status between the centralized optimization layer and the distributed autonomous layer is detected. When a communication interruption is detected, the autonomous operation mode of the distributed autonomous layer is started. Based on the historical operation data stored in the edge controller and the control parameter constraints received last time, the autonomous control command under the communication interruption is obtained.
[0243] Based on the prediction deviation between the prediction results of the global optimization model and the actual measured values, the system detects whether the global optimization model has failed. When the prediction deviation exceeds a preset threshold, the system switches to a rule-based control mode to obtain a safety control command for model failure.
[0244] The system monitors the operating status of hardware components, including a central server, edge controllers, sensors, and actuators. When a failure is detected in any of the hardware components, a backup system switching mechanism is activated, and detailed system operation logs are recorded to achieve fault-tolerant operation of the system.
[0245] Specifically, the first step is to monitor the communication status between the centralized optimization layer and the distributed autonomous layer. During normal operation, the two control architectures continuously exchange information through the communication network: the centralized optimization layer issues global optimization decisions and control parameter constraints, while the distributed autonomous layer reports its operational status and execution feedback. Multiple mechanisms monitor communication health: periodic heartbeat checks (periodic signal exchange every few seconds to tens of seconds to verify the continuity of the communication link), packet integrity checks (using checksums or hash values to verify the integrity of transmitted data and prevent data corruption), and communication latency monitoring (recording and analyzing the time delay of information transmission to identify network congestion or performance degradation). When a communication anomaly is detected, a communication recovery procedure is first attempted, including link reconnection, communication protocol switching, and activation of backup channels. If normal communication cannot be restored within a short time, the system determines it as a communication interruption and immediately triggers an emergency response mechanism. In the event of a communication interruption, the distributed autonomous layer automatically switches to autonomous operation mode, no longer relying on real-time guidance from the centralized optimization layer. Autonomous operation relies on two key types of information: first, historical operational data stored in the edge controller, including recent process parameter records, control behavior records, and performance evaluation records, which reflect the system's behavior patterns under normal operation; and second, the control parameter constraints received last time, which define the reasonable range and safety boundaries of control parameter variation. Based on this information, the edge controller uses a simplified local optimization algorithm to calculate and execute control commands. The autonomous control strategy prioritizes the safe and stable operation of the system, potentially sacrificing some optimization effects for higher reliability. A communication recovery mechanism is also implemented, continuously attempting to re-establish the connection with the centralized layer and recording detailed operating data during autonomous operation to facilitate process analysis and optimization adjustments after communication is restored. This autonomous operation capability enables the maintenance of basic functions in the event of communication interruption, ensuring the continuity and safety of the processing and preventing process interruptions or loss of control due to communication problems.
[0246] Next, the effectiveness of the global optimization model is assessed based on prediction bias. The global optimization model is the core foundation for centralized optimization layer decision-making, and its accuracy and applicability directly affect the quality of optimization decisions. To ensure the model's continued effectiveness, a model performance monitoring mechanism is implemented, periodically comparing the differences between model predictions and actual measurements. The monitoring process first collects actual measurement data of key process parameters, including water quality indicators (such as COD, ammonia nitrogen, and total phosphorus), process parameters (such as dissolved oxygen, MLSS, and pH), and equipment status indicators (such as energy consumption, flow rate, and pressure). Then, the historical values of these parameters are retrospectively input into the global optimization model to obtain the model's prediction results for the current moment. By comparing the predicted and measured values, the system calculates various prediction bias indicators: absolute error (the direct difference between the predicted and measured values), relative error percentage (the proportion of error to the measured value), root mean square error (reflecting overall prediction accuracy), and trend consistency (the degree of matching between the predicted and actual directions of change). A multi-tiered preset threshold system is established, corresponding to different levels of model deviation: warning threshold (requires attention but the model can continue to be used), intervention threshold (requires model parameter adjustment), and failure threshold (the model is unreliable and requires switching control modes). When the prediction deviation exceeds the failure threshold, it indicates that the global optimization model may no longer be applicable to the current system state due to process changes, equipment aging, or abnormal operating conditions. At this point, the system automatically switches to rule-based control mode, a robust control strategy that does not rely on complex models. Rule-based control is based on predefined expert knowledge and rules of thumb, employing an "if-then" structured decision logic to directly determine control behavior based on the current process state. For example, "If dissolved oxygen is below 1.5 mg / L and ammonia nitrogen concentration is high, then increase aeration." These rules consider process safety boundaries and basic optimization principles. Although they may not be as efficient as model optimization, they possess high robustness and interpretability. The system has multiple pre-built rule bases covering different operating conditions and processing objectives. After switching to rule-based control mode, the rule set most suitable for the current state is selected for execution. Simultaneously, continuous model diagnostics and repair attempts are conducted, including data validation, parameter recalibration, and model structure adjustments, in order to restore model effectiveness. This model effectiveness monitoring and emergency control switching enables safe and reliable operation to be maintained even in the event of model failure, preventing process problems caused by erroneous decisions.
[0247] Finally, the operational status of all hardware components is comprehensively monitored to achieve fault-tolerant operation. The hardware architecture of the wastewater treatment control system comprises multi-layered components: a central server (bearing the computation and decision-making functions of the centralized optimization layer), edge controllers (computing units that realize distributed autonomous control), field sensors (sensing devices that collect process parameters and status data), and actuators (mechanical and electrical equipment that implements control actions). A multi-faceted monitoring strategy is adopted to assess the health status of these hardware components in real time. Monitoring content includes performance indicators (such as CPU utilization, memory usage, response time, etc.), self-diagnostic information (built-in status detection and error reporting of equipment), communication quality (signal strength, bit error rate, communication latency, etc.), and abnormal behavior patterns (such as data jumps, control non-response, etc.). When the system detects a failure in any hardware component, a tiered response strategy is immediately triggered. For minor failures (such as performance degradation but complete functionality), an early warning is issued and self-repair attempts are made (such as restarting services, clearing cache, etc.); for moderate failures (partial functional impairment), a functional degradation mode is initiated to maintain the operation of core functions; for severe failures (complete component failure), a backup system switching mechanism is activated. Backup switching is the core mechanism of hardware fault tolerance, with corresponding strategies designed for different components: the central server employs hot or warm standby redundancy, allowing the standby server to quickly take over all functions when the primary server fails; edge controllers use a function takeover mechanism, allowing adjacent controllers to temporarily take over the responsibility area of the failed controller; the sensor system uses data redundancy and soft measurement technology, providing necessary data through software estimation or backup sensors when physical sensors fail; actuators are equipped with a manual bypass system, allowing operators to intervene manually when automatic control fails. During component switching and fault handling, all relevant events and state changes are recorded in detail, forming a complete system operation log. These logs contain information such as the system state before and after the fault, fault characteristics, response actions, and result evaluation, providing valuable data for subsequent fault analysis and system improvement. Through this comprehensive hardware monitoring and fault tolerance mechanism, the system can maintain basic functional operation under various hardware failure conditions, significantly improving the overall system reliability and availability, and ensuring the continuity and safety of the wastewater treatment process.
[0248] The autonomous operation mode is an independent operating state for the edge controller, automatically activated when communication with the centralized optimization layer is interrupted. Unlike the conventional subordinate control mode (where the edge controller strictly executes instructions issued by the centralized layer), the autonomous mode allows edge devices to make independent decisions and controls based on locally stored knowledge and data. The core of autonomous operation is a "local control knowledge base," which contains three key components: a historical control strategy cache (recording recently successful control parameters and action sequences); a simplified process model (describing the core dynamic characteristics of local process units, with computational complexity far lower than the global model); and safety control rules (defining safety boundaries for control parameters and emergency response strategies). When communication is interrupted, the edge controller first assesses the current operating conditions, compares them with historical patterns to determine similar scenarios, and then performs local optimization based on the simplified model, generating control instructions under the constraints of safety rules. Compared to centralized optimization, the autonomous mode typically employs a more conservative control strategy, prioritizing process stability and safety, potentially sacrificing some optimization effectiveness. Autonomous operation also includes a "tiered degradation mechanism" that adjusts the control strategy based on the duration of communication interruption: short-term interruptions (hourly) maintain near-optimal control; medium-term interruptions (daily) gradually shift to robust control; and long-term interruptions (weekly) may switch to a safe maintenance mode. This mechanism enables the system to maintain basic functions for an extended period before communication is restored, and is a key guarantee for the resilience and robustness of distributed control systems.
[0249] Rule-based control is a direct control strategy that does not rely on complex mathematical models. It maps observed states to control actions using predefined logical rules. This approach originates from expert systems, encoding the knowledge and experience of human experts into structured rules, making it suitable for control in environments with model failure or high uncertainty. Rules typically take the form of condition-action: "IF [condition combination] THEN [control action]", where conditions can be simple threshold judgments or complex logical combinations, and control actions can be parameter adjustments, equipment start-up / shutdown, or mode switching. Rule-based control in wastewater treatment systems typically includes four types of rules: basic maintenance rules (ensuring basic treatment functions and safety requirements), efficiency optimization rules (improving energy and resource utilization within safe limits), anomaly response rules (detecting and handling process anomalies and equipment failures), and smooth transition rules (ensuring the gradual and stable nature of control actions). The rule base is organized hierarchically, containing meta-rules (for rule selection and conflict resolution), macro-rules (high-level control strategies), and micro-rules (specific operational instructions). Compared to model predictive control, rule-based control has advantages such as high computational efficiency, robustness to uncertainty, and ease of understanding and maintenance; its disadvantages include lower optimization accuracy and difficulty in handling complex multivariate interactions. In modern control systems, rule-based control is often used as a supplement and backup to model-based control, providing reliable control assurance in the event of model failure or special operating conditions. It is an important component of multi-level fault-tolerant control architecture.
[0250] The standby system switching mechanism is a hardware-level fault-tolerance technology that ensures uninterrupted system functionality in the event of a critical component failure through redundant configuration and automatic switching strategies. Unlike simple backup, standby switching is a dynamic and intelligent process encompassing three key stages: fault detection, state transition, and functional takeover. Fault detection employs multi-dimensional monitoring technologies, including heartbeat detection, performance indicator monitoring, self-diagnostic reports, and cross-validation, ensuring rapid and accurate identification of component failures. State transition is the core of the switching process, involving the secure transfer of current state data, historical records, and operational configurations, ensuring seamless takeover from the point of failure. Functional takeover is the process of activating the standby component and assuming the responsibilities of the original component. Depending on the redundancy configuration type, it is categorized as hot switchover (the standby component continues to run and immediately takes over), warm switchover (the standby component is in standby mode, quickly starts up, and takes over), and cold switchover (the standby component needs to be fully started, and takes over after a delay). The switching mechanism is specifically designed for different hardware components: compute nodes utilize cluster redundancy and load balancing; network communication uses multi-path routing and protocol switching; and field devices are configured with functionally equivalent standby units. The switching process also includes priority management (ensuring critical functions are restored first) and smooth transition strategies (avoiding secondary problems caused by switching shocks). This mechanism is a key guarantee for the high availability of the system, enabling the control system to continue operating under various hardware failure conditions, thus ensuring the safety and stability of the wastewater treatment process.
[0251] Example 13
[0252] like Figure 3 As shown, the present invention also provides a hybrid intelligent optimization decision-making system for wastewater treatment plants, comprising:
[0253] Architecture design module 10 is used to acquire process flow data and control requirement data of wastewater treatment plant, and design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture.
[0254] Model building module 20 is used to collect historical operating data and real-time status data of each processing unit based on the hierarchical control architecture, perform data preprocessing and feature extraction, and build a global optimization model, which reflects the interaction relationship between each processing unit.
[0255] The centralized optimization module 30 is used to calculate the plant-wide resource optimization allocation scheme and control parameter constraints based on the global optimization model. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain a global optimization decision.
[0256] The edge control module 40 is used to deploy edge controllers in each processing unit, collect local sensor data based on the global optimization decision, run optimization algorithms within the range of the control parameter constraints, obtain real-time control commands at the unit level, and control the operation of the devices in each processing unit.
[0257] The communication and collaboration module 50 is used to collect the running status information after execution, and report it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer makes optimization decisions and adjustments for the next cycle, thereby realizing collaborative control between the centralized optimization layer and the distributed autonomous layer.
[0258] The specific implementation methods of the architecture design module, model building module, centralized optimization module, edge control module, and communication and collaboration module in this embodiment can be referred to the corresponding steps in the above method embodiments, and will not be repeated here.
[0259] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A hybrid intelligent optimization decision-making method for wastewater treatment plants, characterized in that, include: Obtain process flow data and control requirement data of the wastewater treatment plant, and design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture. Based on the hierarchical control architecture, historical operating data and real-time status data of each processing unit are collected and preprocessed and feature extracted to construct a global optimization model, which reflects the interaction relationship between each processing unit. Based on the global optimization model, the plant-wide resource optimization allocation scheme and control parameter constraints are calculated. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain the global optimization decision. Edge controllers are deployed in each treatment unit. Based on the global optimization decision, local sensor data is collected. Based on the local sensor data and the control parameter constraints, core treatment units affecting effluent quality are identified. These core treatment units include a biological reactor, a membrane filtration unit, and a disinfection unit, resulting in a set of key treatment units. According to the control problem of this set of key treatment units, the discretized control parameter values within the set interval of the control parameter constraints are modeled as a matroid structure. The basic element set is defined as the set of control parameter values discretized within the set interval of the control parameter constraints at a preset step size. An independent set family is defined as the set of parameter combinations that meet the process requirements, resulting in the wastewater treatment unit set. A process matte model is constructed based on the wastewater treatment process matte model. Matras corresponding to different optimization objectives are built, including energy efficiency matras, treatment efficiency matras, and stability matras. Through matte cross-operations, parameter combinations that simultaneously satisfy multiple constraints are found, resulting in a multi-objective matte structure. Based on the multi-objective matte structure, a parallel basis search algorithm is run to find a set of independent bases, each base corresponding to an optimization dimension, resulting in multiple parallel bases. Based on the multiple parallel bases and the computational resources of the edge controller, the computational resource allocation of each parallel base is dynamically adjusted according to the process state and control requirements, allocating more computational resources to the bases corresponding to key process parameters, thus obtaining an adaptive resource allocation scheme. According to the adaptive resource allocation scheme, an information sharing protocol and a co-evolution mechanism are implemented among parallel bases, enabling each base to learn from and improve each other during the iteration process, resulting in a co-optimized parallel base result. Based on the co-optimized parallel base result, each parallel base is assigned to different computing threads for parallel computation. An anomaly detection mechanism is used to identify and exclude base results that deviate too much. The contribution of each parallel base is evaluated using a Shapley value-based method, and the weight of each parallel base in the weighted voting is determined. The optimization results of each parallel base are integrated through a weighted voting-based fusion algorithm to obtain real-time control instructions at the unit level, and to control the operation of each processing unit. The system collects operational status information after execution and reports it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer then makes optimization decisions and adjustments for the next cycle, thereby achieving collaborative control between the centralized optimization layer and the distributed autonomous layer.
2. The method according to claim 1, characterized in that, The process flow data and control requirement data of the wastewater treatment plant are acquired. Based on this data, a two-layer control architecture is designed, comprising a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture, including: Acquire process flow data and control requirement data of the wastewater treatment plant. The process flow data includes the connection relationship of each treatment unit, the material flow path and energy consumption characteristics. The control requirement data includes effluent water quality standards, treatment capacity requirements and safe operation constraints, to obtain the basic data of the wastewater treatment plant. Based on the basic data of the wastewater treatment plant, the plant-wide resource coordination needs and unit-level real-time response needs are analyzed to determine the system's functional requirements. These functional requirements include global resource collaborative allocation needs, local rapid response control needs, and abnormal situation handling needs, resulting in a list of system functional requirements. Based on the system functional requirements list, a centralized optimization layer is designed to be responsible for plant-wide target setting and resource allocation, and a distributed autonomous layer is designed to be responsible for unit-level optimal autonomous control, resulting in a two-layer control architecture and responsibility boundary definition. Based on the aforementioned two-layer control architecture and responsibility boundary definition, the hardware configuration and installation locations of the central server, edge computing devices, and communication network are planned to obtain the hierarchical control architecture.
3. The method according to claim 1, characterized in that, The process of collecting historical operating data and real-time status data from each processing unit and performing data preprocessing and feature extraction includes: Historical operating data and real-time status data are collected from the sensors and control systems of each treatment unit in the wastewater treatment plant. The historical operating data and real-time status data include water quality parameters, equipment operating parameters and energy consumption data to obtain the raw dataset. The original dataset is subjected to missing value imputation, outlier detection, and data format standardization to obtain a cleaned dataset. The cleaned dataset is then subjected to sensitivity classification identification, and divided into high-sensitivity data, medium-sensitivity data, and low-sensitivity data to obtain the classified dataset. Based on the graded dataset, time features, frequency domain features, and statistical features are extracted to obtain preprocessed plant-wide data.
4. The method according to claim 3, characterized in that, The construction of a global optimization model, which reflects the interaction relationships between processing units, includes: Based on the preprocessed plant-wide data, a process network diagram is constructed to represent the material flow, energy flow, and information flow relationships between each processing unit. The process network diagram includes nodes and edges, with nodes representing processing units and edges representing the connection relationships between processing units, thus obtaining a wastewater treatment process network diagram. Based on the wastewater treatment process network diagram, a simplified activated sludge model and a membrane filtration model are used to establish a local dynamic model for each treatment unit, resulting in a simplified model for each unit. Based on the simplified models of each unit and the coupling relationships between units, model parameters are identified through data-driven machine learning methods, including neural networks or support vector machines, to obtain the global optimization model.
5. The method according to claim 1, characterized in that, The calculation of the plant-wide resource optimization allocation scheme and control parameter constraints, wherein the plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters, yields a global optimization decision, including: Based on the global optimization model, the influent water quality and influent water volume for future periods are predicted to obtain influent prediction data. Based on the influent prediction data and effluent water quality requirements, the energy consumption budget and chemical consumption budget of each treatment unit are calculated by an optimization algorithm. The optimization algorithm aims to minimize the total cost and obtains a resource budget allocation scheme for each unit. Based on the resource budget allocation scheme and process safety constraints of each unit, the reasonable setting range of each key control parameter is calculated. The key control parameters include dissolved oxygen concentration, reverse osmosis system recovery rate and reagent dosage, and the global optimization decision is obtained.
6. The method according to claim 4, characterized in that, The process of obtaining the global optimization model includes: Based on the wastewater treatment process network diagram, tree width analysis is performed on the wastewater treatment process network diagram to identify the key coupling points and bottlenecks of the system and obtain the tree width parameters of the process network. Based on the tree width parameter of the process network, the global optimization model is decomposed into several sub-models. Each sub-model corresponds to a cluster of tree decomposition. Tree decomposition ensures that the connection between sub-models follows a tree structure, resulting in a set of sub-models after tree decomposition. Based on the set of sub-models after tree decomposition, a hierarchical optimization model is constructed. The top-level node of the hierarchical optimization model corresponds to the macro-decision variables of the entire plant, the intermediate-level nodes correspond to the collaborative optimization variables of the processing unit group, and the leaf nodes correspond to the micro-control variables of a single processing unit. This results in a hierarchical optimization model based on a tree structure, which serves as the global optimization model.
7. The method according to claim 6, characterized in that, After obtaining the global optimization model, the process further includes: Based on the set of sub-models after tree decomposition and the wastewater treatment process network diagram, tree coverage metrics are calculated, including node coverage, edge coverage, and critical path retention rate, to obtain tree coverage evaluation results. Based on the tree coverage assessment results and computational resource constraints, the tree width parameter of the process network is dynamically adjusted. When the optimization period is shorter than the preset response time threshold, the tree width parameter value is reduced to reduce computational complexity. When the optimization period is longer than the preset response time threshold, the tree width parameter value is increased to improve model accuracy, thus obtaining the adaptively adjusted tree width parameter. Based on the adaptively adjusted tree width parameter, the step of decomposing the global optimization model into several sub-models is re-executed to obtain an updated hierarchical optimization model based on tree structure. The updated tree-based hierarchical optimization model is mapped to a centralized optimization layer and a distributed autonomous layer. The optimization of the top-level nodes and intermediate nodes of the updated tree-based hierarchical optimization model is implemented in the centralized optimization layer, and the optimization of the leaf nodes is implemented in the distributed autonomous layer, resulting in the mapped tree-based hierarchical optimization model, which serves as the global optimization model.
8. The method according to claim 3, characterized in that, The reporting to the centralized optimization layer via a two-way communication mechanism includes: Based on the running status information after execution and the hierarchical dataset, differential privacy sensitivity is calculated for each type of sensitive data. The differential privacy sensitivity represents the maximum impact of a change in a single record on the query results, thus obtaining the sensitivity parameters for each type of data. Based on the sensitivity parameters and privacy budget parameters of the various types of data, the standard deviation of Gaussian noise is calculated, and a time-adaptive noise calibration mechanism is designed to obtain the calibrated Gaussian noise parameters. Based on the running status information after execution and the calibrated Gaussian noise parameters, Gaussian noise is applied before the edge controller reports sensitive data. Different levels of privacy protection are applied to data with different sensitivity levels to obtain privacy-protected status reporting data. Based on the privacy-protected status reporting data, it is transmitted to the centralized optimization layer through the communication network to realize privacy-protected status reporting.
9. The method according to claim 8, characterized in that, Before reporting sensitive data to the edge controller, Gaussian noise is applied, and different levels of privacy protection are applied to data with different sensitivity levels. After obtaining the privacy-protected status reporting data, the process further includes: Based on the privacy-protected state reporting data, a noise-aware optimization objective function is run in the centralized optimization layer to incorporate data uncertainty into the decision-making process, thereby obtaining an optimization objective that takes noise into account. Based on the optimization objective that takes noise into account, a probabilistic inference method is used to estimate and filter out noise, calculate the decision confidence interval, and obtain a near-optimal decision result. Based on the near-optimal decision results, the decision quality loss index is calculated by comparing the difference in the optimization objective function values before and after privacy protection, and the decision error evaluation results are obtained. Based on the decision error assessment results, a quantitative relationship model between privacy protection strength and decision quality is established. The calibrated Gaussian noise parameters are dynamically adjusted according to the comparison between the decision quality loss index and the preset quality loss threshold to obtain an optimized decision that balances privacy protection and decision quality, which serves as the global optimized decision for the next cycle.
10. The method according to claim 1, characterized in that, It also includes fault tolerance and emergency response procedures, including: The communication status between the centralized optimization layer and the distributed autonomous layer is detected. When a communication interruption is detected, the autonomous operation mode of the distributed autonomous layer is started. Based on the historical operation data stored in the edge controller and the control parameter constraints received last time, the autonomous control command under the communication interruption is obtained. Based on the prediction deviation between the prediction results of the global optimization model and the actual measured values, the system detects whether the global optimization model has failed. When the prediction deviation exceeds a preset threshold, the system switches to a rule-based control mode to obtain a safety control command for model failure. The system monitors the operating status of hardware components, including a central server, edge controllers, sensors, and actuators. When a failure is detected in any of the hardware components, a backup system switching mechanism is activated, and detailed system operation logs are recorded to achieve fault-tolerant operation of the system.
11. A hybrid intelligent optimization decision-making system for wastewater treatment plants, characterized in that, include: The architecture design module is used to acquire process flow data and control requirement data of the wastewater treatment plant, and to design a two-layer control architecture based on the process flow data and control requirement data. The two-layer control architecture includes a centralized optimization layer and a distributed autonomous layer, resulting in a hierarchical control architecture. The model building module is used to collect historical operating data and real-time status data of each processing unit based on the hierarchical control architecture, perform data preprocessing and feature extraction, and build a global optimization model that reflects the interaction relationship between each processing unit. The centralized optimization module is used to calculate the plant-wide resource optimization allocation scheme and control parameter constraints based on the global optimization model. The plant-wide resource optimization allocation scheme allocates resource budgets to each processing unit, and the control parameter constraints determine the setting range of key control parameters to obtain the global optimization decision. The edge control module is used to deploy edge controllers in each processing unit. Based on the global optimization decision, it collects local sensor data and identifies the core processing units affecting the effluent quality based on the local sensor data and the control parameter constraints. The core processing units include a biological reactor, a membrane filtration unit, and a disinfection unit, resulting in a set of key processing units. According to the control problem of the set of key processing units, the discretized control parameter values within the set interval of the control parameter constraints are modeled as a matroid structure. The basic element set is defined as the set of control parameter values discretized within the set interval of the control parameter constraints at a preset step size. The independent set family is defined as the set of parameter combinations that meet the process requirements. A wastewater treatment process matroid model is obtained. Based on the wastewater treatment process matroid model, matroids corresponding to different optimization objectives are constructed, including energy efficiency matroid, treatment efficiency matroid, and stability matroid. Through matroid cross-operation, parameter combinations that simultaneously satisfy multiple constraints are found, resulting in a multi-objective matroid structure. Based on the multi-objective matroid structure, a parallel basis search algorithm is run to find a set of mutually independent bases, each base corresponding to an optimization dimension, resulting in multiple parallel bases. Based on the multiple parallel bases and the computing resources of the edge controller, the computing resource allocation of each parallel base is dynamically adjusted according to the process state and control requirements, allocating more computing resources to the bases corresponding to key process parameters, resulting in an adaptive resource allocation scheme. According to the adaptive resource allocation scheme, an information sharing protocol and a co-evolution mechanism are implemented among parallel bases, enabling each base to learn from and improve each other during the iteration process, resulting in a co-optimized parallel base result. Based on the co-optimized parallel base result, each parallel base is assigned to different computing threads for parallel computation. An anomaly detection mechanism is used to identify and exclude base results that deviate too much. The contribution of each parallel base is evaluated using a Shapley value-based method, and the weight of each parallel base in the weighted voting is determined. The optimization results of each parallel base are integrated through a weighted voting-based fusion algorithm to obtain real-time control instructions at the unit level, and to control the operation of each processing unit. The communication and collaboration module is used to collect the running status information after execution and report it to the centralized optimization layer through a two-way communication mechanism. The centralized optimization layer makes optimization decisions and adjustments for the next cycle, realizing collaborative control between the centralized optimization layer and the distributed autonomous layer.
Citation Information
Patent Citations
Multi-working-condition double-layer optimization control method for sewage treatment process based on task clustering
CN116881742A
Water-energy-medicine collaborative optimization method and system for sewage plant
CN122047656A