Water treatment system and method based on computer vision and multi-agent reinforcement learning

By employing computer vision and multi-agent reinforcement learning, the problems of measurement lag and multivariate coupling in the A2O wastewater treatment process were solved, enabling real-time perception of microbial status and multi-objective optimization control, thereby improving treatment efficiency and system stability.

CN121554086APending Publication Date: 2026-02-24HUADIAN WATER TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511685679.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing A2O wastewater treatment processes face challenges such as point-based measurement and spatial heterogeneity, sensor detection lag, lack of biological state perception, and strong coupling of multiple variables, making it difficult to achieve precise multi-objective collaborative optimization control.

Method used

A system based on computer vision and multi-agent reinforcement learning is adopted. The system collects images and sensor data through a multi-source perception module, performs semantic segmentation and morphological analysis using a feature extraction and fusion module, generates a fused state vector, and makes decisions by combining a multi-agent reinforcement learning model to adjust the aeration rate, internal recirculation ratio and external recirculation ratio.

Benefits of technology

It enables comprehensive and proactive perception of process status, significantly improves treatment efficiency, reduces energy consumption, and maintains long-term system stability while ensuring that the effluent meets standards, thereby enhancing system reliability and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121554086A_ABST
    Figure CN121554086A_ABST
Patent Text Reader

Abstract

The invention relates to a water treatment system and method based on computer vision and multi-agent reinforcement learning. The water treatment system based on computer vision and multi-agent reinforcement learning comprises a multi-source sensing module used for synchronously acquiring a mixed solution image sequence of an A2O process aerobic tank and sensor data of each process unit; the feature extraction and fusion module is used for performing semantic segmentation and morphological analysis on the image sequence, extracting microbial visual feature parameters, and fusing the visual feature parameters with sensor data to generate a fusion state vector; the intelligent decision-making module is a module based on a multi-agent reinforcement learning model and is used for outputting a combined control action vector according to the fusion state vector; and the execution control module is used for adjusting the aeration rate, the internal reflux ratio and the external reflux ratio of the A2O process according to the combined control action vector. According to the water treatment system disclosed by the invention, a complex biochemical process is realized, closed-loop control of multi-target collaborative optimization is carried out, the treatment efficiency is improved, and the energy consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for wastewater treatment, and more specifically, to a water treatment system and method based on computer vision and multi-agent reinforcement learning. Background Technology

[0002] A 2 Anaerobic-Anoxic-Oxic (ANOVA) is a secondary wastewater treatment process and a mainstream biological nitrogen and phosphorus removal technology. It achieves simultaneous nitrogen and phosphorus removal through three stages of biological reaction: anaerobic, anoxic, and aerobic. Its stable operation depends on precise control of the anaerobic, anoxic, and aerobic environments. Currently, ANOVA... 2 The control of the O process mainly relies on online sensors (such as dissolved oxygen (DO), oxidation-reduction potential (ORP), and sludge concentration (MLSS)) and traditional PID controllers or empirical rules.

[0003] However, this method has inherent drawbacks, such as: 1. The contradiction between point measurement and spatial heterogeneity: The sensor can only reflect local information at the probe installation point and cannot characterize the overall macroscopic state of the reaction tank, especially the distribution of the microbial community. 2. Control lag: By the time the sensor detects a change in water quality, the biochemical reaction has already occurred for some time, resulting in a control response lag. 3. Lack of biological state perception: It cannot obtain key microbial morphological information reflecting sludge health and activity (such as floc structure and filamentous bacteria abundance), which is an important basis for early warning of sludge bulking and process optimization. Existing manual microscopic observation methods are inefficient, highly subjective, and cannot provide real-time feedback. 4. Challenges of strong coupling of multiple variables: A 2 In the O process, variables such as aeration, internal reflux, and external reflux affect each other, making it difficult to achieve coordinated optimization using traditional single-loop control.

[0004] However, how to deeply integrate unstructured visual information with structured sensor data and build a closed-loop control system that can understand complex biochemical processes and perform multi-objective collaborative optimization remains a technical challenge that urgently needs to be solved in this field. Summary of the Invention

[0005] The technical problem this invention aims to solve is how to deeply integrate unstructured visual information with structured sensor data to construct a closed-loop control system capable of understanding complex biochemical processes and performing multi-objective collaborative optimization, addressing the challenges in wastewater treatment systems such as the contradiction between point measurement and spatial heterogeneity, the lag in sensor detection of water quality changes, the lack of biological state perception, and the strong coupling of multiple variables.

[0006] To address the aforementioned technical problems, according to one aspect of the present invention, a water treatment system based on computer vision and multi-agent reinforcement learning is provided, comprising: a multi-source sensing module for synchronously acquiring A...2 The system includes: a mixed liquor image sequence from the aerobic tank of the O process and sensor data from each process unit; a feature extraction and fusion module for semantic segmentation and morphological analysis of the image sequence, extracting microbial visual feature parameters, and fusing these parameters with sensor data to generate a fused state vector; an intelligent decision-making module based on a multi-agent reinforcement learning model for outputting a joint control action vector based on the fused state vector; and an execution control module for adjusting A based on the joint control action vector. 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

[0007] According to an embodiment of the present invention, in the feature extraction and fusion module, the visual feature parameters may include the average diameter d of the flocs. avg Filament abundance index FSVI and floc activity index AI; fused state vector S t It is composed of visual feature parameters and sensor data.

[0008] According to an embodiment of the present invention, the average diameter d of the flocs avg The calculation formula can be:

[0009] ,

[0010] Where, d avg Where A is the average diameter of the flocs, N is the number of flocs, and A is the average diameter of the flocs. i Let i be the area of ​​the i-th floc.

[0011] The formula for calculating the abundance index of filamentous fungi is:

[0012] ,

[0013] Wherein, FSVI is the abundance index of filamentous bacteria, L fil A represents the total length of filamentous bacteria. floc This represents the total area of ​​the flocculent body.

[0014] According to an embodiment of the present invention, the multi-agent reinforcement learning model in the intelligent decision-making module can employ the MADDPG algorithm, wherein the multi-agent reinforcement learning model includes three agents: aeration control, internal reflux control, and external reflux control; the reward function R t Includes a sludge state reward term R driven by visual feature parameters. sludge .

[0015] According to an embodiment of the present invention, the reward function R t Specifically, it could be:

[0016] ,

[0017] in, to R represents the weighting coefficients. effluen As a reward for the quality of the effluent, R energy As an energy reward, R sludge As a reward for sludge status, R stable As a stability bonus.

[0018] According to another aspect of the present invention, a water treatment method based on computer vision and multi-agent reinforcement learning is provided, implemented based on the water treatment system based on computer vision and multi-agent reinforcement learning as described above. The water treatment method based on computer vision and multi-agent reinforcement learning includes the following steps:

[0019] S1, Synchronous Acquisition A 2 Image sequence of mixed liquor in the aerobic tank of the O process and sensor data of each process unit;

[0020] S2. Perform semantic segmentation and morphological analysis on the image sequence to extract microbial visual feature parameters, and fuse the visual feature parameters with sensor data to generate a fused state vector S. t ;

[0021] S3, merge the state vector S t The input is fed into a pre-trained multi-agent reinforcement learning model to obtain the joint control action vector A. t ;

[0022] S4. Based on the joint control action vector A t Adjust A 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

[0023] According to an embodiment of the present invention, in step S2, a U-Net-based deep learning model can be used for semantic segmentation, and the loss function of the deep learning model is a combination of Dice Loss and Focal Loss.

[0024] ,

[0025] in, .

[0026] According to an embodiment of the present invention, in step S3, the multi-agent reinforcement learning model can be trained by a reward function that includes a visual feature reward term to collaboratively optimize effluent quality, energy consumption, and sludge health status.

[0027] According to an embodiment of the present invention, in step S4, control instructions can be executed by the PLC, and the instruction values ​​can be safely limited.

[0028] According to an embodiment of the present invention, during deployment, it may first be in A 2The multi-agent reinforcement learning model is trained in the process simulation environment, and then run in the real system first in the suggestion mode, and then switched to the closed-loop control mode.

[0029] Compared with the prior art, the technical solution provided by the embodiments of the present invention can achieve at least the following beneficial effects:

[0030] According to the present invention, a water treatment system and method based on computer vision and multi-agent reinforcement learning are deployed on A 2 An industrial camera in the O-reactor captured a sequence of images of the mixed liquid. A deep learning semantic segmentation model was used to extract microbial morphological features, including the average diameter d of the flocs. avg The visual features, including the filamentous bacteria abundance index (FSVI) and the floc activity index (AI), are spatiotemporally aligned and feature-fused with multi-source data such as dissolved oxygen (DO), oxidation-reduction potential (ORP), and mixed liquor suspended solids concentration (MLSS) measured by traditional sensors, to construct a comprehensive state vector S. t A multi-agent decision-making model is constructed based on reinforcement learning algorithms, with S... t To input the optimized control commands that generate aeration volume ΔDO, internal recirculation ratio ΔRi, and external recirculation ratio ΔRe in real time, A is achieved. 2 Closed-loop optimized operation of the O process.

[0031] This invention effectively solves the problems of traditional control methods that rely on point measurement, have slow response, and cannot detect the state of microorganisms, thus significantly improving processing efficiency and reducing energy consumption.

[0032] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can achieve comprehensive and advanced perception of process status, directly acquire microbial morphological information through computer vision, overcome the limitations of point sensors, and provide early warning of process anomalies such as sludge bulking earlier than traditional methods.

[0033] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can achieve multi-objective collaborative optimization control. By designing a multi-objective reward function that includes water quality, energy consumption and sludge health, and by utilizing multi-agent reinforcement learning, collaborative optimization of strongly coupled multi-variables is achieved. Under the premise of ensuring that the effluent meets the standards, energy consumption is significantly reduced and the long-term stability of the system is maintained.

[0034] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can form a data-driven intelligent closed loop, constructing a complete data flow closed loop from "visual perception" to "intelligent decision-making" and then to "precise execution". The system has self-learning and self-adaptive capabilities, reducing the dependence on human experience.

[0035] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can improve the reliability and safety of the system. Through hard safety constraints and mode switching logic at the PLC level, the process safety of the intelligent control system under abnormal conditions is ensured. Attached Figure Description

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of the present invention and are not intended to limit the present invention.

[0037] Figure 1 This is a flowchart illustrating a water treatment system and method based on computer vision and multi-agent reinforcement learning according to an embodiment of the present invention;

[0038] Figure 2 This is a diagram illustrating the overall architecture and data flow of a water treatment system based on computer vision and multi-agent reinforcement learning according to an embodiment of the present invention.

[0039] Figure 3 This is a detailed flowchart illustrating visual feature extraction according to an embodiment of the present invention;

[0040] Figure 4 This is a structural diagram illustrating a multi-agent reinforcement learning decision-making model according to an embodiment of the present invention;

[0041] Figure 5 This illustrates a camera according to an embodiment of the present invention in A. 2 Schematic diagram of the deployment in the aerobic tank. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” and similar terms used in the specification and claims of this patent application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a limitation of quantity, but rather indicate the presence of at least one.

[0044] Figure 1 This is a flowchart illustrating a water treatment system and method based on computer vision and multi-agent reinforcement learning according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the overall architecture and data flow of a water treatment system based on computer vision and multi-agent reinforcement learning according to an embodiment of the present invention.

[0045] like Figure 1 and Figure 2 As shown, the water treatment system based on computer vision and multi-agent reinforcement learning includes: a multi-source perception module, a feature extraction and fusion module, an intelligent decision-making module, and an execution control module.

[0046] The multi-source sensing module is used to synchronously collect A 2 Image sequence of mixed liquor in the aerobic tank of the O process and sensor data of each process unit.

[0047] The feature extraction and fusion module is used to perform semantic segmentation and morphological analysis on image sequences, extract visual feature parameters of microorganisms, and fuse the visual feature parameters with sensor data to generate a fused state vector.

[0048] The intelligent decision-making module is based on a multi-agent reinforcement learning model and is used to output a joint control action vector based on the fused state vector.

[0049] The execution control module is used to adjust A according to the joint control action vector. 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

[0050] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can achieve comprehensive and advanced perception of process status, directly acquire microbial morphological information through computer vision, overcome the limitations of point sensors, and provide early warning of process anomalies such as sludge bulking earlier than traditional methods.

[0051] The multi-source sensing module includes a visual acquisition unit and a sensor acquisition unit.

[0052] The visual acquisition unit in A 2 Waterproof industrial cameras and matching light sources are deployed at the beginning, middle, and end of the aerobic tank in the O process for continuously acquiring image sequences of the mixed liquor. t .

[0053] The sensor acquisition unit is used to synchronously acquire A 2 Online sensor data X from each unit of the O process t Including the oxidation-reduction potential (ORP) of the anaerobic tank. anaerobic Dissolved oxygen (DO) in the anoxic tank anoxic and redox potential (ORP) anoxicDissolved oxygen (DO) and mixed liquor suspended solids (MLSS) in the aerobic tank, and influent flow rate (Q) in and effluent ammonia nitrogen (NH4) + -Neff and nitrate nitrogen NO3 - -Neff.

[0054] Visual monitoring point deployment optimized. Due to the gradient distribution of microbial community structure in different sections of the aerobic tank, IP68 waterproof industrial cameras (such as Basler ace 2) are installed at the beginning, middle, and end of the aerobic tank. These cameras utilize a 520nm green LED array light source to enhance the contrast with the brown color of the activated sludge. They are installed at a 45° angle to the camera lens and equipped with an automatic cleaning device (compressed air blowing + mechanical scraper) to ensure stable image quality. Camera data acquisition resolution: 2048×1536 pixels, sampling frequency: 1 frame / second, image format: RAW→JPEG compression (quality factor 85%).

[0055] Construction of multi-source sensor networks. A 2 Each unit of the O process requires a different redox environment. In the anaerobic stage, ORP reflects phosphorus release conditions, while in the anoxic stage, nitrate concentration affects denitrification. Therefore, an ORP sensor (±5mV accuracy) and a pH sensor are installed in the anaerobic tank; a trace DO sensor (0-0.5mg / L range) and an ORP sensor are installed in the anoxic tank; and an optical DO sensor (0-20mg / L range) and a laser MLSS sensor are installed in the aerobic tank. Online sensors are also installed at the influent and effluent. , COD analyzer; data acquisition interval is 10 seconds / time, communication protocol is Modbus TCP / IP, and data storage is a time series database (InfluxDB).

[0056] Spatial synchronization and data preprocessing. Multi-source data fusion requires strict time alignment. Furthermore, microbial morphological changes and water quality parameter changes are temporally correlated. Therefore, the following technical measures can be taken: deploy an NTP time synchronization server (accuracy ±1ms), configure an industrial-grade switch (supporting the IEEE 1588 precision clock protocol), and set up an edge computing gateway for data caching; use a unified timestamp format during data acquisition: YYYY-MM-DD HH:MM:SS.sss; perform sensor data interpolation based on the image acquisition time; and remove obviously abnormal data based on the 3σ criterion for outlier detection.

[0057] Data quality verification and calibration. Sensor drift can affect long-term monitoring accuracy, and image quality directly impacts the reliability of feature extraction. Corresponding equipment and technical measures include: weekly on-site sensor calibration; daily parallel water sample collection for laboratory comparison; setting up an image quality assessment module (resolution, contrast, and brightness detection); storing calibration records in a metadata database during data acquisition; image quality scoring (0-100 points); triggering an alarm if the score is below 80 points; establishing a sensor drift compensation model; and generating data quality flags.

[0058] Through these steps, S1 not only completes the data acquisition task but also provides high-quality, standardized, and multimodal input data for feature extraction and fusion in S2, ensuring that the entire system possesses scientific rigor, accuracy, and reliability from the data source. The standardized image sequence It output by S1 is directly used as input to the U-Net model in S2, and the aligned sensor data Xt output by S1 is directly input to the data fusion submodule of S2. The data quality flags generated by S1 are used for confidence evaluation of feature extraction in S2.

[0059] Figure 3 This is a detailed flowchart illustrating visual feature extraction according to an embodiment of the present invention.

[0060] like Figure 3 As shown, in the feature extraction and fusion module, the visual feature parameters include the average diameter d of the flocs. avg Filament abundance index FSVI and floc activity index AI; fused state vector S t It is composed of visual feature parameters and sensor data.

[0061] According to one or more embodiments of the present invention, the average diameter d of the flocs avg The calculation formula is:

[0062] ,

[0063] Where, d avg Where A is the average diameter of the flocs, N is the number of flocs, and A is the average diameter of the flocs. i Let i be the area of ​​the i-th floc.

[0064] The formula for calculating the abundance index of filamentous fungi is:

[0065] ,

[0066] Wherein, FSVI is the abundance index of filamentous bacteria, L fil A represents the total length of filamentous bacteria. floc This represents the total area of ​​the flocculent body.

[0067] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can achieve multi-objective collaborative optimization control. By designing a multi-objective reward function that includes water quality, energy consumption and sludge health, and by utilizing multi-agent reinforcement learning, collaborative optimization of strongly coupled multi-variables is achieved. Under the premise of ensuring that the effluent meets the standards, energy consumption is significantly reduced and the long-term stability of the system is maintained.

[0068] The feature extraction and fusion module includes: performing semantic segmentation and morphological analysis on image sequences, extracting visual feature parameters of microorganisms, and fusing the visual feature parameters with sensor data to generate a fused state vector S. t .

[0069] Image preprocessing and quality enhancement. The equipment and technical measures include: utilizing a GPU cluster (NVIDIA A100) for parallel image processing and accelerating image processing algorithms using the OpenCV library; specifically, Gaussian filtering effectively suppresses high-frequency noise while preserving microbial edge information; the CLAHE algorithm improves uneven lighting in underwater images; and the Lab color space better aligns with human visual perception, facilitating feature extraction. Data processing includes: Gaussian filtering for noise reduction using a 5×5 kernel size and σ=1.5; histogram equalization using the CLAHE algorithm with clipLimit=2.0 and tileGridSize=8×8; color space conversion from RGB to Lab; and contrast enhancement.

[0070] Semantic segmentation based on U-Net. The encoder-decoder structure of U-Net is suitable for biomedical image segmentation, where the ResNet50 backbone provides powerful feature extraction capabilities, and multi-class segmentation accurately reflects the microbial community structure. The equipment technical measures are as follows: the encoder of the pre-trained U-Net model adopts ResNet50, the decoder is transposed convolution, the model is deployed on the TensorRT inference engine, and FP16 is used for accuracy acceleration; in terms of data processing: the input image size is adjusted to 512×512 pixels, the pixel-level classification output is background (0), dense bacterial flocs (1), loose flocs (2), filamentous bacteria (3), and the output segmentation confidence map is set with a threshold of 0.75.

[0071] Morphological feature quantitative analysis. Using the scikit-image library for morphological operations and calling the CUDA-accelerated connected component analysis algorithm, the analysis revealed that floc size distribution reflects sludge settling performance, filamentous bacteria length is positively correlated with sludge bulking risk, and floc edge sharpness is correlated with microbial activity. Data processing included: floc analysis; connected component labeling (8-neighborhood connectivity); area filtering (removing noise points with less than 50 pixels); and calculation of equivalent diameter (d). avg,tFilamentous fungal analysis: Skeletonization: Zhang-Suen parallel thinning algorithm; Length calculation: number of skeleton pixels × calibration coefficient; Abundance index: FSVI t Bacterial floc activity analysis: Edge detection: Canny operator, threshold (50, 150), texture features: Local Binary Pattern (LBP) entropy calculation, activity index: AI = 0.6 × edge sharpness + 0.4 × texture complexity.

[0072] Spatiotemporal alignment of multi-source data. Apache Spark Structured Streaming is used for streaming data processing to address the time delay between changes in microbial morphology and water quality parameters. A time-series database is deployed for data caching, and appropriate time windows are used to ensure the accuracy of data correlation. The specific data processing steps are as follows: Time window alignment: linear interpolation of sensor data within a 5-second window before and after the image timestamp; Data resampling: standardizing to a 1-minute time series interval; Missing value handling: linear interpolation for completion, with more than 3 consecutive missing points marked as anomalies.

[0073] Feature standardization and fusion. Feature standardization is performed using scikit-learn's StandardScaler, and high-dimensional vector concatenation is done using NumPy to eliminate the influence of different scales, ensuring fair feature weights. Multimodal feature fusion provides a more comprehensive description of the process state. Data processing is as follows: Z-score standardization:

[0074] ,

[0075] in, The characteristic mean, The standard deviation is denoted as .

[0076] Feature vector construction:

[0077] Feature dimensions: 9-dimensional vector, including 3 visual features + 6 sensor features.

[0078] Data quality verification and output. A feature rationality check module is set up to prevent abnormal data from affecting decision-making, implement a real-time data quality monitoring dashboard, and ensure system reliability. Data processing: Range check: davg∈[10,500]μm, FSVI∈[0,10], AI∈[0,1]; Consistency verification: Logical consistency between visual features and sensor data; Output format: JSON format, including timestamp, feature vector, and quality flag.

[0079] Through this detailed design process, S2 successfully transforms raw image and sensor data into feature vectors with clear physical and biological meanings, providing high-quality, interpretable input features for subsequent intelligent decision-making.

[0080] The visual feature extraction submodule uses a deep learning model based on the U-Net architecture to perform semantic segmentation on image It, identify bacterial flocs, flocs and filamentous bacteria, and calculate the following morphological feature parameters: average diameter of flocs, filamentous bacteria abundance index FSVI, and model loss function.

[0081] The data fusion submodule performs time alignment and normalization processing on the visual feature parameters and the sensor data, and then concatenates them into a fused state vector: .

[0082] S2 outputs the normalized fusion state vector S t The feature quality labels generated by S2 are directly used as input states for the S3 reinforcement learning model, and are used for the decision confidence evaluation of S3.

[0083] Figure 4 This is a structural diagram illustrating a multi-agent reinforcement learning decision-making model according to an embodiment of the present invention.

[0084] like Figure 4 As shown, the multi-agent reinforcement learning model in the intelligent decision-making module adopts the MADDPG algorithm. This model includes three agents: aeration control, internal reflux control, and external reflux control. The reward function R... t Includes a sludge state reward term R driven by visual feature parameters. sludge .

[0085] According to one or more embodiments of the present invention, the reward function R t Specifically:

[0086] ,

[0087] in, to R represents the weighting coefficients. effluen As a reward for the quality of the effluent, R energy As an energy reward, R sludge As a reward for sludge status, R stable As a stability bonus.

[0088] Action space:

[0089]

[0090] Where At is the joint control action vector, and △DO setpointΔRi is the adjustment amount for the dissolved oxygen setpoint, ΔRe is the adjustment amount for the internal reflux ratio, and ΔRe is the adjustment amount for the external reflux ratio.

[0091] Intelligent decision-making and optimal control will integrate the state vector S t The input is fed into a pre-trained multi-agent reinforcement learning model to obtain the joint control action vector A. t .

[0092] State vector preprocessing and validation. A microservice for input data validation is deployed to perform data integrity checks, ensuring no missing values ​​in the 9-dimensional feature vector. A real-time data anomaly detection module is set up to perform time-series continuity validation, detect data jumps, and filter abnormal fluctuations. Simultaneously, S... t Are the values ​​of each dimension within a reasonable range: DO: [0,5] mg / L, ORP: [-400,100] mV, MLSS: [1000,6000] mg / L

[0093] d avg :[10,500]μm, FSVI:[0,10], AI:[0,1].

[0094] Multi-agent policy network inference. Deep neural networks can learn complex nonlinear control policies, and the multi-agent architecture effectively solves the coupling problem between control variables. It loads three pre-trained Actor networks (PyTorch models), deploys a model inference service (TensorRT acceleration), and optimizes GPU memory for batch processing with dynamic memory allocation. Its network architecture has an input layer of 9 neurons, corresponding to S... t The structure has 9 dimensions, with 2 fully connected hidden layers, each with 256 neurons and ReLU activation, and 3 output neurons corresponding to ΔDO and ΔR. i ΔR e The following reasoning is performed: State standardization: using the mean and variance of the training set; Forward propagation: S t →Actor network→Original action output; Action post-processing: Tanh activation function mapped to the range [-1,1].

[0095] Motion space mapping and constraint processing. Motion scaling maps the neural network output to the actual engineering range, safety constraints prevent control commands from exceeding equipment capabilities, and motion smoothing ensures stable process operation. The motion scaling and constraint processing module implements this, setting physical limiting logic for the actuators. Data processing is as follows: Motion scaling: ΔDOsetpoint = original output × 0.3 mg / L, ΔRi = original output × 10%, ΔRe = original output × 10%; Safety constraints: DOsetpoint ∈ [0.5, 4.0] mg / L, Ri ∈ [100%, 300%], Re ∈ [50%, 100%]; Motion smoothing: First-order low-pass filtering avoids drastic fluctuations in setpoints, with a rate of change limit of |ΔDO| ≤ 0.5 mg / L / minute.

[0096] Real-time multi-objective reward calculation. By deploying a reward calculation microservice, a dynamic weight adjustment mechanism for multiple objectives is implemented, allowing reinforcement learning to guide the agent to learn the optimal policy through reward signals. Data processing includes: reward function calculation: ; Itemized Rewards: Based on the effluent quality meeting standards, R energy =-k×Power blower Energy consumption penalty term, R sludge =-penalty×I(FSVI>3.0) Sludge state penalty, R stable =-λ×||At-At-1||² Stability bonus; Weight adaptive: The weight coefficient is dynamically adjusted according to the influent load, and priority is given to ensuring that the effluent meets the standards under abnormal operating conditions.

[0097] Online learning and model updates. To break down data correlations through experience replay and improve learning efficiency, the target network enhances learning stability. Online learning enables the system to adapt to process changes. An experience replay buffer (Redis cluster) is deployed, an online model update service is set up, and an A / B testing framework is implemented. Data processing is as follows: Experience storage: storing a quadruple: (St, At, Rt, St+1), buffer capacity: 100,000 experiences, priority sampling: based on TD error; Model updates:

[0098] Sampling batch size: 256, learning rate: Actor1e-4, Critic1e-3, soft update coefficient τ: 0.01, update frequency: once every 1000 steps.

[0099] Decision confidence assessment and anomaly handling. The decision quality assessment module evaluates the reliability of decisions and sets up a backup control strategy switching mechanism to ensure continuous and stable process operation. Data processing includes: confidence assessment: Critic network value estimation variance, state space coverage detection, and action exploration degree assessment; anomaly handling: switching to PID control when confidence is low, activating the expert rule base when the model is abnormal, and retaining the last valid command in case of communication failure.

[0100] Through the detailed design steps outlined above, S3 successfully achieved the intelligent transformation from process status perception to optimized control decision-making, providing a foundation for the entire A... 2 The O process provides an adaptive, multi-objective, safe, and reliable intelligent control core. The joint control action vector A output by S3... t =[ΔDO setpoint , ΔR i , ΔR e The decision confidence flag generated by S3 is sent directly to the PLC control system of S4 and used for the safety interlocking logic of S4.

[0101] Control execution and system feedback: Adjust A according to the joint control action vector At. 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

[0102] Control command reception and parsing. An industrial communication gateway (supporting OPC UA / MQTT protocols) is deployed according to industrial communication protocols to ensure reliable data transmission. The PLC is configured with a dedicated data receiving function block, and command verification and redundant transmission mechanisms are set up. Command verification prevents the execution of erroneous commands, and incremental control avoids sudden jumps in setpoints. Data processing includes: Protocol parsing: parsing JSON-formatted control command A. t = [ΔDO setpoint , ΔR i , ΔR e Data verification: CRC check, serial number check, timestamp freshness verification (<2 seconds); Command conversion: convert percentage change to actual set value; DO setpointnew =DO setpointcurrent + ΔDO setpoint ;R inew = R icurrent + ΔR i ;R enew = R ecurrent + ΔR e .

[0103] Actuator safety limiting and interlocking. Safety constraints protect equipment from damage, rate-of-change limits maintain process stability, interlocking logic prevents misoperation, the PLC is configured with hard safety logic (independent of program loops), hardware limit switches and overload protection are set, and safety relays are installed to achieve emergency stop. Data processing is as follows: Parameter hard limit: DO setpoint ∈[0.5,4.0]mg / L, R i ∈[100%, 300%], R e ∈[50%, 100%]; Rate of change limit: |ΔDO setpoint | ≤ 0.3 mg / L / dose, |ΔR i | ≤ 15% / time, |ΔR e | ≤ 10% / time; Equipment interlock logic: restrict sludge discharge when MLSS is low (<1500mg / L), prioritize adjustment of the return pump when the water level is high, and automatically switch to standby when equipment fails.

[0104] Precise control of the aeration system. Aeration intensity design is guided by oxygen transfer efficiency theory, and the dissolved oxygen (DO) setpoint is determined by microbial oxygen consumption kinetics. Variable frequency control achieves precise aeration and energy saving. A blower variable frequency drive (ABB ACS880 series) is used, along with a fine bubble aeration system and dissolved oxygen cascade PID control. Data processing: Variable frequency control: based on DO... setpoint Calculate blower frequency, frequency output range: 25-50 Hz, dead zone compensation: ±0.1 mg / L; aeration intensity adjustment: aeration volume based on MLSS and temperature compensation, zone control: independent adjustment of the first, middle and last sections of the aerobic tank, over-aeration prevention logic: automatic frequency reduction when DO>3.0mg / L; energy consumption optimization: real-time power monitoring, optimal efficiency operating point tracking, vibration isolation zone control.

[0105] The recirculation system is coordinated and controlled. The recirculation flow rate is calculated based on the mass balance principle, external recirculation adjustment is guided by sludge age theory, and internal recirculation demand is determined by denitrification kinetics. Both the internal and external recirculation pumps are controlled by variable frequency drives (VFDs), and flow meters and level gauges provide real-time monitoring. Data processing is as follows: Internal recirculation control: Flow rate setting: Q ir = R i × Q in Denitrification efficiency monitoring: based on ORP and NO3⁻-N in the anoxic tank; Carbon source optimization: dynamically adjusting R based on the C / N ratio. i External reflux control: Flow rate setting: Q er = R e × Q in Sludge age control: based on MLSS and sludge discharge volume; Secondary sedimentation tank stability: maintain appropriate sludge layer height; Coordination strategy: internal recirculation prioritizes denitrification, external recirculation maintains biomass balance, and in case of conflict, prioritize ensuring effluent meets standards.

[0106] Execution status monitoring and feedback. Based on the dynamic performance control theory of the evaluation system, the predictive maintenance theory of equipment failure, and the statistical process control of the monitoring system stability, equipment status sensors (vibration, temperature, current), actuator position feedback devices, and a real-time data acquisition system are set up. Data processing includes: execution verification: setpoint-actual value deviation monitoring, response time statistics, and execution success rate calculation; equipment health diagnosis: motor current trend analysis, bearing vibration spectrum monitoring, and seal leakage detection; performance evaluation: control accuracy: DO control error < ±0.2 mg / L, response speed: setpoint tracking lag < 30 seconds, stability index: overshoot < 10%.

[0107] System closed-loop and adaptive adjustment. Based on closed-loop control theory, system stability is ensured; adaptive control responds to changes in operating conditions; performance evaluation guides continuous optimization; a feedback data acquisition network is deployed; and a control performance self-evaluation module is set up to achieve adaptive parameter tuning. Data processing includes: feedback data collection: process parameters: real-time values ​​of DO, ORP, and MLSS; water quality indicators: , TP change trend, visual features: davg, FSVI, AI update data; control effect evaluation: effluent compliance rate statistics, energy efficiency calculation, sludge status score; parameter adaptation: automatic adjustment of control parameters based on operating conditions, activation of enhanced control during load shocks, and long-term performance degradation compensation.

[0108] Through this detailed design process, S4 successfully transforms intelligent decision-making into precise physical control actions, and achieves closed-loop optimization and continuous improvement of the entire system through a robust feedback mechanism. The execution effect of S4 is monitored in real time by the sensor network and vision system of S1. The control action results of S4 serve as new process state inputs to S1's data acquisition. After the execution of S4's control commands, changes in process parameters are collected and verified by S1, and the microbial morphological changes provided by S1 provide feedback on the long-term control effect of S4. The execution data of S4 provides new training experience for the reinforcement learning of S3, and S1-S4 form a complete "perception-decision-execution-learning" intelligent closed loop.

[0109] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can form a data-driven intelligent closed loop, constructing a complete data flow closed loop from "visual perception" to "intelligent decision-making" and then to "precise execution". The system has self-learning and self-adaptive capabilities, reducing the dependence on human experience.

[0110] According to another aspect of the present invention, a water treatment method based on computer vision and multi-agent reinforcement learning is provided, implemented based on the water treatment system based on computer vision and multi-agent reinforcement learning as described above. The water treatment method based on computer vision and multi-agent reinforcement learning includes the following steps:

[0111] S1, Synchronous Acquisition A 2 Image sequence of mixed liquor in the aerobic tank of the O process and sensor data of each process unit;

[0112] S2. Perform semantic segmentation and morphological analysis on the image sequence to extract microbial visual feature parameters, and fuse the visual feature parameters with sensor data to generate a fused state vector S. t ;

[0113] S3, merge the state vector S t The input is fed into a pre-trained multi-agent reinforcement learning model to obtain the joint control action vector A. t ;

[0114] S4. Based on the joint control action vector A t Adjust A 2 Aeration rate ΔDO and internal reflux ratio ΔR of the O process i and external reflux ratio ΔR e .

[0115] According to the present invention, a water treatment system and method based on computer vision and multi-agent reinforcement learning are deployed on A 2 An industrial camera in the O-reactor captured a sequence of images of the mixed liquid. A deep learning semantic segmentation model was used to extract microbial morphological features, including the average diameter d of the flocs. avg The visual features, including the filamentous bacteria abundance index (FSVI) and the floc activity index (AI), are spatiotemporally aligned and feature-fused with multi-source data such as dissolved oxygen (DO), oxidation-reduction potential (ORP), and mixed liquor suspended solids concentration (MLSS) measured by traditional sensors, to construct a comprehensive state vector S. t A multi-agent decision-making model is constructed based on reinforcement learning algorithms, with S... t To input the optimized control commands that generate aeration volume ΔDO, internal recirculation ratio ΔRi, and external recirculation ratio ΔRe in real time, A is achieved. 2 Closed-loop optimized operation of the O process.

[0116] According to one or more embodiments of the present invention, in step S2, a U-Net-based deep learning model is used for semantic segmentation, and the loss function of the deep learning model is a combination of Dice Loss and Focal Loss.

[0117] ,

[0118] in, .

[0119] According to one or more embodiments of the present invention, in step S3, the multi-agent reinforcement learning model is trained by a reward function that includes a visual feature reward term in order to collaboratively optimize effluent quality, energy consumption and sludge health status.

[0120] According to one or more embodiments of the present invention, in step S4, control instructions are executed by the PLC, and the instruction values ​​are subject to safety limits.

[0121] The water treatment system and method based on computer vision and multi-agent reinforcement learning according to the present invention can improve the reliability and safety of the system. Through hard safety constraints and mode switching logic at the PLC level, the process safety of the intelligent control system under abnormal conditions is ensured.

[0122] Figure 5 This illustrates a camera according to an embodiment of the present invention in A. 2 Schematic diagram of the deployment in the aerobic tank.

[0123] like Figure 5 As shown, during deployment, first in A 2 The multi-agent reinforcement learning model is trained in the process simulation environment, and then run in the real system first in the suggestion mode, and then switched to the closed-loop control mode.

[0124] The above description is merely an exemplary embodiment of the present invention and is not intended to limit the scope of protection of the present invention, which is determined by the appended claims.

Claims

1. A water treatment system based on computer vision and multi-agent reinforcement learning, comprising: Multi-source sensing module, used for synchronous acquisition of A 2 Image sequence of mixed liquor in the aerobic tank of the O process and sensor data of each process unit; The feature extraction and fusion module is used to perform semantic segmentation and morphological analysis on image sequences, extract visual feature parameters of microorganisms, and fuse the visual feature parameters with sensor data to generate a fused state vector. The intelligent decision-making module is a module based on a multi-agent reinforcement learning model, used to output a joint control action vector according to the fused state vector; The execution control module is used to adjust A according to the joint control action vector. 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

2. The water treatment system based on computer vision and multi-agent reinforcement learning as described in claim 1, wherein, In the feature extraction and fusion module, the visual feature parameters include the average diameter d of the flocs. avg Filament abundance index FSVI and floc activity index AI; the fused state vector S t It is formed by splicing the visual feature parameters and the sensor data.

3. The water treatment system based on computer vision and multi-agent reinforcement learning as described in claim 2, wherein, The average diameter d of the flocs avg The calculation formula is: , Where, d avg Where A is the average diameter of the flocs, N is the number of flocs, and A is the average diameter of the flocs. i Let i be the area of ​​the i-th floc. The formula for calculating the abundance index of filamentous bacteria is as follows: , Wherein, FSVI is the abundance index of filamentous bacteria, L fil A represents the total length of filamentous bacteria. floc This represents the total area of ​​the flocculent body.

4. The water treatment system based on computer vision and multi-agent reinforcement learning as described in claim 1, wherein, The multi-agent reinforcement learning model in the intelligent decision-making module adopts the MADDPG algorithm. This model includes three agents: aeration control, internal reflux control, and external reflux control. The reward function R... t Includes a sludge state reward term R driven by the aforementioned visual feature parameters. sludge .

5. The water treatment system based on computer vision and multi-agent reinforcement learning as described in claim 4, wherein, The reward function R t Specifically: , in, to R represents the weighting coefficients. effluen As a reward for the quality of the effluent, R energy As an energy reward, R sludge As a reward for sludge status, R stable As a stability bonus.

6. A water treatment method based on computer vision and multi-agent reinforcement learning, implemented based on the water treatment system based on computer vision and multi-agent reinforcement learning as described in any one of claims 1 to 5, the water treatment method based on computer vision and multi-agent reinforcement learning comprising the following steps: S1, Synchronous Acquisition A 2 Image sequence of mixed liquor in the aerobic tank of the O process and sensor data of each process unit; S2. Perform semantic segmentation and morphological analysis on the image sequence, extract microbial visual feature parameters, and fuse the visual feature parameters with the sensor data to generate a fused state vector S. t ; S3, convert the fusion state vector S t The input is fed into a pre-trained multi-agent reinforcement learning model to obtain the joint control action vector A. t ; S4. According to the joint control action vector A t Adjust A 2 The aeration rate, internal reflux ratio, and external reflux ratio of the O process.

7. The water treatment method based on computer vision and multi-agent reinforcement learning as described in claim 6, wherein, In step S2, a U-Net-based deep learning model is used for semantic segmentation. The loss function of the deep learning model is a combination of DiceLoss and Focal Loss. , in, .

8. The water treatment method based on computer vision and multi-agent reinforcement learning as described in claim 6, wherein, In step S3, the multi-agent reinforcement learning model is trained using a reward function that includes visual feature reward terms to collaboratively optimize effluent quality, energy consumption, and sludge health.

9. The water treatment method based on computer vision and multi-agent reinforcement learning as described in claim 6, wherein, In step S4, control instructions are executed by the PLC, and the instruction values ​​are subject to safety limits.

10. The water treatment method based on computer vision and multi-agent reinforcement learning as described in claim 6, wherein, During deployment, first in A 2 The multi-agent reinforcement learning model is trained in the process simulation environment, and then run in the real system first in the suggestion mode, and then switched to the closed-loop control mode.

Citation Information

Cited By

  • Sludge age self-adaptive regulation and control method and system based on dissolved oxygen trend recognition

    CN121948678A