Sewage treatment data management method and system based on artificial intelligence

By adopting artificial intelligence-based data management methods in sewage treatment systems, using technologies such as multi-source data fusion network and Bayesian network, the problems of lag and high cost of traditional sewage treatment systems are solved, and intelligent management and efficiency improvement of system abnormalities are achieved.

CN120234741AActive Publication Date: 2025-07-01CHINA NAT INST OF STANDARDIZATION

Patent Information

Application Number
CN202510703093.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional sewage treatment systems rely on manual monitoring, and have problems such as lagging reactions, high cost and strong subjectivity in judgment, making it difficult to effectively manage abnormal working conditions of sewage treatment systems.

Method used

Using artificial intelligence-based sewage treatment data management method, data is collected in real time through distributed sensor networks, key features are extracted using multi-source data fusion networks, and combined with Bayesian networks and random forest models to realize intelligent management of abnormal working conditions and generation of regulatory strategies.

Benefits of technology

It realizes intelligent management of sewage treatment system abnormalities throughout the process, improves operation stability and treatment efficiency, and reduces energy consumption and operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234741A_ABST
    Figure CN120234741A_ABST
Patent Text Reader

Abstract

The invention discloses a sewage treatment data management method and system based on artificial intelligence, and relates to the technical field, and the method comprises the steps: obtaining real-time operation data of a sewage treatment system, extracting key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network, inputting the key features into a pre-trained state evaluation model to generate a sewage treatment state; according to the sewage treatment state, judging whether the current operation state is in an abnormal working condition, if the current operation state is in the abnormal working condition, identifying an abnormal reason based on the real-time operation data and the sewage treatment state through a diagnosis model based on a Bayesian network, and determining whether the current operation state is in the abnormal working condition based on the abnormal reason. Inputting the key features into a strategy generation model based on a random forest, and generating a targeted regulation and control strategy; the technical problem of data processing is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to a sewage treatment data management method and system based on artificial intelligence. Background Art

[0002] Sewage treatment is an important link in ensuring water resource security and ecological environment protection. In modern sewage treatment systems, various process parameters such as pH value, dissolved oxygen, chemical oxygen demand (COD), biochemical oxygen demand (BOD), ammonia nitrogen content, etc. need to be monitored and controlled in real time. The sewage treatment process is complex and changeable, affected by various factors such as fluctuations in influent water quality, changes in environmental temperature, and changes in microbial activity, and is prone to abnormal working conditions. If these abnormalities cannot be detected and corresponding measures are not taken in time, it will lead to a decrease in treatment efficiency, energy waste, and even may cause the treated water quality to fail to meet the discharge standards, causing secondary pollution to the environment.

[0003] Traditional sewage treatment systems mainly rely on manual experience for monitoring and adjustment, and have problems such as lagging response, high labor costs, and strong subjectivity in judgment. With the development of Internet of Things technology and sensor technology, a large amount of real-time data has been generated in sewage treatment systems. These data contain rich information on the operating status of the system, and there is an urgent need for an intelligent data management method that can automatically monitor abnormal situations, analyze the causes of abnormalities, and provide control strategies in a timely manner to ensure the stable and efficient operation of sewage treatment systems.

[0004] Therefore, the present invention proposes a sewage treatment data management method and system based on artificial intelligence. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention proposes a sewage treatment data management method and system based on artificial intelligence, realizing the full-process intelligent management of abnormalities in sewage treatment systems.

[0006] To achieve the above object, a sewage treatment data management method based on artificial intelligence is proposed, including the following steps: Step 1: Obtain the real-time operation data of the sewage treatment system; Step 2: Extract the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network; Step 3: Input the key features into a pre-trained state evaluation model to generate the sewage treatment state; according to the sewage treatment state, judge whether the current operating state is in an abnormal working condition; Step 4: If in an abnormal working condition, identify the cause of the abnormality through a diagnostic model based on a Bayesian network, based on the real-time operation data and the sewage treatment state; Step 5: Based on the abnormal reason, input the key features into a policy generation model based on random forest. This model classifies different abnormal types by integrating multiple decision trees and generates targeted regulation policies in combination with historical processing experience and current working condition parameters; The acquisition of the real-time operation data of the sewage treatment system includes the following steps: Step 11: Collect the real-time parameter data of each process unit of the sewage treatment system through a distributed sensor network; Step 12: Structurally process the collected real-time parameter data through a field control unit adopting an edge computing architecture to form structured real-time operation data; The method for extracting the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network includes the following steps: Step 21: Collect the historical operation data of the sewage treatment system and construct a multi-source heterogeneous data training sample library; Step 22: Construct a multi-source data fusion network model. The multi-source data fusion network model adopts a multi-stream architecture design, including a feature extraction module, a time series modeling module, a multi-source fusion module, and a feature output module; The structure of the multi-source data fusion network model includes: In the feature extraction module, for the data stream of water quality parameters, a one-dimensional convolutional neural network structure is adopted, which includes 3 consecutive one-dimensional convolutional blocks. Each convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation function, and a max pooling layer; For the data stream of operation parameters, a multi-scale residual network structure is adopted, which includes 4 residual blocks. Each residual block contains two one-dimensional convolutional layers and a skip connection to capture the device operation characteristics at different time scales; For the data stream of process control parameters, a gated recurrent unit (GRU) network is adopted, which includes a 2-layer bidirectional GRU structure, and the number of hidden units in each layer is 64. The GRU network can effectively model the long-term dependence relationship and time series change trend of process control parameters, and the bidirectional structure considers both historical and future information at the same time, improving the comprehensiveness of feature extraction.

[0007] In the time series modeling module, the time attention mechanism is applied to the features extracted from the three data streams respectively to learn the importance weights of the features at different time points. That is, the time attention mechanism generates attention weights by calculating the similarity between the query vector and the features at each time point, and then performs weighted summation on the features at each time point to obtain a feature representation with time importance; In the multi-source fusion module, an adaptive feature fusion strategy is adopted, which includes three steps: feature alignment, feature interaction, and feature fusion. In the feature alignment step, the features output by the three data streams through the temporal modeling module are mapped to the same feature space through a fully connected layer; in the feature interaction step, through the cross-stream attention mechanism, the mutual influence between the corresponding feature spaces of different data streams is calculated; in the feature fusion step, through the gated fusion unit, according to the correlation of the feature spaces output by each data stream feature after the feature interaction step, the fusion weights are adaptively determined to generate a comprehensive feature representation.

[0008] In the feature output module, the comprehensive feature representation is reduced in dimension and feature selected through a two-layer fully connected network to extract the most representative key features.

[0009] Step 23: Based on the constructed multi-source heterogeneous data training sample library, adopt an end-to-end supervised learning method to train the multi-source data fusion network model and optimize the network parameters; Step 24: Use the trained multi-source data fusion network to process real-time operation data and extract the key features of the sewage treatment system; Inputting the key features into the pre-trained state evaluation model to generate the sewage treatment state includes the following steps: Step 31: Synchronously construct a training sample library for the state evaluation model with the multi-source heterogeneous data training sample library; the training sample library contains two types of data: normal working condition data and abnormal working condition data; Step 32: Design a state evaluation model with a network structure including a feature extraction layer, a temporal modeling layer, and a state classification layer; The structure of the state evaluation model includes: In the feature extraction layer, a multi-layer perceptron network is used to perform non-linear transformation and dimensionality reduction on the input key features. The feature extraction layer contains three fully connected layers, and the activation function uses the ReLU function to enhance the expression ability of the state evaluation model for non-linear relationships. A Dropout layer is added between the fully connected layers to prevent the model from overfitting.

[0010] In the temporal modeling layer, a bidirectional long short-term memory network is used to capture the time-dependent relationship of the sewage treatment system parameters. The temporal modeling layer contains two bidirectional long short-term memory networks; In the state classification layer, the output features of the temporal modeling layer pass through two fully connected layers and then output the state classification result of the sewage treatment system through the Softmax function.

[0011] Step 33: Based on the constructed training sample library, adopt a supervised learning method to train the state evaluation model with the goal of optimizing state classification; Step 34: The state evaluation model receives the key features from Step 2 and generates a state classification result of the sewage treatment state in real time.

[0012] The steps for identifying the root cause of anomalies through a diagnostic model based on a Bayesian network, based on real-time operation data and sewage treatment states, include the following: Step 41: Based on the historical operation data of the sewage treatment system and the recorded abnormal events annotated by experts, construct a second training sample library of the Bayesian network diagnostic model that contains the causal relationships between abnormal types, causes, and symptoms; Step 42: Based on the second training sample library, adopt a method that combines structure learning and expert knowledge to construct a three-layer hierarchical Bayesian network structure that includes a root cause node layer, an intermediate state node layer, and an observed feature node layer; The Bayesian network structure adopts a hierarchical design and includes three layers of nodes: The first layer is the root cause node layer, which represents the fundamental abnormal causes with probabilities greater than a preset probability threshold, including the influent water quality abnormal node group, the equipment failure node group, and the process parameter abnormal node group; The second layer is the intermediate state node layer, which represents the system state changes caused by the root causes, including the biochemical reaction state node group, the hydraulic state node group, and the solid-liquid separation state node group; The third layer is the observed feature node layer, which represents the abnormal features that can be observed through the sensor network, including the water quality feature node group, the equipment operation feature node group, and the process feature node group.

[0013] Step 43: Based on the annotated data in the second training sample library, adopt the maximum likelihood estimation and expectation maximization algorithms to learn the conditional probability table of the causal relationships between the nodes in the Bayesian network, and handle the data sparsity problem through the Bayesian estimation method; Step 44: Based on the real-time operation data as the observed evidence, adopt the joint tree algorithm for probability inference, calculate the posterior probability distribution of each root cause node, and output the abnormal cause with the highest probability; The steps for integrating multiple decision trees to classify different abnormal types and generating targeted regulation strategies by combining historical treatment experience and current working condition parameters include the following: Step 51: Construct a third training sample library of the policy generation model that contains historical abnormal events and their successful treatment experiences; Step 52: Construct a policy generation model based on a random forest that includes a feature preprocessing module, an abnormal classification module, and a policy generation module; In the feature preprocessing module, the input key feature vector and the abnormal cause are combined and processed to form the input features for policy generation; In the abnormal classification module, a random forest classifier is used to perform fine-grained abnormal type recognition on the features output by the feature preprocessing module; In the policy generation module, for the fine-grained abnormal types identified by the abnormal classification module, a random forest regressor is used to generate specific regulation parameter values; Step 53: Based on the third training sample library, use the ensemble learning method to train the random forest policy generation model; Step 54: Input the abnormal cause and key features into the trained random forest policy generation model to generate targeted regulation policies; A sewage treatment data management system based on artificial intelligence is proposed, including a real-time data collection module, a key feature extraction module, an abnormal working condition judgment module, an abnormal cause identification module, and a policy generation module; among them, each module is connected electrically; The real-time data collection module obtains the real-time operation data of the sewage treatment system and sends the real-time operation data to the key feature extraction module; The key feature extraction module extracts the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network and sends the key features to the abnormal working condition judgment module; The abnormal working condition judgment module inputs the key features into a pre-trained state evaluation model to generate the sewage treatment state; according to the sewage treatment state, it judges whether the current operation state is in an abnormal working condition. If it is judged to be an abnormal working condition, it sends the real-time operation data and the sewage treatment state to the abnormal cause identification module; The abnormal cause identification module uses a diagnostic model based on a Bayesian network to identify the abnormal cause based on the real-time operation data and the sewage treatment state, and sends the abnormal cause and the real-time operation data to the policy generation module; The policy generation module, based on the abnormal cause, inputs the key features into a random forest-based policy generation model. This model classifies different abnormal types by integrating multiple decision trees and combines historical processing experience and current working condition parameters to generate targeted regulation policies.

[0014] Compared with the prior art, the beneficial effects of the present invention are: Through the sensor network distributed in each process link of sewage treatment, multi-source data of water quality parameters are collected in real time to form the real-time operation data of the sewage treatment system. Then, the pre-trained multi-source data fusion network is used to process the real-time operation data, extract and fuse features of data from different sources and different dimensions, and screen out the key features that can reflect the system operation state, solving the problems of high data dimension and information redundancy in traditional methods, improving the efficiency and accuracy of subsequent analysis. The extracted key features are input into the pre-trained state evaluation model to evaluate the operation state of the current sewage treatment system and determine whether it is in an abnormal working condition. If it is judged to be in an abnormal working condition, the diagnosis model based on the Bayesian network is started. Based on the identified abnormal reasons, the system inputs the key features into the strategy generation model based on the random forest. By integrating multiple decision trees, refined classification of different types of abnormalities is carried out, and combined with historical treatment experience and current working condition parameters, targeted regulation strategies are generated. A closed-loop intelligent management system is formed, realizing the full-process intelligent management of sewage treatment system abnormalities from data collection, feature extraction, state evaluation, abnormal diagnosis to strategy generation, significantly improving the operation stability and treatment efficiency of the sewage treatment system, and reducing energy consumption and operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of the sewage treatment data management method based on artificial intelligence in Embodiment 1 of the present invention; Figure 2 It is a module connection relationship diagram of the sewage treatment data management system based on artificial intelligence in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Embodiment 1

[0017] As Figure 1 shown, the sewage treatment data management method based on artificial intelligence includes the following steps: Step 1: Obtain the real-time operation data of the sewage treatment system; Step 2: Extract the key features of the sewage treatment system from the real-time operation data through the pre-trained multi-source data fusion network; Step 3: Input the key features into the pre-trained state evaluation model to generate the sewage treatment state; according to the sewage treatment state, judge whether the current operation state is in an abnormal working condition; Step 4: If in an abnormal operating condition, identify the cause of the abnormality through a diagnostic model based on a Bayesian network, based on real-time operation data and sewage treatment status; Step 5: Based on the cause of the abnormality, input the key features into a policy generation model based on a random forest. This model classifies different types of abnormalities by integrating multiple decision trees and generates targeted regulation strategies in combination with historical treatment experience and current operating condition parameters; In an embodiment of the present invention, the obtaining of the real-time operation data of the sewage treatment system includes the following steps: Step 11: Collect real-time parameter data of each process unit of the sewage treatment system through a distributed sensor network; Specifically, the distributed sensor network is composed of various types of sensors, including water quality parameter sensors, equipment operation status sensors, and process control parameter sensors.

[0018] Among them, water quality parameter sensors are arranged at the inlet, each sewage treatment unit, and the outlet of the sewage treatment system to continuously monitor key water quality indicators such as pH value, dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen, and suspended solid concentration; the sewage treatment unit refers to a set of sewage treatment equipment and water quality parameter monitoring equipment used for real-time treatment of sewage between the inlet and the outlet; Among them, the equipment operation status sensors are installed on core equipment such as aeration equipment, stirring equipment, and pumping stations to monitor operation parameters such as the running current, rotation speed, vibration frequency, and temperature of the equipment; Among them, the process control parameter sensors monitor process control parameters such as the hydraulic load, sludge concentration, reflux ratio, and aeration volume of each sewage treatment unit.

[0019] In the specific implementation process of the present invention, during the data collection process, each sensor continuously collects parameter data according to a preset sampling frequency. The sampling frequency of the water quality parameter sensors is set to once every 5 minutes to ensure timely capture of the water quality change trend; the sampling frequency of the equipment operation status sensors is once every 30 seconds to monitor the equipment operation status in real time; the sampling frequency of the process control parameter sensors is once every 1 minute to ensure real-time adjustment of the process parameters.

[0020] Each time the collected key water quality indicators, the operation parameters, and the process control parameters form real-time parameter data; Step 12: Structurally process the collected real-time parameter data through a field control unit adopting an edge computing architecture to form structured real-time operation data; Specifically, a field control unit is configured for each sewage treatment unit, which is responsible for receiving data from all water quality parameter sensors, equipment operation status sensors, and process control parameter sensors within the sewage treatment unit, and performing signal conversion, data verification, and outlier detection.

[0021] During the signal conversion process, the analog or digital signals output by the water quality parameter sensors, equipment operation status sensors, and process control parameter sensors are uniformly converted into a standard signal format. In the data verification stage, a redundant sensor cross-verification method is adopted. By comparing the readings of multiple sensors for the same parameter, possible sensor failures or abnormal readings are identified and marked. Outlier detection uses an anomaly detection algorithm based on statistical methods to identify data points outside the normal range and mark them according to the degree of anomaly.

[0022] After structured processing, the field control unit organizes the data into structured real-time operation data packets, which at least include the following information: timestamp, accurate to the second level; parameter identifier, which uniquely identifies each monitoring parameter; parameter value, the actual measured value after verification; quality label, indicating the reliability level of the data; equipment identifier, identifying the specific equipment or location where the data source is located.

[0023] Furthermore, the method of extracting the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network includes the following steps: Step 21: Collect historical operation data of the sewage treatment system and construct a multi-source heterogeneous data training sample library; Specifically, in this embodiment, during the process of collecting the historical operation data, the operation data of at least 12 consecutive months is extracted from the background data management system of the sewage treatment plant in the same way as in step one, so as to ensure that the samples cover the treatment conditions under different seasons, different weather conditions, and different influent water qualities. That is, the collected data includes three types of heterogeneous data: water quality parameter data, equipment operation status data, and process control parameter data in the historical sewage treatment process.

[0024] In the embodiment of the present invention, the multi-source heterogeneous data training sample library is organized according to the time series. Each sample contains the operation data within a time window. Generally, the length of the time window is set to 2 hours, and the sliding step is 10 minutes. For each time window, the water quality parameters, equipment operation status, and process control parameters within the time window are extracted to form a multi-source heterogeneous data sample. At the same time, the treatment effect indicators corresponding to this time window are manually recorded, including the compliance of the effluent water quality, energy consumption level, and treatment efficiency, etc., as sample labels.

[0025] Step 22: Construct a multi-source data fusion network model. The multi-source data fusion network model is designed with a multi-stream architecture, including a feature extraction module, a temporal modeling module, a multi-source fusion module, and a feature output module; Specifically, the multi-source data fusion network model adopts an overall architecture of "multi-stream - fusion - output". According to the characteristics of three types of heterogeneous data, namely water quality parameters, operation parameters, and process control parameters in the sewage treatment system, three parallel feature extraction streams are constructed to process different types of data respectively. Then, the features extracted by each stream are integrated through the fusion module, and finally the key features of the sewage treatment system are output.

[0026] In the feature extraction module, for the data stream of water quality parameters, a one-dimensional convolutional neural network structure is adopted, which includes 3 consecutive one-dimensional convolutional blocks. Each convolutional block consists of a one-dimensional convolutional layer with a kernel size of 5, a batch normalization layer, a ReLU activation function, and a max pooling layer. The number of output channels of the first convolutional block is 32, and the number of channels of each subsequent block doubles. This structure can effectively capture the spatial correlation and short-term change patterns among water quality parameters.

[0027] For the data stream of operation parameters, a multi-scale residual network structure is adopted, which includes 4 residual blocks. Each residual block contains two one-dimensional convolutional layers and a skip connection, and the kernel sizes are 3 and 5 respectively to capture the device operation characteristics at different time scales. The residual connection helps to alleviate the gradient vanishing problem of the deep network to improve the training stability of the model. This structure is suitable for processing the high-frequency fluctuations and mutation characteristics in the device operation parameter data.

[0028] For the data stream of process control parameters, a gated recurrent unit (GRU) network is adopted, which includes a 2-layer bidirectional GRU structure, and the number of hidden units in each layer is 64. The GRU network can effectively model the long-term dependence relationship and temporal change trend of process control parameters, and the bidirectional structure considers both historical and future information at the same time, improving the comprehensiveness of feature extraction.

[0029] In the time series modeling module, the time attention mechanism is applied to the features extracted from the three data streams respectively to learn the importance weights of the features at different time points. That is, the time attention mechanism generates attention weights by calculating the similarity between the query vector and the features at each time point, and then performs weighted summation on the features at each time point to obtain a feature representation with time importance. This mechanism enables the model to adaptively focus on the key time points in the sewage treatment process, such as the moment of water quality mutation, equipment start-stop, or process parameter adjustment. Among them, the time point feature refers to the data state representation at a specific moment during the operation of the sewage treatment system. Specifically, for the water quality parameter data stream, the time point feature represents the numerical states of water quality indicators such as pH value, dissolved oxygen, and COD at that moment; for the equipment operation state data stream, the time point feature represents the states of operation parameters such as equipment current, rotation speed, and vibration frequency at that moment; for the process control parameter data stream, the time point feature represents the states of process parameters such as hydraulic load, sludge concentration, and reflux ratio at that moment.

[0030] In the multi-source fusion module, an adaptive feature fusion strategy is adopted, which includes three steps: feature alignment, feature interaction, and feature fusion. In the feature alignment step, the features output by the three data streams through the time series modeling module are mapped to the same feature space through a fully connected layer; in the feature interaction step, through the cross-stream attention mechanism, the mutual influence between the corresponding feature spaces of different data streams is calculated, such as the influence of equipment operation state on water quality parameters, or the influence of process control parameters on equipment operation; in the feature fusion step, through the gated fusion unit, according to the correlation of the feature spaces output by the features of each data stream after the feature interaction step, the fusion weights are adaptively determined to generate a comprehensive feature representation.

[0031] In the feature output module, the comprehensive feature representation is reduced in dimension and feature selection is performed through a two-layer fully connected network to extract the most representative key features. The output dimension of the first fully connected layer is 128, and the ReLU activation function is used; the output dimension of the second fully connected layer is 64, corresponding to the final key feature dimension.

[0032] Step 23: Based on the constructed multi-source heterogeneous data training sample library, adopt an end-to-end supervised learning method to train the multi-source data fusion network model and optimize the network parameters; Specifically, the training of the multi-source data fusion network model aims to optimize the sewage treatment effect. The training process is divided into the following steps: First, the multi-source heterogeneous data training sample library is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set is used for learning model parameters, the validation set is used for model selection and hyperparameter adjustment, and the test set is used for evaluating the performance of the final model.

[0033] In the design of the loss function, a multi-task learning framework is adopted to optimize multiple objectives simultaneously. The main loss functions include: the loss for effluent water quality prediction, which calculates the difference between the predicted value and the actual value using the mean squared error (MSE); the loss for treatment efficiency prediction, which also uses MSE for calculation; and the loss for abnormal condition detection, which calculates the difference between the predicted abnormal probability and the actual label using binary cross-entropy. The three loss functions are combined by weighted summation, and the weights are set according to the importance of each task.

[0034] The optimization algorithm of the multi-source data fusion network model uses the Adam optimizer. The initial learning rate is set to 0.001, and a learning rate decay strategy is used, multiplying the learning rate by 0.9 every 10 training epochs. The batch size is set to 64, and the number of training epochs is 100. During the training process, an early stopping strategy is adopted. When the loss on the validation set has not improved for 5 consecutive epochs, the training process is stopped to avoid overfitting. At the same time, the model parameters with the best performance on the validation set are saved as the final model.

[0035] Step 24: Use the trained multi-source data fusion network to process the real-time operation data and extract the key features of the sewage treatment system; Specifically, in the actual application stage, the structured real-time operation data obtained in Step 1 is input into the trained multi-source data fusion network. Through the feature extraction module, time series modeling module, multi-source fusion module, and feature output module in Step 22, a 64-dimensional key feature vector is generated.

[0036] The key feature vector contains the core information of the current operating state of the sewage treatment system. In the embodiments of the present invention, it specifically includes: a water quality feature sub-vector (20-dimensional), which characterizes the key characteristics and changing trends of the influent and effluent water quality; an equipment feature sub-vector (20-dimensional), which characterizes the operating state and health of the core equipment; a process feature sub-vector (20-dimensional), which characterizes the configuration of process parameters and the regulation effect; and a system overall feature sub-vector (4-dimensional), which characterizes the overall operating efficiency and stability of the sewage treatment system.

[0037] It should be noted that, in view of the multi-source heterogeneous characteristics of the data in the sewage treatment system, the multi-source data fusion network designs a feature extraction architecture with multiple streams in parallel. Each data stream adopts a neural network structure suitable for the characteristics of this type of data, such as a one-dimensional convolutional network for water quality parameters, a residual network for equipment status, and a GRU network for process parameters, realizing efficient feature extraction for different types of data. Secondly, the multi-source data fusion network introduces a temporal attention mechanism and a cross-stream attention mechanism to adaptively focus on critical moments in the time series and the mutual influence between different data sources, so as to be able to capture the dynamic change process of sewage treatment. Thirdly, an adaptive feature fusion strategy is adopted, and the fusion weights of different data sources are dynamically adjusted through a gated fusion unit, enabling the multi-source data fusion network to flexibly integrate multi-source information according to the characteristics of the current working conditions, improving the pertinence and accuracy of feature extraction. The multi-source data fusion network is trained through a multi-task learning framework, simultaneously optimizing multiple objectives such as effluent water quality prediction, treatment efficiency prediction, and abnormal working condition detection, so that the extracted key features have stronger comprehensive representation ability and can support subsequent tasks such as state assessment, abnormal diagnosis, and strategy generation.

[0038] Further, inputting the key features into a pre-trained state assessment model to generate the sewage treatment state includes the following steps: Step 31: Synchronously construct a training sample library for the state assessment model with the multi-source heterogeneous data training sample library; the training sample library contains two types of data: normal working condition data and abnormal working condition data; Specifically, in a preferred embodiment, the construction process of the training sample library for the state assessment model is carried out synchronously with the collection of the multi-source heterogeneous data training sample library of the multi-source data fusion network in step two.

[0039] In the stage of collecting normal working condition data, select the data of several consecutive days from the operation data collected during the stable operation of the sewage treatment system as the basic sample set. Each basic sample record includes a timestamp, water quality parameters of each sewage treatment unit, equipment operation status parameters, process control parameters, and corresponding treatment effect indicators (such as the compliance of the effluent water quality).

[0040] In the stage of collecting abnormal working condition data, extract the marked abnormal event data from the historical operation data as the abnormal working condition samples, and mark the degree of abnormality of the abnormal working condition samples, such as "normal operation" and "abnormal", to record all parameter changes during the system response process.

[0041] In a preferred embodiment, for each sample of abnormal operating conditions, it can also be annotated by domain experts to determine the degree of abnormality and the scope of influence; the abnormal conditions are divided into sudden changes in influent water quality, equipment failures, process parameter deviations, etc.; the scope of influence is marked with the affected process units and possible chain reactions.

[0042] Step 32: Design a state evaluation model with a network structure including a feature extraction layer, a time series modeling layer, and a state classification layer; Specifically, in the feature extraction layer, a multi-layer perceptron network is used to perform non-linear transformation and dimensionality reduction processing on the input key features. The feature extraction layer contains three fully connected layers, each layer containing 128, 64, and 32 neurons respectively. The activation function uses the ReLU function to enhance the state evaluation model's ability to express non-linear relationships. A Dropout layer (dropout rate is 0.3) is added between the fully connected layers to prevent the model from overfitting.

[0043] In the time series modeling layer, a bidirectional long short-term memory network (Bi-LSTM) is used to capture the time-dependent relationships of the sewage treatment system parameters. The time series modeling layer contains two bidirectional long short-term memory networks, each layer containing 64 LSTM units, which can consider the information of past and future time points simultaneously to more comprehensively understand the parameter change trend. The output of the Bi-LSTM layer is processed by an attention mechanism to adaptively assign weights to the features at different time points and highlight the state changes at critical moments.

[0044] In the state classification layer, the output features of the time series modeling layer pass through two fully connected layers (containing 32 and 16 neurons respectively), and then the state classification results of the sewage treatment system are output through the Softmax function. In this embodiment, the state classification results include two categories: "normal operation" and "abnormal".

[0045] Step 33: Based on the constructed training sample library, adopt the supervised learning method to train the state evaluation model with the goal of optimizing state classification; Specifically, first, the training sample library is divided into a training set and a validation set according to a ratio of 8:2.

[0046] In the model training stage, the mini-batch gradient descent algorithm is adopted, the batch size is set to 64, the initial value of the learning rate is 0.001, and a learning rate decay strategy is used. The learning rate is reduced to 0.9 times the original value every 50 epochs. An early stopping strategy is used during the training process. When the performance index on the validation set has not improved for 10 consecutive epochs, the training process is stopped to prevent overfitting.

[0047] The loss function of the model is the state classification loss, and the state classification loss uses the cross-entropy loss function to measure the difference between the predicted state and the true state; Step 34: The state evaluation model receives the key features from Step 2 and generates a state classification result of the sewage treatment state in real time.

[0048] Further, the method for judging whether the current operating state is in an abnormal condition according to the sewage treatment state is as follows: If the sewage treatment state is expressed as "abnormal", it is judged that the current sewage treatment is in an abnormal condition; Further, the method for identifying the cause of the abnormality based on the diagnostic model based on the Bayesian network, based on the real-time operation data and the sewage treatment state, includes the following steps: Step 41: Based on the historical operation data of the sewage treatment system and the abnormal event records annotated by experts, construct a second training sample library of the Bayesian network diagnostic model including the causal relationships of abnormal types, causes, and symptoms; Specifically, in the embodiment of the present invention, in the construction stage of the second training sample library, all confirmed abnormal event data are screened out from the historical operation data of the sewage treatment system, and the following information is synchronously supplemented for each abnormal condition sample: the time period when the abnormality occurs, the real-time operation data before and after the abnormality occurs, the type and root cause of the abnormality, the propagation path of the abnormality, the scope and degree of the impact of the abnormality, the measures taken to handle the abnormality and their effects. It can be understood that different from the training sample library of the state evaluation model, the second training sample library pays more attention to the causal chain record of abnormal events to ensure that the propagation law of abnormalities in the system can be accurately captured.

[0049] In the preferred embodiment of the present invention, the abnormal types can be divided into three categories according to the source: abnormal influent water quality, abnormal equipment operation, and abnormal process parameters. Abnormal influent water quality includes, but is not limited to: sudden increase in influent COD, ammonia nitrogen load shock, abnormal pH value, toxic substance shock, etc.; abnormal equipment operation includes, but is not limited to: aeration equipment failure, agitation equipment failure, reflux pump failure, instrument failure, etc.; abnormal process parameters include, but is not limited to: insufficient dissolved oxygen, abnormal sludge concentration, hydraulic load fluctuation, improper reflux ratio, etc.

[0050] For each abnormal event data in the second training sample library, an expert team composed of sewage treatment process experts and equipment experts makes annotations to determine the causal relationship chain among the root cause, secondary cause, and manifestation symptoms of the abnormality. For example, the failure of the aeration equipment (root cause) may lead to insufficient dissolved oxygen (secondary cause), which in turn causes a decrease in the activity of nitrifying bacteria (intermediate state), and finally leads to exceeding the standard of ammonia nitrogen in the effluent (manifestation symptom). These detailed causal chain annotations provide precise prior knowledge for the Bayesian network structure learning, which complements the training sample library of the state evaluation model that only focuses on the judgment of abnormal states.

[0051] Step 42: Based on the second training sample library, adopt a method combining structure learning and expert knowledge to construct a three-layer hierarchical Bayesian network structure including a root cause node layer, an intermediate state node layer, and an observed feature node layer; Specifically, the Bayesian network structure adopts a hierarchical design and includes three layers of nodes: The first layer is the root cause node layer, which represents the fundamental abnormal causes with probabilities greater than a preset probability threshold, including an influent water quality anomaly node group, an equipment failure node group, and a process parameter anomaly node group. Among them, the influent water quality anomaly node group includes nodes such as sudden increase in influent COD, ammonia nitrogen load shock, and abnormal pH value; the equipment failure node group includes nodes such as aeration equipment failure, agitation equipment failure, and reflux pump failure; the process parameter anomaly node group includes nodes such as abnormal dissolved oxygen control, abnormal sludge concentration, and abnormal hydraulic load.

[0052] The second layer is the intermediate state node layer, which represents the system state changes caused by the root causes, including a biochemical reaction state node group, a hydraulic state node group, and a solid-liquid separation state node group. Among them, the biochemical reaction state node group includes nodes such as nitrification reaction inhibition, decreased denitrification efficiency, and blocked organic matter degradation; the hydraulic state node group includes nodes such as abnormal hydraulic retention time and uneven mixing; the solid-liquid separation state node group includes nodes such as decreased sludge settling performance and sludge floating.

[0053] The third layer is the observed feature node layer, which represents the abnormal features that can be observed through the sensor network, including a water quality feature node group, an equipment operation feature node group, and a process feature node group. Among them, the water quality feature node group includes nodes such as abnormal effluent COD, abnormal effluent ammonia nitrogen, and abnormal effluent total phosphorus; the equipment operation feature node group includes nodes such as abnormal equipment current, abnormal equipment vibration, and abnormal equipment temperature; the process feature node group includes nodes such as abnormal dissolved oxygen concentration, abnormal MLSS, and abnormal SVI.

[0054] In the process of constructing the Bayesian network structure, a method combining structure learning and expert knowledge is adopted. First, based on the prior knowledge provided by domain experts, determine the initial structure of the causal relationship between nodes; then, use a scoring-based structure learning algorithm, such as the K2 algorithm or the Bayesian Information Criterion (BIC) algorithm, to learn and optimize the network structure from the second training sample library. During the structure learning process, keep the strong causal relationships defined by experts unchanged and only optimize and adjust the weak causal relationships.

[0055] Step 43: Based on the labeled data in the second training sample library, adopt the maximum likelihood estimation and expectation maximization algorithm to learn the conditional probability table of the causal relationship between each node in the Bayesian network, and handle the data sparsity problem through the Bayesian estimation method; Specifically, after determining the Bayesian network structure, based on the labeled data in the second training sample library, it is necessary to learn the conditional probability table of each node, that is, the probability distribution of the current node in each possible state given the state of the parent node.

[0056] For discrete nodes, the maximum likelihood estimation method is used to statistically calculate the conditional probability from the second training sample library. For example, for the conditional probability table from the "aeration equipment failure" node (taking values of "yes" or "no") to the "insufficient dissolved oxygen" node (taking values of "yes" or "no"), it is determined by statistically calculating the proportion of insufficient dissolved oxygen when the aeration equipment fails and the proportion of insufficient dissolved oxygen when the aeration equipment is normal in the second training sample library.

[0057] For conditionally sparse probabilities, the Bayesian estimation method is used to introduce a prior distribution to smooth the probability estimation and avoid the zero-probability problem. In specific implementation, for each conditional probability, a small pseudo-count (such as 0.5) is added to ensure that even state combinations not observed in the training data have non-zero probabilities.

[0058] For a discrete child node with a continuous parent node, the soft discretization method is used to map the continuous value to the probability of a discrete state through the sigmoid function. For example, for the mapping from "dissolved oxygen concentration" (continuous value) to "nitrification reaction inhibition" (discrete state), the function P(nitrification reaction inhibition = "yes" | dissolved oxygen = x) = 1 / (1 + exp(α(x - β))) is used, where α and β are parameters learned from the training data.

[0059] In a preferred embodiment of the present invention, the conditional probability learning adopts the expectation maximization algorithm, which can handle the problem of missing data in the training samples. The EM algorithm estimates the distribution of the missing data in the E step and maximizes the likelihood function including the estimated missing data in the M step until convergence.

[0060] Step 44: Based on the real-time operation data as the observed evidence, the joint tree algorithm is used for probabilistic inference, calculate the posterior probability distribution of each root cause node, and output the abnormal cause with the highest probability; Specifically, in the actual application stage, when it is judged in step three that the current operating state is in an abnormal condition, the diagnostic model of the Bayesian network trained based on the second training sample library receives the real-time operation data obtained in step one as input information.

[0061] That is, first, map the real-time operation data to the evidence of the observation feature node layer in the Bayesian network. For example, if the real-time operation data shows that the ammonia nitrogen concentration in the effluent is 15 mg / L, exceeding the standard limit of 10 mg / L, then set the evidence of the "abnormal ammonia nitrogen in the effluent" node to "yes"; if the dissolved oxygen concentration is 0.8 mg / L, lower than the normal range of 2-4 mg / L, then set the evidence of the "abnormal dissolved oxygen concentration" node to "yes".

[0062] Then, based on the set evidence, use the Bayesian network inference algorithm to calculate the posterior probability distribution of each node in the root cause node layer. In the embodiment of the present invention, the joint tree algorithm is used for accurate inference, and probability propagation calculation is realized by converting the Bayesian network into a joint tree structure.

[0063] Finally, according to the posterior probabilities of each node in the root cause node layer, identify and output the abnormal cause with the highest probability. Specifically, select the root cause node with the highest posterior probability as the main abnormal cause, and at the same time consider that the posterior probability exceeds a preset threshold such as 0.3, while other root cause nodes are used as secondary abnormal causes. Directly output these abnormal causes and their corresponding posterior probabilities, where the abnormal cause with the highest posterior probability is used as the main abnormal cause.

[0064] It should be noted that the diagnostic model based on the Bayesian network has the following advantages: First, the graph structure of the Bayesian network intuitively expresses the causal relationship between the components and parameters in the sewage treatment system, making the diagnostic results interpretable; Second, the Bayesian network can handle incomplete observation data, and can still perform effective inference even if some sensor data is missing; Third, through the probability inference mechanism, the Bayesian network can quantify uncertainty and provide a confidence evaluation for each possible abnormal cause, avoiding misjudgment that may be caused by deterministic diagnostic methods; Finally, the Bayesian network model can continuously update the conditional probability table with the accumulation of new data to achieve continuous optimization of the diagnostic ability.

[0065] Furthermore, the integration of multiple decision trees to classify different abnormal types and combine historical treatment experience and current working condition parameters to generate targeted regulation strategies includes the following steps: Step 51: Construct a third training sample library of the strategy generation model that includes historical abnormal events and their successful treatment experiences; Specifically, in the construction stage of the third training sample library, screen out all the abnormal event data that have been successfully processed from the historical operation data of the sewage treatment system. For each abnormal event data, synchronously supplement the following information: key features at the time of the abnormality, abnormal causes, regulation strategies adopted to handle the abnormality, and effect evaluation after the implementation of the regulation strategy.

[0066] In a preferred embodiment of the present invention, the regulation strategies can be classified into three major categories according to the objects of action: water quality parameter regulation strategies, equipment operation regulation strategies, and process parameter regulation strategies. The water quality parameter regulation strategies include, but are not limited to: adjusting the influent pH value, increasing the carbon source dosage, increasing the phosphorus removal agent dosage, etc.; the equipment operation regulation strategies include, but are not limited to: adjusting the operation frequency of the aeration equipment, adjusting the operation time of the stirring equipment, adjusting the flow rate of the reflux pump, etc.; the process parameter regulation strategies include, but are not limited to: adjusting the dissolved oxygen set value, adjusting the sludge concentration, adjusting the hydraulic retention time, adjusting the reflux ratio, etc.

[0067] In the third training sample library, each sample record includes the following fields: timestamp, anomaly type, anomaly cause, key feature vector at the time of anomaly occurrence, adopted regulation strategy, and regulation effect evaluation indicators (such as the improvement degree of the effluent water quality after regulation, energy consumption change, treatment efficiency improvement, etc.).

[0068] Step 52: Construct a strategy generation model based on random forest, including a feature preprocessing module, an anomaly classification module, and a strategy generation module; Specifically, in the feature preprocessing module, the input key feature vector and the anomaly cause are combined and processed to form the input features for strategy generation. The specific processing methods include: concatenating the 64-dimensional key feature with the one-hot encoding vector of the anomaly cause to form an enhanced feature vector; applying principal component analysis (PCA) to the enhanced feature vector for dimensionality reduction; and performing standardization processing on the dimensionality-reduced features to make the dimensions of each feature consistent.

[0069] In the anomaly classification module, a random forest classifier is used to perform fine-grained anomaly type recognition on the features output by the feature preprocessing module. The random forest classifier consists of 100 decision trees. Each decision tree randomly selects a feature subset and a sample subset during training, and determines the final anomaly type through a voting method. The maximum depth of the decision tree is set to 10 to balance the complexity and generalization ability of the model. The output of the anomaly classification module is the fine-grained category of the anomaly, such as "insufficient nitrification caused by high ammonia nitrogen load", "insufficient dissolved oxygen caused by aeration equipment failure", etc.

[0070] In the strategy generation module, for the fine-grained anomaly types identified by the anomaly classification module, a random forest regressor is used to generate specific regulation parameter values. The random forest regressor consists of 100 regression decision trees. Each regression decision tree independently predicts the values of each regulation parameter, and finally obtains the final parameter prediction value through an averaging method. The maximum depth of the regression decision tree is set to 12 to capture the complex non-linear relationships between the parameters. The output of the strategy generation module is a set of specific regulation parameter values, such as "adjust the dissolved oxygen set value in the aeration tank to 3.5 mg / L", "adjust the reflux ratio to 150%", etc.

[0071] Step 53: Based on the third training sample library, train a random forest policy generation model using the ensemble learning method; Specifically, the training process of the random forest policy generation model is divided into the following steps: First, divide the third training sample library into a training set, a validation set, and a test set according to a ratio of 8:1:1. The training set is used for learning model parameters, the validation set is used for model selection and hyperparameter tuning, and the test set is used to evaluate the performance of the final model.

[0072] In the training of the anomaly classification module, the cross-entropy loss function is used to measure the difference between the predicted anomaly type and the true anomaly type. During the training process, each decision tree is constructed using a randomly selected subset of features and a randomly selected subset of samples. The size of the feature subset is set to the square root of the total number of features, and the sample subset is obtained by sampling with replacement and has the same size as the original training set. Each decision tree is trained independently, and finally, the prediction results of all decision trees are integrated by majority voting.

[0073] In the training of the policy generation module, the mean squared error loss function is used to measure the difference between the predicted control parameter value and the true control parameter value. Similarly, each regression decision tree is constructed using a random feature subset and a sample subset. The size of the feature subset is set to one-third of the total number of features, and the sample subset is obtained by sampling with replacement. Each regression decision tree is trained independently, and finally, the prediction results of all regression decision trees are integrated by averaging.

[0074] In the model hyperparameter optimization stage, a grid search combined with cross-validation is used to optimize the key hyperparameters of the random forest, including the number of decision trees, the maximum depth, the minimum number of samples in a leaf node, etc. For each combination of hyperparameters, evaluate the model performance on the validation set, and select the combination of hyperparameters with the best performance as the configuration of the final model.

[0075] In the model evaluation stage, use the test set to evaluate the performance of the final model. For the anomaly classification module, calculate the classification accuracy, precision, recall, and F1 score; for the policy generation module, calculate the mean squared error, mean absolute error, and R² score. At the same time, the sewage treatment experts conduct an artificial evaluation of the control strategies generated by the model to ensure the rationality and operability of the strategies.

[0076] Step 54: Input the anomaly cause and key features into the trained random forest policy generation model to generate targeted control strategies; Specifically, in the actual application stage, the abnormal causes output by the Bayesian network diagnosis model in step four and the key features extracted by the multi-source data fusion network in step two are input into the trained random forest policy generation model. After being processed by the feature preprocessing module, abnormal classification module, and policy generation module, specific regulation strategies for the current abnormal working condition are generated.

[0077] The regulation strategy at least includes the following contents: description of the abnormal type, list of parameters recommended to be adjusted and their target values, priority ranking of the adjustments, and expected regulation effects. Among them, the description of the abnormal type provides an accurate description of the current abnormal situation; the list of parameters recommended to be adjusted includes water quality parameters, equipment operation parameters, or process control parameters that need to be adjusted and their specific target values; the priority ranking of the adjustments is sorted according to the importance of the parameters and the urgency of the adjustments to guide the regulation sequence of the operators; the expected regulation effects predict the improvement of the system performance after the regulation.

[0078] In the preferred embodiment of the present invention, for different types of abnormalities, the random forest policy generation model will generate different types of regulation strategies. For example, for the abnormal situation of the effluent COD exceeding the standard caused by the sudden increase of the influent COD, the generated regulation strategies may include: increasing the aeration volume, extending the hydraulic retention time, adjusting the reflux ratio, etc.; for the abnormal situation of insufficient dissolved oxygen caused by the failure of the aeration equipment, the generated regulation strategies may include: enabling the standby aeration equipment, reducing the influent flow rate, increasing the mixed liquor reflux, etc.; for the abnormal situation of sludge bulking caused by the abnormal sludge concentration, the generated regulation strategies may include: adjusting the sludge discharge amount, adding flocculants, adjusting the aeration mode, etc. Embodiment 2

[0079] As Figure 2 shown, the sewage treatment data management system based on artificial intelligence includes a real-time data collection module, a key feature extraction module, an abnormal working condition judgment module, an abnormal cause identification module, and a policy generation module; among them, each module is connected electrically. The real-time data collection module obtains the real-time operation data of the sewage treatment system and sends the real-time operation data to the key feature extraction module. The key feature extraction module extracts the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network and sends the key features to the abnormal working condition judgment module. The abnormal working condition judgment module inputs the key features into a pre-trained state evaluation model to generate the sewage treatment state; according to the sewage treatment state, it judges whether the current operation state is in an abnormal working condition. If it is judged to be in an abnormal working condition, it sends the real-time operation data and the sewage treatment state to the abnormal cause identification module. The abnormal cause identification module, through a diagnostic model based on a Bayesian network, identifies the abnormal cause based on real-time operation data and sewage treatment status, and sends the abnormal cause and real-time operation data to the policy generation module; The policy generation module, based on the abnormal cause, inputs the key features into a policy generation model based on a random forest. This model classifies different abnormal types by integrating multiple decision trees, and combines historical treatment experience and current working condition parameters to generate targeted regulation policies.

[0080] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A sewage treatment data management method based on artificial intelligence, characterized in that, It includes the following steps: Step 1: Obtain the real-time operation data of the sewage treatment system; Step 2: Extract the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network; Step 3: Input the key features into a pre-trained state evaluation model to generate the sewage treatment state; based on the sewage treatment state, judge whether the current operation state is in an abnormal condition; Step 4: If it is in an abnormal condition, identify the abnormal cause through a diagnostic model based on the Bayesian network, based on the real-time operation data and the sewage treatment state; Step 5: Based on the abnormal cause, input the key features into a policy generation model based on the random forest. This model classifies different abnormal types by integrating multiple decision trees and combines historical treatment experience and current working condition parameters to generate targeted regulation strategies.

2. The method for managing sewage treatment data based on artificial intelligence according to claim 1, wherein, The obtaining of the real-time operation data of the sewage treatment system includes the following steps: Step 11: Collect the real-time parameter data of each process unit of the sewage treatment system through a distributed sensor network; Step 12: Structurally process the collected real-time parameter data through a field control unit adopting an edge computing architecture to form structured real-time operation data.

3. The method for managing sewage treatment data based on artificial intelligence according to claim 2, wherein, The method of extracting the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network includes the following steps: Step 21: Collect the historical operation data of the sewage treatment system and construct a multi-source heterogeneous data training sample library; Step 22: Construct a multi-source data fusion network model. The multi-source data fusion network model adopts a multi-stream architecture design, including a feature extraction module, a time series modeling module, a multi-source fusion module, and a feature output module; Step 23: Based on the constructed multi-source heterogeneous data training sample library, adopt an end-to-end supervised learning method to train the multi-source data fusion network model and optimize the network parameters; Step 24: Use the trained multi-source data fusion network to process the real-time operation data and extract the key features of the sewage treatment system.

4. The method for managing sewage treatment data based on artificial intelligence according to claim 3, wherein, The structure of the multi-source data fusion network model includes: In the feature extraction module, for the data stream of water quality parameters, a one-dimensional convolutional neural network structure is adopted, which includes 3 consecutive one-dimensional convolutional blocks; each convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation function, and a max pooling layer; For the data stream of operation parameters, a multi-scale residual network structure is adopted, which includes 4 residual blocks; each residual block contains two one-dimensional convolutional layers and a skip connection to capture the device operation features at different time scales; For the data stream of process control parameters, a gated recurrent unit network is adopted, which includes a 2-layer bidirectional GRU structure; the GRU network can effectively model the long-term dependence relationship and time series change trend of process control parameters, and the bidirectional structure considers both historical and future information at the same time, improving the comprehensiveness of feature extraction; In the time series modeling module, the time attention mechanism is applied to the features extracted from the three data streams respectively to learn the importance weights of the features at different time points. That is, the time attention mechanism generates attention weights by calculating the similarity between the query vector and the features at each time point, and then performs weighted summation on the features at each time point to obtain a feature representation with time importance. In the multi-source fusion module, an adaptive feature fusion strategy is adopted, which includes three steps: feature alignment, feature interaction, and feature fusion. In the feature alignment step, the features output by the three data streams through the time series modeling module are mapped to the same feature space through a fully connected layer. In the feature interaction step, the cross-stream attention mechanism is used to calculate the mutual influence between the corresponding feature spaces of different data streams. In the feature fusion step, the gated fusion unit adaptively determines the fusion weights according to the correlation of the feature spaces output by each data stream feature after the feature interaction step, and generates a comprehensive feature representation. In the feature output module, the comprehensive feature representation is reduced in dimension and feature selection is performed according to the fusion weights through a two-layer fully connected network to extract key features.

5. The method for managing sewage treatment data based on artificial intelligence according to claim 4, characterized in that, The steps of inputting the key features into a pre-trained state evaluation model to generate the sewage treatment state are as follows: Step 31: Synchronously construct a training sample library for the state evaluation model with the multi-source heterogeneous data training sample library. The training sample library contains two types of data: normal working condition data and abnormal working condition data. Step 32: Design a state evaluation model with a network structure including a feature extraction layer, a time series modeling layer, and a state classification layer. Step 33: Based on the constructed training sample library, adopt a supervised learning method to train the state evaluation model with the goal of optimizing state classification. Step 34: The state evaluation model receives the key features from Step 2 and generates a state classification result of the sewage treatment state in real time.

6. The method for managing sewage treatment data based on artificial intelligence according to claim 5, wherein, The structure of the state evaluation model includes: In the feature extraction layer, a multi-layer perceptron network is used to perform nonlinear transformation and dimensionality reduction on the input key features. The feature extraction layer contains three fully connected layers, and the activation function uses the ReLU function to enhance the expression ability of the state evaluation model for nonlinear relationships. A Dropout layer is added between the fully connected layers to prevent the model from overfitting. In the time series modeling layer, a bidirectional long short-term memory network is used to capture the time-dependent relationship of the sewage treatment system parameters. The time series modeling layer contains two bidirectional long short-term memory networks. In the state classification layer, the output features of the time series modeling layer pass through two fully connected layers, and then the state classification result of the sewage treatment system is output through the Softmax function.

7. The method for managing sewage treatment data based on artificial intelligence according to claim 6, wherein The steps of identifying the abnormal cause based on the diagnostic model based on the Bayesian network, based on the real-time operation data and the sewage treatment state, are as follows: Step 41: Based on the historical operation data of the sewage treatment system and the abnormal event records annotated by experts, construct a second training sample library of the Bayesian network diagnostic model containing the causal relationships of abnormal types, causes, and symptoms. Step 42: Based on the second training sample library, adopt a method combining structure learning and expert knowledge to construct a three-layer hierarchical Bayesian network structure including a root cause node layer, an intermediate state node layer, and an observation feature node layer; Step 43: Based on the labeled data in the second training sample library, adopt the maximum likelihood estimation and expectation maximization algorithm to learn the conditional probability table of the causal relationships between the nodes in the Bayesian network, and handle the data sparsity problem through the Bayesian estimation method; Step 44: Based on the real-time operation data as the observation evidence, adopt the joint tree algorithm for probability inference, calculate the posterior probability distribution of each root cause node, and output the abnormal cause with the highest probability.

8. The method for managing sewage treatment data based on artificial intelligence according to claim 7, characterized in that The Bayesian network structure adopts a hierarchical design and includes three layers of nodes: The first layer is the root cause node layer, which represents the fundamental abnormal causes with probabilities greater than a preset probability threshold, including the inlet water quality abnormal node group, the equipment failure node group, and the process parameter abnormal node group; The second layer is the intermediate state node layer, which represents the system state changes caused by the root causes, including the biochemical reaction state node group, the hydraulic state node group, and the solid-liquid separation state node group; The third layer is the observation feature node layer, which represents the abnormal features that can be observed through the sensor network, including the water quality feature node group, the equipment operation feature node group, and the process feature node group.

9. The artificial intelligence-based sewage treatment data management method according to claim 8, wherein the integration of multiple decision trees to classify different abnormal types, and combine historical treatment experience and current working condition parameters to generate targeted regulation strategies includes the following steps: Step 51: Construct a third training sample library of the strategy generation model including historical abnormal events and their successful treatment experiences; Step 52: Construct a strategy generation model based on a random forest, including a feature preprocessing module, an abnormal classification module, and a strategy generation module; In the feature preprocessing module, the input key feature vector and the abnormal cause are combined and processed to form the input features for strategy generation; In the abnormal classification module, a random forest classifier is used to perform fine-grained abnormal type recognition on the features output by the feature preprocessing module; In the strategy generation module, for the fine-grained abnormal types identified by the abnormal classification module, a random forest regressor is used to generate specific regulation parameter values; Step 53: Based on the third training sample library, adopt an ensemble learning method to train the random forest strategy generation model; Step 54: Input the abnormal cause and key features into the trained random forest strategy generation model to generate targeted regulation strategies.

10. An artificial intelligence-based sewage treatment data management system for implementing the artificial intelligence-based sewage treatment data management method according to any one of claims 1-9, characterized in that, It includes a real-time data collection module, a key feature extraction module, an abnormal working condition judgment module, an abnormal cause identification module, and a strategy generation module; among them, each module is electrically connected; The real-time data collection module obtains the real-time operation data of the sewage treatment system and sends the real-time operation data to the key feature extraction module; The key feature extraction module extracts the key features of the sewage treatment system from the real-time operation data through a pre-trained multi-source data fusion network and sends the key features to the abnormal working condition judgment module; The abnormal condition judgment module inputs the key features into a pre-trained state evaluation model to generate the sewage treatment state; according to the sewage treatment state, it judges whether the current operating state is in an abnormal condition. If it is judged to be an abnormal condition, it sends the real-time operation data and the sewage treatment state to the abnormal cause identification module; The abnormal cause identification module, through a diagnostic model based on a Bayesian network, identifies the abnormal cause based on the real-time operation data and the sewage treatment state, and sends the abnormal cause and the real-time operation data to the policy generation module; The policy generation module, based on the abnormal cause, inputs the key features into a policy generation model based on a random forest. This model classifies different abnormal types by integrating multiple decision trees, and combines historical treatment experience and current working condition parameters to generate targeted regulation policies.

Citation Information

Patent Citations

  • Wind turbine generator remote fault diagnosis method and system based on cloud platform

    CN117930815A

  • Wind turbine generator fault diagnosis method and system based on deep learning

    CN118332442A

  • Fault detection visualization processing system and method for industrial switch

    CN119254613A

  • Electric power system fault diagnosis and early warning system based on AI

    CN120011874A

  • Intelligent systems and methods for process and asset health diagnosis, anomoly detection and control in wastewater treatment plants or drinking water plants

    WO2019071384A1

Cited By

  • Method and system for assisting in managing water efficiency in whole process of water system

    CN120782593A

  • Method and system for assisting management of water efficiency in a water system

    CN120782593B

  • Water quality up-to-standard monitoring method for rainwater treatment and reuse

    CN120802815A

  • Self-adaptive dynamic energy testing method based on combined heat and power generation

    CN120822120A

  • Intelligent decision fusion system for multi-stage process cooperation of sewage plant

    CN121680296A