Sewage treatment state monitoring method and system based on image feature analysis

A wastewater treatment status monitoring method based on image feature analysis and multimodal data fusion solves the problem of real-time monitoring and intelligent control of the flocculation process in wastewater treatment, achieving real-time performance and accuracy in the wastewater treatment process and reducing reagent consumption.

CN121543027APending Publication Date: 2026-02-17LANZHOU PETROCHEMICAL VOCATIONAL & TECH UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202610057748.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In traditional wastewater treatment, machine vision technology cannot achieve real-time and intelligent detection of wastewater flocs. Existing technologies cannot effectively monitor the flocculation process during wastewater treatment, leading to excessive or insufficient dosage of chemicals, which affects the stability of effluent quality.

Method used

A wastewater treatment status monitoring method based on image feature analysis is adopted. Data is acquired using industrial cameras and sensor clusters, multimodal data stitching and feature fusion are performed, and attention mechanism and recurrent neural network are combined to conduct real-time status assessment and chemical dosage control.

Benefits of technology

It enables real-time monitoring and intelligent control of the wastewater treatment process, improves the accuracy and real-time performance of the flocculation process, and reduces reagent consumption and operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543027A_ABST
    Figure CN121543027A_ABST
Patent Text Reader

Abstract

The invention provides a sewage treatment state monitoring method and system based on image feature analysis. The method comprises the following steps: acquiring a plurality of image frame data and a plurality of sensor time sequence data of a flocculation basin in a first time period; performing data splicing on the image frame data and the plurality of sensor time sequence data to obtain a plurality of multi-modal data at different moments; inputting each piece of multi-modal data into a preset floc detection model, so that the floc detection model performs feature fusion on each piece of data in the multi-modal data through an attention mechanism, and generates a frame-level floc feature vector corresponding to the multi-modal data according to a feature fusion result; and inputting each frame-level floc feature vector and the plurality of sensor time sequence data into a preset sewage treatment state evaluation model, so that the sewage treatment state evaluation model outputs the current sewage treatment state according to the input data, and the accuracy and real-time performance of sewage treatment state monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal data fusion and wastewater status monitoring technology, and in particular to a wastewater treatment status monitoring method and system based on image feature analysis. Background Technology

[0002] Flocculation is a crucial step in wastewater treatment. Its core principle is to destabilize and aggregate small suspended particles and colloidal impurities in the water that are difficult to settle, forming larger and denser flocs. This creates the necessary conditions for subsequent physical separation processes such as sedimentation and filtration. Traditional wastewater treatment flocculation monitoring and control heavily relies on human experience and discrete sensor data. Operators typically rely on visual observation of floc formation, combined with readings from a few online instruments such as influent flow meters, pH meters, and turbidity meters, to make empirical judgments and manually adjust chemical dosages. This method has significant limitations: firstly, human observation is highly subjective, cannot be quantified, and is difficult to sustain; secondly, conventional sensors can only provide macroscopic parameters of the effluent or a specific point, failing to perceive the morphology, particle size distribution, and dynamic evolution of flocs within the flocculation reactor; and thirdly, image data and sensor data are analyzed in isolation, leading to delayed and incomplete state assessments. Therefore, traditional methods are difficult to adapt to real-time fluctuations in water quality and quantity, and the dosage of chemicals is often excessive or insufficient, which not only wastes the chemicals but also affects the stability of the effluent water quality.

[0003] To overcome the aforementioned bottlenecks, machine vision technology has been introduced in recent years. This technology captures images of flocs using cameras and extracts static features such as floc size and quantity using image processing algorithms. However, these methods are mostly limited to offline or post-event analysis, lacking real-time performance. Furthermore, the feature extraction methods are relatively simple, with weak anti-interference capabilities, and they haven't been deeply integrated with real-time sensor data. While some research has attempted to integrate multiple sensors, these efforts mostly remain at the data level, providing simple display or alarms, lacking a complete technical framework capable of deeply fusing multi-source heterogeneous information at the feature level and using this information for real-time status assessment and intelligent decision-making. Achieving millisecond-level synchronization of multimodal data, constructing a dynamic analysis model that integrates visual and process knowledge, and forming a real-time closed loop of "perception-analysis-decision-execution" remain key challenges for achieving intelligent and precise control of the flocculation process. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a wastewater treatment status monitoring method and system based on image feature analysis, thereby improving the accuracy and real-time performance of wastewater treatment status monitoring.

[0005] In a first aspect, embodiments of this application provide a wastewater treatment status monitoring method based on image feature analysis, comprising: Acquire several image frame data and several sensor time series data of the flocculation tank in the first time period; Based on the time sequence, each of the image frame data and the time sequence data of the several sensors are respectively spliced ​​to obtain multimodal data at several different times; Each of the multimodal data is input into a preset floc detection model, so that the floc detection model performs feature fusion on each data in the multimodal data through an attention mechanism, and generates a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. The floc detection model is constructed based on the target detection model. Each frame-level floc feature vector and the time-series data from the several sensors are input into a preset wastewater treatment status assessment model. The wastewater treatment status assessment model generates hidden state vectors at each time step based on the input data through a gating mechanism, and outputs the current wastewater treatment status based on each hidden state vector. The hidden state vectors are used to capture the dependency relationship between floc morphology features and sensor data features. The wastewater treatment status assessment model is constructed based on a recurrent neural network model.

[0006] This application provides a wastewater treatment status monitoring method based on image feature analysis. By fusing multimodal data, it achieves comprehensive dynamic perception of the flocculation process, thereby enabling accurate assessment of the wastewater treatment status. Specifically, this embodiment first acquires image frame data of the flocculation tank and sensor time-series data, concatenates them in chronological order to form multimodal data, and then effectively fuses image features and sensor features through a floc detection model, extracting floc feature vectors at each frame level to achieve accurate localization and feature extraction of each floc in the image. In this process, sensor time-series data mainly serves as supplementary information to the image data to improve the accuracy of floc detection. Then, these feature vectors and sensor time-series data are input into an evaluation model based on a recurrent neural network. A gating mechanism is used to extract hidden state vectors, capturing long-term dependencies between various features, and finally outputting the current wastewater treatment status. This application overcomes the limitations of the separation between image data and sensor data in traditional monitoring, achieving deep fusion of visual information and sensor data, thereby significantly improving the accuracy of wastewater treatment status monitoring. Furthermore, traditional monitoring technologies typically assess the current wastewater treatment status based on the final effluent quality indicators, at which point wastewater treatment is already complete, thus exhibiting a certain degree of lag. In contrast, this embodiment utilizes data from the flocculation tank during the wastewater treatment process to monitor the current wastewater treatment status in real time, thereby improving the real-time performance of wastewater treatment status monitoring.

[0007] Furthermore, the acquisition of several image frame data and several sensor time-series data of the flocculation tank within the first time period includes: The flocculation tank is acquired using a pre-set industrial camera within a first time period, along with several initial image frames. The flocculation tank acquires several initial sensor time-series data within a first time period through a preset sensor cluster. The several initial sensor time-series data include influent flow rate time-series data, pH value time-series data, conductivity time-series data, raw water turbidity time-series data, and temperature time-series data. Denoising, enhancement, and color space conversion operations are performed sequentially on each of the initial image frame data to obtain each image frame data; The initial sensor time series data are sequentially filtered, outlier removal, missing value interpolation, and normalization to obtain the time series data for each sensor.

[0008] This application further clarifies the specific steps of data acquisition and preprocessing. Raw images and multi-dimensional sensor data are acquired through industrial cameras and sensor clusters, and the initial data undergoes a series of processing steps including denoising, enhancement, filtering, and normalization. This design significantly improves data quality and reliability, laying a solid foundation for subsequent model analysis. Specifically, image denoising and enhancement effectively reduce the interference of factors such as lighting changes and water turbidity on visual analysis, while filtering and outlier removal of sensor data avoid the negative impact of noise and abnormal sampling on time-series analysis. Through a systematic preprocessing workflow, this embodiment ensures the consistency and availability of multimodal data from the source, thereby enhancing the robustness and environmental adaptability of the entire monitoring system and improving the accuracy of subsequent wastewater treatment status monitoring.

[0009] In one possible implementation, the step of concatenating each of the image frame data with the time-series data from the plurality of sensors, based on time sequence, to obtain multimodal data at several different times includes: The image frame data is traversed sequentially. For any image frame data, according to the timestamp of the image frame data, the corresponding sensor sub-data is extracted from the time series data of each sensor, and the sensor sub-data is concatenated with the image frame data to obtain the multimodal data with the corresponding timestamp. After the traversal is completed, multimodal data at several different times are obtained.

[0010] This application provides a method for stitching multimodal data, which precisely matches and stitches image frames with corresponding sensor sub-data at specific times through timestamp alignment. This process achieves millisecond-level multi-source data synchronization, overcoming the information lag problem caused by asynchronous data acquisition in traditional methods. The time-sequential traversal and stitching mechanism ensures that each image frame can be organically combined with concurrent sensor data (such as influent flow rate, pH value, turbidity, etc.), thereby constructing a multimodal sample with temporal consistency. This refined data organization provides an accurate data foundation for subsequent temporal feature fusion and dynamic state assessment, significantly improving the model's ability to characterize process evolution trends.

[0011] In one possible implementation, the floc detection model fuses features from various data points in the multimodal data using an attention mechanism, and generates a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result, including: Semantic features are extracted from the image frame data in the multimodal data through the visual backbone network inside the model to obtain the corresponding semantic image feature map. The model uses a fully connected network to encode the features of each sensor sub-data in the multimodal data, thereby obtaining the corresponding sensor feature vector. The sensor feature vector is expanded into a sensor feature map with the same spatial dimension as the semantic image feature map by spatial copying or broadcasting. The semantic image feature map and the sensor feature map are fused using an attention mechanism to generate a multimodal joint feature map; The model uses an internal detection head to perform target detection on the multimodal joint feature map, thereby determining the position of each floc in the multimodal joint feature map. Based on the position of each floc, the corresponding region features are cropped from the semantic image feature map, and pooling is performed on each region feature to obtain the morphological features of each floc. The frame-level flocculent feature vector is obtained by aggregating and statistically analyzing each of the morphological features according to a number of preset index types.

[0012] This application provides a method for floc detection using a model. It extracts semantic features from images through a visual backbone network, encodes sensor features through a fully connected network, and then uses an attention mechanism after spatial expansion to achieve interaction and fusion of image and sensor features. This method not only achieves deep fusion of multimodal information but also dynamically highlights key features through attention weights, thereby improving the accuracy and robustness of floc detection. Furthermore, the detection head locates the floc positions and extracts regional features from the feature map, which are then pooled and statistically aggregated to form frame-level feature vectors, ensuring that each frame of data represents the overall state of the flocs at the current moment. This design allows the model to simultaneously perceive the visual morphology of the flocs and the technological environment, providing rich and structured feature representations for subsequent time-series analysis.

[0013] Furthermore, the wastewater treatment status monitoring method also includes: When the first semantic image feature map corresponding to the last moment in the first time period is obtained, several sensor data corresponding to the last moment in the first time period and the current dosage of the flocculation tank are input into a preset physical information neural network, so that the physical information neural network calculates the theoretical number of flocs and the theoretical average volume of flocs based on the several sensor data and the current dosage. Based on the first semantic image feature map, determine the corresponding number of observed flocs and the average volume of observed flocs; Calculate the first deviation between the theoretical number of flocs and the observed number of flocs; calculate the second deviation between the theoretical average volume of flocs and the observed average volume of flocs; Based on the first deviation and the second deviation, determine whether there is any abnormal data acquisition of each sensor in the flocculation tank; If an abnormal data acquisition is detected, the time-series data of the several sensors are corrected based on several historical sensor data to obtain several corrected sensor time-series data. Based on each of the image frame data and each of the corrected sensor time-series data, the floc detection model is used to regenerate each frame-level floc feature vector. Sensor abnormality information is generated and sent to the designated device.

[0014] This application provides a method for verifying and correcting sensor data. Considering that in wastewater treatment scenarios, sensors are often immersed in complex, corrosive, and adhesive wastewater for extended periods, making them highly susceptible to biofouling, chemical scaling, or physical blockage, leading to measurement drift, decreased sensitivity, or even complete failure, this embodiment utilizes image data from the last moment to verify the data collected by the sensors. First, a physical information neural network is used to calculate the theoretical characteristics of flocs, and a first semantic image feature map is used to calculate the observed characteristics of the flocs. Then, by comparing the theoretical and observed characteristics of the flocs, it is determined whether the current sensor cluster is abnormal. If an abnormality is detected, the sensor data is corrected and the abnormality information is reported, further ensuring the accuracy of wastewater treatment status monitoring under abnormal operating conditions.

[0015] In one possible implementation, the wastewater treatment status assessment model generates hidden state vectors at various time points based on input data using a gating mechanism, and outputs the current wastewater treatment status based on each of the hidden state vectors, including: Each frame-level floc feature vector is concatenated with the time-series data from the several sensors to obtain a time-series joint feature vector. Based on the time sequence, the hidden state vectors are extracted sequentially from the joint feature vectors at different times in the time sequence joint feature vectors through the gating unit inside the model, so as to obtain the corresponding hidden state vectors. During the extraction of the hidden state vector, for any current joint feature vector, corresponding forgotten information, newly added information, output information, and candidate memory cells are generated through the forget gate, input gate, output gate, and memory cell, respectively; the memory cells of the previous time step are updated according to the forgotten information, the newly added information, and the candidate memory cells to obtain the current memory cell; and the hidden state vector of the current time step is generated according to the current memory cell and the hidden state vector of the previous time step. Calculate the attention weights of each hidden state vector, and perform weighted fusion of each hidden state vector according to each attention weight to obtain the context vector; The model uses a fully connected classification head to generate probability distribution data of the wastewater treatment status based on the context vector, and then outputs the current wastewater treatment status based on the probability distribution data.

[0016] This application provides a method for assessing the state of wastewater treatment using a model. The method progressively processes the temporal joint feature vector step-by-step through gating units. During processing, gating structures such as forget gates, input gates, and output gates are used to dynamically update the memory state, thereby capturing long-term dependencies in the sequence. Based on this, an attention mechanism is introduced to weightedly fuse the hidden states at each time step, generating a context vector that centrally reflects key temporal information. Finally, a fully connected classification head outputs the state probability distribution. This design enables the model to comprehensively consider data change trends over long periods, rather than just focusing on data at the current moment. It accurately identifies state transitions and abnormal patterns in the wastewater treatment process, significantly improving the temporal awareness and decision reliability of state assessment, and enhancing the accuracy and real-time performance of wastewater treatment state monitoring.

[0017] Furthermore, the step of constructing the floc detection model based on the target detection model includes: An initial floc labeling model is constructed based on the YOLOv8 model. The initial floc labeling model includes a visual backbone network, a fully connected network, a dimension synchronization module, a feature fusion module, and a detection head. Acquire several historical multimodal data sets and manually mark the floc positions in each historical multimodal data set to construct the first training dataset; The initial floc labeling model is trained using the first training dataset, and the various model parameters in the initial floc labeling model are updated to obtain the floc labeling model; The floc labeling model is combined with a pre-trained feature extraction module to construct the floc detection model, wherein the feature extraction module is used to extract frame-level floc feature vectors from the corresponding semantic image feature map according to each floc position.

[0018] This application provides a method for constructing a floc detection model. Based on the advanced object detection framework YOLOv8, it incorporates a fully connected network, a dimension synchronization module, and a feature fusion module, enabling it to simultaneously process image and structured sensor data. By training with historical multimodal data and fusing a pre-trained feature extraction module, the model not only possesses strong generalization ability but can also be optimized for floc detection tasks, improving the accuracy of subsequent wastewater treatment status monitoring. It is important to note that the feature fusion module and the floc labeling model in this application are trained separately, providing a modular and phased model training approach. This modular and scalable design allows the model to quickly adapt to different sensor configurations and water treatment scenarios, facilitating practical engineering deployment.

[0019] Furthermore, the process of constructing the wastewater treatment status assessment model based on a recurrent neural network model includes: An initial evaluation model is constructed based on the LSTM model, which includes a gating unit, a feature fusion unit, and a fully connected classification head. Historical multimodal datasets from different time periods are acquired, and the joint feature vectors of historical time series from different time periods are obtained through the floc detection model. Based on the wastewater quality indicators of the wastewater treatment effluent in each corresponding time period, the wastewater treatment status corresponding to each of the historical time series joint feature vectors is manually marked, and then a second training dataset is constructed. The initial evaluation model is trained using the second training dataset, and the various model parameters in the initial evaluation model are updated to obtain the wastewater treatment status evaluation model.

[0020] This application provides a training method for a wastewater treatment status assessment model. It utilizes a pre-trained floc detection model, generates historical time-series joint feature vectors for different time periods from a historical multimodal dataset, and constructs a second training dataset by combining manually labeled wastewater treatment status data. The second training dataset is then used to perform end-to-end training on the initial LSTM-based assessment model, enabling the model to learn the complex mapping relationship from multimodal time-series features to status labels. This process fully leverages the image features and feature association knowledge contained in historical data and sensor data, giving the model adaptability and status reasoning capabilities to different operating conditions. Therefore, in practical applications, it can provide accurate and reliable status assessment results, improving the accuracy of wastewater treatment status monitoring.

[0021] In one possible implementation, the wastewater treatment status monitoring method further includes controlling the dosage of chemicals in the flocculation tank during a second time period based on the current wastewater treatment status, specifically: Based on the linear extrapolation method, the predicted influent flow rate change trend and the predicted raw water turbidity change trend of the flocculation tank in the second time period are predicted by using the influent flow rate time series data and raw water turbidity time series data from the time series data of the aforementioned sensors. Based on the current wastewater treatment status, the estimated influent flow rate change trend, and the estimated raw water turbidity change trend, a query is performed in a preset rule base to determine the corresponding dosing rule, and the target dosing amount is determined according to the dosing rule. The dosing rule includes antecedent conditions and a conclusion. The antecedent conditions include the wastewater treatment status, the influent flow rate change trend, and the raw water turbidity change trend, and the conclusion is the dosing amount. Based on the target dosage and the current dosage, the dosage of the flocculation tank during the second time period is controlled by a preset control strategy.

[0022] This embodiment further expands the dosing control function based on status monitoring. Based on the current wastewater treatment status and predicted influent flow rate and raw water turbidity trends, it determines the optimal dosing dosage through rule base queries and automatically adjusts the dosage using control strategies. This design forms a closed-loop control system of "sensing-analysis-decision-execution," achieving full automation from status monitoring to wastewater treatment control. By combining linear extrapolation prediction with rule-based reasoning, this embodiment can proactively address fluctuations in water quality and quantity, achieving precise dosing. While ensuring stable effluent quality, it significantly reduces reagent consumption and operating costs, thereby improving wastewater treatment efficiency.

[0023] Secondly, embodiments of this application provide a wastewater treatment status monitoring system based on image feature analysis, including an acquisition module, a data fusion module, a floc detection module, and a status assessment module; The acquisition module is used to acquire several image frame data and several sensor time series data of the flocculation tank in the first time period. The first data fusion module is used to concatenate each of the image frame data with the time-series data of the several sensors based on time order to obtain multimodal data at several different times. The floc detection module is used to input each of the multimodal data into a preset floc detection model, so that the floc detection model can perform feature fusion on each data in the multimodal data through an attention mechanism, and generate a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. The floc detection model is constructed based on the target detection model. The state assessment module is used to input the frame-level floc feature vectors and the time-series data of the several sensors into a preset wastewater treatment state assessment model, so that the wastewater treatment state assessment model generates hidden state vectors at each time step according to the input data through a gating mechanism, and outputs the current wastewater treatment state according to each hidden state vector. The hidden state vectors are used to capture the dependency relationship between floc morphological features and sensor data features. The wastewater treatment state assessment model is constructed based on a recurrent neural network model. Attached Figure Description

[0024] Figure 1 A flowchart illustrating a wastewater treatment status monitoring method based on image feature analysis, provided in an embodiment of this application; Figure 2 A schematic diagram of the model structure of the floc detection model in a wastewater treatment status monitoring method based on image feature analysis provided in this application embodiment; Figure 3 This is a schematic diagram of a wastewater treatment status monitoring system based on image feature analysis, provided as an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0027] Example 1: As Figure 1 As shown, Embodiment 1 provides a wastewater treatment status monitoring method based on image feature analysis, including steps S1-S4: Step S1, acquiring several image frame data and several sensor time-series data of the flocculation tank within a first time period; Step S2, based on the time sequence, concatenating each of the image frame data and the several sensor time-series data to obtain several multimodal data at different times; Step S3, inputting each of the multimodal data into a preset floc detection model, so that the floc detection model performs feature fusion on each data in the multimodal data through an attention mechanism, and generates data based on the feature fusion result. The frame-level floc feature vectors corresponding to the multimodal data are obtained by constructing the floc detection model based on the target detection model; Step S4: Input each of the frame-level floc feature vectors and the time-series data of the several sensors into a preset wastewater treatment status assessment model, so that the wastewater treatment status assessment model generates hidden state vectors at each time step according to the input data through a gating mechanism, and outputs the current wastewater treatment status according to each of the hidden state vectors. The hidden state vectors are used to capture the dependency relationship between floc morphological features and sensor data features. The wastewater treatment status assessment model is obtained by constructing the model based on a recurrent neural network model.

[0028] This application provides a wastewater treatment status monitoring method based on image feature analysis. By fusing multimodal data, it achieves comprehensive dynamic perception of the flocculation process, thereby enabling accurate assessment of the wastewater treatment status. Specifically, this embodiment first acquires image frame data of the flocculation tank and sensor time-series data, concatenates them in chronological order to form multimodal data, and then effectively fuses image features and sensor features through a floc detection model, extracting floc feature vectors at each frame level to achieve accurate localization and feature extraction of each floc in the image. In this process, sensor time-series data mainly serves as supplementary information to the image data to improve the accuracy of floc detection. Then, these feature vectors and sensor time-series data are input into an evaluation model based on a recurrent neural network. A gating mechanism is used to extract hidden state vectors, capturing long-term dependencies between various features, and finally outputting the current wastewater treatment status. This application overcomes the limitations of the separation between image data and sensor data in traditional monitoring, achieving deep fusion of visual information and sensor data, thereby significantly improving the accuracy of wastewater treatment status monitoring. Furthermore, traditional monitoring technologies typically assess the current wastewater treatment status based on the final effluent quality indicators, at which point wastewater treatment is already complete, thus exhibiting a certain degree of lag. In contrast, this embodiment utilizes data from the flocculation tank during the wastewater treatment process to monitor the current wastewater treatment status in real time, thereby improving the real-time performance of wastewater treatment status monitoring.

[0029] Furthermore, in step S1, acquiring several image frame data and several sensor time-series data of the flocculation tank within a first time period includes: acquiring several initial image frame data of the flocculation tank within a first time period using a preset industrial camera; acquiring several initial sensor time-series data of the flocculation tank within a first time period using a preset sensor cluster, wherein the several initial sensor time-series data include influent flow rate time-series data, pH value time-series data, conductivity time-series data, raw water turbidity time-series data, and temperature time-series data; sequentially performing denoising, enhancement, and color space conversion operations on each of the initial image frame data to obtain each image frame data; and sequentially performing filtering, outlier removal, missing value interpolation, and normalization processing on each of the initial sensor time-series data to obtain each sensor time-series data.

[0030] This application further clarifies the specific steps of data acquisition and preprocessing. Raw images and multi-dimensional sensor data are acquired through industrial cameras and sensor clusters, and the initial data undergoes a series of processing steps including denoising, enhancement, filtering, and normalization. This design significantly improves data quality and reliability, laying a solid foundation for subsequent model analysis. Specifically, image denoising and enhancement effectively reduce the interference of factors such as lighting changes and water turbidity on visual analysis, while filtering and outlier removal of sensor data avoid the negative impact of noise and abnormal sampling on time-series analysis. Through a systematic preprocessing workflow, this embodiment ensures the consistency and availability of multimodal data from the source, thereby enhancing the robustness and environmental adaptability of the entire monitoring system and improving the accuracy of subsequent wastewater treatment status monitoring.

[0031] In a preferred embodiment, a high-definition industrial camera (e.g., 1920×1080 resolution, 30fps) is pre-installed above the flocculation tank of the wastewater treatment plant, and a sensor cluster is deployed at the inlet of the flocculation tank and inside the tank body, including: an influent flow meter for real-time monitoring of influent flow; a pH sensor for monitoring the acidity and alkalinity of the water; a conductivity sensor for monitoring the ion concentration of the water; a turbidity sensor for monitoring the turbidity of the raw water; and a temperature sensor for monitoring the water temperature.

[0032] The initial time period is set to 10 minutes. During the data acquisition process, data is synchronously collected within 10 minutes using the aforementioned acquisition equipment. Specifically, an industrial camera continuously captures images of the flocculation tank surface, generating an initial image frame sequence containing timestamps. The sensor cluster collects data on influent flow rate, pH value, conductivity, raw water turbidity, and temperature at a frequency of once per second, forming several initial sensor time-series data with timestamps. For each initial image frame in the initial image frame sequence, a Gaussian filtering algorithm is first used to reduce random noise in the image. Then, histogram equalization technology is used to enhance image contrast, making the floc outline clearer. Finally, the image is converted from the RGB color space to a grayscale or Lab color space more suitable for feature extraction. For each initial sensor time-series data, a moving average filter is first used to smooth the time-series data and eliminate instantaneous fluctuations. Then, based on the 3σ principle, obviously abnormal data points are identified and removed, while missing data points are filled using linear interpolation. Finally, all sensor data are scaled to the [0, 1] interval to eliminate the influence of dimensions.

[0033] In one possible implementation, in step S2, the step of concatenating each of the image frame data with the time-series data of the plurality of sensors based on time order to obtain multimodal data at several different times includes: sequentially traversing each of the image frame data, wherein, for any image frame data, according to the timestamp of the image frame data, extracting each sensor sub-data corresponding to the timestamp from each of the sensor time-series data, and concatenating each sensor sub-data with the image frame data to obtain multimodal data at the corresponding timestamp; after the traversal is completed, the multimodal data at several different times is obtained.

[0034] This application provides a method for stitching multimodal data, which precisely matches and stitches image frames with corresponding sensor sub-data at specific times through timestamp alignment. This process achieves millisecond-level multi-source data synchronization, overcoming the information lag problem caused by asynchronous data acquisition in traditional methods. The time-sequential traversal and stitching mechanism ensures that each image frame can be organically combined with concurrent sensor data (such as influent flow rate, pH value, turbidity, etc.), thereby constructing a multimodal sample with temporal consistency. This refined data organization provides an accurate data foundation for subsequent temporal feature fusion and dynamic state assessment, significantly improving the model's ability to characterize process evolution trends.

[0035] In a preferred embodiment, the system iterates through each preprocessed image frame data in chronological order: for each image frame, based on its timestamp, it extracts sub-data such as influent flow rate, pH value, conductivity, turbidity, and temperature corresponding to the same moment from the preprocessed sensor time-series data. The extracted sensor sub-data is used as an additional information channel and concatenated with the image frame data along the channel dimension to form the multimodal data for that moment. After the iteration is complete, a time-aligned multimodal data sequence is obtained.

[0036] In one possible implementation, in step S3, the floc detection model fuses features from various data points in the multimodal data using an attention mechanism, and generates a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. This includes: extracting semantic features from image frame data in the multimodal data using a visual backbone network within the model to obtain a corresponding semantic image feature map; encoding features from various sensor sub-data points in the multimodal data using a fully connected network within the model to obtain a corresponding sensor feature vector; and expanding the sensor feature vector to match the semantic map through spatial replication or broadcasting. The semantic image feature map and the sensor feature map have the same spatial dimensions as the feature map. An attention mechanism is used to fuse the semantic image feature map and the sensor feature map to generate a multimodal joint feature map. An internal detection head is used to perform target detection on the multimodal joint feature map to determine the position of each floc in the multimodal joint feature map. Based on the position of each floc, corresponding region features are cropped from the semantic image feature map, and pooling is performed on each region feature to obtain the morphological features of each floc. The morphological features are aggregated and statistically analyzed according to several preset index types to obtain the frame-level floc feature vector.

[0037] This application provides a method for floc detection using a model. It extracts semantic features from images through a visual backbone network, encodes sensor features through a fully connected network, and then uses an attention mechanism after spatial expansion to achieve interaction and fusion of image and sensor features. This method not only achieves deep fusion of multimodal information but also dynamically highlights key features through attention weights, thereby improving the accuracy and robustness of floc detection. Furthermore, the detection head locates the floc positions and extracts regional features from the feature map, which are then pooled and statistically aggregated to form frame-level feature vectors, ensuring that each frame of data represents the overall state of the flocs at the current moment. This design allows the model to simultaneously perceive the visual morphology of the flocs and the technological environment, providing rich and structured feature representations for subsequent time-series analysis.

[0038] In a preferred embodiment, the multimodal data obtained in step S2 are input one by one into the trained floc detection model. Within the model, a visual backbone network extracts the image semantic feature map, and a fully connected network encodes the sensor feature vector. Since the sensor feature vector and the image semantic feature map have different dimensions, spatial copying or broadcasting techniques are needed to expand the dimension of the sensor feature vector to obtain a sensor feature map with the same spatial dimension as the semantic image feature map. Then, the sensor feature map and the image feature map are fused using the SE attention module to generate a multimodal joint feature map. The detection head performs target detection on this feature map, outputting the bounding boxes (x, y, w, h) and confidence scores (conf) for all detected flocs in the current frame, where x and y represent the x and y coordinates of the bounding box center point, respectively, and w and h represent the width and height of the bounding box, respectively.

[0039] Existing object detection models typically output the final object detection bounding boxes and confidence scores. This application further adds a feature extraction module to the backend of the detection head to extract the morphological features of each floc from the semantic image feature map and perform aggregation statistics to obtain a frame-level floc feature vector. Specifically, firstly, based on the bounding box coordinates of each detected floc, the RoI Align operation is used to accurately crop the corresponding region's image features from the semantic image feature map output by the visual backbone network. Then, global average pooling (GAP) is performed on the region feature map of each floc to obtain a fixed-length feature vector (e.g., 512-dimensional), which represents the morphological features of a single floc. Extraction from the semantic feature map is chosen because it more purely preserves visual information, while the fused feature map has already been used for detection and may contain task bias.

[0040] Finally, the feature vectors of all detected flocs in the current frame are statistically aggregated according to the following preset metrics: Average Feature Vector: Calculate the mean of all floc feature vectors. Maximum Feature Vector: Calculate the maximum value of all floc feature vectors in each dimension. Number of Flocs: The total number of detected flocs. Average Confidence: The average confidence of all detection boxes. These statistics are then concatenated to form a comprehensive frame-level floc feature vector.

[0041] Furthermore, the wastewater treatment status monitoring method further includes: when obtaining the first semantic image feature map corresponding to the last moment in the first time period, inputting several sensor data corresponding to the last moment in the first time period and the current dosage of the flocculation tank into a preset physical information neural network, so that the physical information neural network calculates the theoretical number of flocs and the theoretical average volume of flocs based on the several sensor data and the current dosage; determining the corresponding observed number of flocs and the observed average volume of flocs based on the first semantic image feature map; calculating the first deviation between the theoretical number of flocs and the observed number of flocs; calculating the second deviation between the theoretical average volume of flocs and the observed average volume of flocs; determining whether there is a data acquisition anomaly in each sensor in the flocculation tank based on the first deviation and the second deviation; if it is determined that there is a data acquisition anomaly, correcting the time-series data of the several sensors based on several historical sensor data to obtain several corrected sensor time-series data, and regenerating each frame-level floc feature vector through the floc detection model based on each image frame data and each corrected sensor time-series data; generating sensor anomaly information and sending it to a designated device.

[0042] This application provides a method for verifying and correcting sensor data. Considering that in wastewater treatment scenarios, sensors are often immersed in complex, corrosive, and adhesive wastewater for extended periods, making them highly susceptible to biofouling, chemical scaling, or physical blockage, leading to measurement drift, decreased sensitivity, or even complete failure, this embodiment utilizes image data from the last moment to verify the data collected by the sensors. First, a physical information neural network is used to calculate the theoretical characteristics of flocs, and a first semantic image feature map is used to calculate the observed characteristics of the flocs. Then, by comparing the theoretical and observed characteristics of the flocs, it is determined whether the current sensor cluster is abnormal. If an abnormality is detected, the sensor data is corrected and the abnormality information is reported, further ensuring the accuracy of wastewater treatment status monitoring under abnormal operating conditions.

[0043] In a preferred embodiment, at the end of each 10-minute processing cycle, the system initiates a verification process. The semantic image feature map obtained from the floc detection model of the last frame of that cycle is input into a lightweight regression sub-network to quickly estimate the number of observed flocs in that frame. ) and average volume ( The volume is calculated using the equivalent diameter predicted by regression. Simultaneously, the sensor data [flow rate, pH, conductivity, turbidity, temperature] at the last moment of the cycle and the current dosage are input into a pre-trained Physical Information Neural Network (PINN). This PINN encodes constraints from flocculation kinetic equations (such as the Smoluchowski equation) and can output the theoretical amount of flocs that should form under the current conditions based on the input operating parameters. ) and theoretical average volume ( Considering the inherent uncertainties in sensor measurements and image analysis, the theoretical value should not be a fixed point, but rather a reasonable range (confidence interval). The deviation rate calculates the degree to which the observed value falls outside this range ("exceeding the limit"). In this embodiment, the first and second deviation rates are calculated using the following formulas:

[0044]

[0045] in, and These represent the theoretical number of flocs and the theoretical average volume calculated by PINN, respectively. and These represent the number of observed flocs and the average observed volume, respectively. This represents the standard deviation of the theoretical quantity prediction. This can be provided by PINN's uncertainty quantification module or derived statistically based on the error distribution between theoretical and actual values ​​in historical data. K is the confidence coefficient, typically taken as 2 (approximately 95% confidence interval) or 3 (approximately 99.7% confidence interval). This refers to the permissible relative error in volume. For example, This indicates that a reasonable fluctuation of ±25% in volume is allowed. This value can be determined based on process stability and image measurement accuracy. In the above formula, it is first determined whether the number of observations is within the reasonable fluctuation range of the theoretical number. If it is within the range, the deviation is 0, and the sensor data is considered reliable. If it is outside the range, the relative distance between the observed value and the nearest interval boundary is calculated as the deviation. At this time, it is necessary to determine whether N_obs or V_obs is less than or greater than the reasonable fluctuation range. If it is less than the reasonable fluctuation range, then the "" in the formula... "Take the minus sign; if it exceeds the reasonable fluctuation range, then the formula's 'minus sign'..." "Take the plus sign. This is actually a calculation." or The closest distance to the interval.

[0046] If the first deviation exceeds a first preset threshold and the second deviation exceeds a second preset threshold, then the sensor data acquisition is deemed abnormal. In this case, normal data from the same time period within the past 24 hours of the sensor cluster is used, and a pre-trained ARIMA time series model is used to predict the "normal" value for the current moment, which is then used to replace the abnormal sensor data.

[0047] Finally, the frame-level floc feature vectors are regenerated using the data from each image frame and the corrected sensor data. At the same time, an alarm message is sent to the host computer monitoring system, indicating the sensor ID and deviation of the suspected abnormality, and prompting maintenance personnel to check.

[0048] In one possible implementation, in step S4, the wastewater treatment status assessment model generates hidden state vectors at various time points based on the input data through a gating mechanism, and outputs the current wastewater treatment status based on each hidden state vector. This includes: concatenating each frame-level flocculent feature vector with the time-series data from the several sensors to obtain a time-series joint feature vector; extracting hidden state vectors from the joint feature vectors at different time points in the time-series joint feature vector sequentially through a gating unit within the model, based on the time order, to obtain each corresponding hidden state vector; during the extraction of the hidden state vectors, for any current joint feature vector, passing through a forget gate, an input gate, and an output gate respectively. The gates and memory units generate corresponding forgotten information, added information, output information, and candidate memory cells; the memory cells of the previous time step are updated according to the forgotten information, added information, and candidate memory cells to obtain the current memory cell; the hidden state vector of the current time step is generated according to the current memory cell and the hidden state vector of the previous time step; the attention weight of each hidden state vector is calculated, and the hidden state vectors are weighted and fused according to the attention weight to obtain the context vector; the probability distribution data of the sewage treatment status is generated according to the context vector through the fully connected classification head inside the model, and then the current sewage treatment status is output according to the probability distribution data.

[0049] This application provides a method for assessing the state of wastewater treatment using a model. The method progressively processes the temporal joint feature vector step-by-step through gating units. During processing, gating structures such as forget gates, input gates, and output gates are used to dynamically update the memory state, thereby capturing long-term dependencies in the sequence. Based on this, an attention mechanism is introduced to weightedly fuse the hidden states at each time step, generating a context vector that centrally reflects key temporal information. Finally, a fully connected classification head outputs the state probability distribution. This design enables the model to comprehensively consider data change trends over long periods, rather than just focusing on data at the current moment. It accurately identifies state transitions and abnormal patterns in the wastewater treatment process, significantly improving the temporal awareness and decision reliability of state assessment, and enhancing the accuracy and real-time performance of wastewater treatment state monitoring.

[0050] In a preferred embodiment, the frame-level floc feature vectors output by the floc detection model are concatenated with the previously acquired time-series data from each sensor along the feature dimension to form a final time-series joint feature vector sequence. This sequence is then input into a trained wastewater treatment status assessment model, which processes the sequence step-by-step over time.

[0051] For each time step t, the input x_t: Based on the hidden state h_{t-1} of the previous time step and the input x_t of the current time step, the forget gate determines which old information to discard from long-term memory, resulting in f_t. Based on the hidden state h_{t-1} of the previous time step and the input x_t of the current time step, the input gate determines which new information will be stored in long-term memory, resulting in i_t. Based on the hidden state h_{t-1} of the previous time step and the input x_t of the current time step, the output gate determines which information to output to the current hidden state h_t, resulting in o_t. Based on the hidden state h_{t-1} of the previous time step and the input x_t of the current time step, the memory cells generate possible new memory content C'_t for the current time step. Based on the forgotten information, the newly added information, and the candidate memory cells, the memory cells of the previous time step are updated to obtain the current memory cell: C_t = f_t ⊙ C_{t-1} + i_t ⊙ C'_t. Here, ⊙ represents element-wise multiplication. This formula allows the model to selectively forget old memories and selectively add new memories. Finally, based on the current memory cell and the hidden state vector of the previous time step, the hidden state vector of the current time step is generated: h_t = o_t ⊙ tanh(C_t), where tanh() represents the hyperbolic tangent activation function, which compresses any real number input into the interval (-1, 1).

[0052] After processing the entire sequence, T hidden state vectors {h_1, h_2, ..., h_T} are obtained, each encoding historical process information up to that moment. The attention weights α_i of these T hidden state vectors are calculated; moments with higher weights represent moments where the state at that moment has a greater impact on the final judgment. The weighted sum is used to obtain the context vector. This context vector c is input into a fully connected classifier head, and the Softmax function outputs a 4-dimensional probability distribution, such as [0.02, 0.15, 0.80, 0.03], representing the wastewater treatment status (excellent, good, moderate, poor), respectively. The system selects the category with the highest probability (here, "moderate") as the output result for the current wastewater treatment status.

[0053] Furthermore, the step of constructing the floc detection model based on the target detection model includes: constructing an initial floc labeling model based on the YOLOv8 model, wherein the initial floc labeling model includes a visual backbone network, a fully connected network, a dimension synchronization module, a feature fusion module, and a detection head; acquiring several historical multimodal data sets and manually labeling the floc positions in each historical multimodal data set to construct a first training dataset; training the initial floc labeling model using the first training dataset and updating each model parameter in the initial floc labeling model to obtain the floc labeling model; and combining the floc labeling model with a pre-trained feature extraction module to construct the floc detection model, wherein the feature extraction module is used to extract frame-level floc feature vectors from the corresponding semantic image feature map according to each floc position.

[0054] This application provides a method for constructing a floc detection model. Based on the advanced object detection framework YOLOv8, it incorporates a fully connected network, a dimension synchronization module, and a feature fusion module, enabling it to simultaneously process image and structured sensor data. By training with historical multimodal data and fusing a pre-trained feature extraction module, the model not only possesses strong generalization ability but can also be optimized for floc detection tasks, improving the accuracy of subsequent wastewater treatment status monitoring. It is important to note that the feature fusion module and the floc labeling model in this application are trained separately, providing a modular and phased model training approach. This modular and scalable design allows the model to quickly adapt to different sensor configurations and water treatment scenarios, facilitating practical engineering deployment.

[0055] In a preferred embodiment, an initial flocculent labeling model is constructed, using YOLOv8n as the visual backbone, with its CSPDarknet portion used to extract high-level semantic feature maps from the image. A new fully connected encoding network, consisting of two linear layers, is added to encode the input 5D sensor feature vector into a 256-dimensional feature vector. Then, a dimension synchronization module is designed to expand the encoded 256-dimensional sensor feature vector into a sensor feature map with the exact same spatial dimensions (e.g., [H / 32, W / 32, 256]) as the feature map output by the visual backbone network through a "copy-tiling" operation. The feature fusion module employs a Squeeze-and-Excitation (SE) attention mechanism. First, the visual feature map and the sensor feature map are concatenated along the channel dimension. Then, a global average pooling and two fully connected layers are used to generate channel attention weights. Finally, these weights are used to reweight the concatenated feature map to generate a multimodal joint feature map. Finally, the detection head uses the same detection head as YOLOv8, which is responsible for predicting the bounding box and confidence of the floc on the joint feature map.

[0056] We collected a large amount of video and synchronous sensor data from the factory's normal operation over the past three months, and used the LabelImg tool to perform fine-grained bounding box annotations on the flocs in the video keyframes. A total of 50,000 images were annotated to construct the first training dataset. The annotated data was then divided into training, validation, and test sets in an 8:1:1 ratio.

[0057] During training, the pre-trained weights of the visual backbone network are fixed. Initially, only the fully connected encoding network, dimension synchronization module, feature fusion module, and detection head are trained to learn how to utilize sensor information to assist in floc detection. After the loss function converges, the last few layers of the backbone network are unfrozen for end-to-end fine-tuning training. The entire training process uses the AdamW optimizer with an initial learning rate of 1e-3 and a cosine annealing strategy. The loss function is the standard classification + regression loss of YOLOv8. After training, a high-performance floc labeling model is obtained. This model is then combined with a pre-trained feature extraction module (composed of RoI Align layers and global average pooling layers) to form the final floc detection model. The final model structure is as follows: Figure 2 As shown.

[0058] Furthermore, the step of constructing the wastewater treatment status assessment model based on the recurrent neural network model includes: constructing an initial assessment model based on an LSTM model, wherein the initial assessment model includes a gating unit, a feature fusion unit, and a fully connected classification head; acquiring historical multimodal datasets for different time periods, and obtaining historical time-series joint feature vectors for different time periods through the floc detection model; manually labeling the wastewater treatment status corresponding to each of the historical time-series joint feature vectors according to the wastewater treatment effluent quality indicators for each corresponding time period, thereby constructing a second training dataset; and training the initial assessment model using the second training dataset, updating each model parameter in the initial assessment model, and obtaining the wastewater treatment status assessment model.

[0059] This application provides a training method for a wastewater treatment status assessment model. It utilizes a pre-trained floc detection model, generates historical time-series joint feature vectors for different time periods from a historical multimodal dataset, and constructs a second training dataset by combining manually labeled wastewater treatment status data. The second training dataset is then used to perform end-to-end training on the initial LSTM-based assessment model, enabling the model to learn the complex mapping relationship from multimodal time-series features to status labels. This process fully leverages the image features and feature association knowledge contained in historical data and sensor data, giving the model adaptability and status reasoning capabilities to different operating conditions. Therefore, in practical applications, it can provide accurate and reliable status assessment results, improving the accuracy of wastewater treatment status monitoring.

[0060] In a preferred embodiment, a two-layer stacked LSTM model is constructed as the initial evaluation model. Its input dimensions are the frame-level flocculent feature vector dimension (1026 dimensions) plus the sensor data dimension (5 dimensions), totaling 1031 dimensions. The hidden layer state dimension is set to 512. Finally, a fully connected classification head is connected, outputting four neurons corresponding to the four categories of wastewater treatment status: excellent, good, average, and poor.

[0061] During the training phase, historical data is used to generate a large number of continuous historical time-series joint feature vector sequences (each sequence corresponds to a 10-minute processing cycle) through a pre-trained floc detection model. Based on the effluent water quality (COD, ammonia nitrogen, TP, SS) measured after the corresponding time period, experts formulate rules to classify it into four categories: "excellent, good, medium, and poor," which serve as the label for that sequence. An example of a conventional water quality status classification rule that can be directly used to generate training labels is as follows:

[0062] This rule is based on four key effluent water quality indicators: Chemical Oxygen Demand (COD, mg / L), Ammonia Nitrogen (NH3-N, mg / L), Total Phosphorus (TP, mg / L), and Suspended Solids (SS, mg / L). The classification logic adopts a "single-factor veto system," meaning the final overall status level is determined based on the worst-case level for each indicator. Specific thresholds are set with reference to the Class A and Class B standards of the "Discharge Standard of Pollutants for Municipal Wastewater Treatment Plants" (GB 18918-2002), combined with actual process control objectives.

[0063] A second training dataset containing tens of thousands of sequence-label pairs was constructed. The LSTM model was then trained under supervision using this second training dataset, employing cross-entropy loss and the Adam optimizer. The training objective was to enable the model to learn to identify flocculation process state patterns leading to different effluent qualities from temporal dynamics, ultimately obtaining the desired wastewater treatment state assessment model.

[0064] In one possible implementation, the wastewater treatment status monitoring method further includes controlling the dosage of the flocculation tank in a second time period based on the current wastewater treatment status. Specifically, this involves: using linear extrapolation, predicting the estimated influent flow rate change trend and the estimated raw water turbidity change trend of the flocculation tank in the second time period using influent flow rate time-series data and raw water turbidity time-series data from several sensors; determining the corresponding dosing rule by querying a preset rule base based on the current wastewater treatment status, the estimated influent flow rate change trend, and the estimated raw water turbidity change trend, and determining the target dosage based on the dosing rule, wherein the dosing rule includes antecedent conditions and a conclusion, the antecedent conditions including wastewater treatment status, influent flow rate change trend, and raw water turbidity change trend, and the conclusion being the dosage; and controlling the dosage of the flocculation tank in the second time period based on the target dosage and the current dosage using a preset control strategy.

[0065] This embodiment further expands the dosing control function based on status monitoring. Based on the current wastewater treatment status and predicted influent flow rate and raw water turbidity trends, it determines the optimal dosing dosage through rule base queries and automatically adjusts the dosage using control strategies. This design forms a closed-loop control system of "sensing-analysis-decision-execution," achieving full automation from status monitoring to wastewater treatment control. By combining linear extrapolation prediction with rule-based reasoning, this embodiment can proactively address fluctuations in water quality and quantity, achieving precise dosing. While ensuring stable effluent quality, it significantly reduces reagent consumption and operating costs, thereby improving wastewater treatment efficiency.

[0066] In a preferred embodiment, the rule conclusion can also be directly designed as the target dosing frequency adjustment amount when designing the rule base based on historical data. Then, the dosing amount in the flocculation tank during the second time period is controlled by the target dosing frequency. Specifically, based on the time series data of influent flow rate and raw water turbidity within the current cycle (10 minutes), a linear regression extrapolation method is used to predict the changing trends of influent flow rate and raw water turbidity in the next 10-minute cycle, such as rising, falling, or remaining stable. Then, the current wastewater treatment status (e.g., "medium"), the predicted influent flow rate change trend (e.g., "rising"), and the predicted raw water turbidity change trend (e.g., "rising") are combined into a query condition and queried in the preset rule base. The preset expert rule base is stored in the form of "IF-THEN", one specific form being: IF{wastewater treatment status; influent flow rate change trend; raw water turbidity change trend}-THEN{target dosing pump frequency adjustment amount}. A complete rule base example is as follows:

[0067] Here, * represents a wildcard, matching any trend. If multiple rules match simultaneously during the matching process, the non-wildcard rule takes precedence.

[0068] Finally, based on the conclusions drawn from the matching rules, and according to the target dosing pump frequency adjustment amount, control commands are sent to the dosing pump PLC of the flocculation tank via an industrial communication protocol (such as Modbus) to automatically adjust the dosing amount, thereby achieving precise control of the flocculation process in the next cycle.

[0069] Example 2: Figure 3As shown in Embodiment 2, a wastewater treatment status monitoring system based on image feature analysis is provided, including an acquisition module 10, a data fusion module 20, a floc detection module 30, and a status assessment module 40. The acquisition module 10 acquires several image frame data and several sensor time-series data within a first time period from the flocculation tank. The first data fusion module 20 concatenates each of the image frame data and the several sensor time-series data according to time order to obtain several multimodal data at different times. The floc detection module 30 inputs each of the multimodal data into a preset floc detection model, so that the floc detection model uses an attention mechanism to select each floc from the multimodal data. The data is fused to generate a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. The floc detection model is constructed based on the target detection model. The state assessment module 40 is used to input each frame-level floc feature vector and the time-series data of several sensors into a preset wastewater treatment state assessment model, so that the wastewater treatment state assessment model generates hidden state vectors at each time step based on the input data through a gating mechanism, and outputs the current wastewater treatment state based on each hidden state vector. The hidden state vector is used to capture the dependency relationship between floc morphological features and sensor data features. The wastewater treatment state assessment model is constructed based on a recurrent neural network model.

[0070] In summary, this embodiment achieves precise, real-time perception and intelligent control of the flocculation process state through deep multimodal fusion and intelligent temporal analysis. Specifically, this embodiment first addresses the problems of image and sensor data separation and judgment lag in traditional methods by achieving millisecond-level multi-source data synchronization through timestamp alignment. Furthermore, it utilizes an improved YOLOv8 model combined with an attention mechanism to deeply fuse visual information and process parameters (such as flow rate, pH, and turbidity) at the feature level, significantly improving the accuracy and anti-interference capability of floc detection. Next, it employs an LSTM network to capture the temporal dependency between floc morphology features and sensor data, enabling dynamic and forward-looking assessment of the wastewater treatment state and overcoming the lag of relying on post-event judgments based on final effluent quality. In addition, this embodiment innovatively introduces a physical information neural network and an ARIMA model to construct a sensor data anomaly verification and correction mechanism, enhancing the system's robustness under complex and harsh operating conditions and improving the accuracy and real-time performance of wastewater treatment state monitoring. Ultimately, by combining the state assessment results with the influent trend prediction through the rule-based intelligent decision-making module, a closed-loop control of "perception-analysis-decision-execution" is formed, realizing adaptive and precise adjustment of the dosage. While ensuring stable effluent quality, it effectively reduces chemical consumption and operating costs, and comprehensively improves the intelligence level and operating efficiency of the sewage treatment process.

[0071] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.

Claims

1. A wastewater treatment status monitoring method based on image feature analysis, characterized in that, include: Acquire several image frame data and several sensor time series data of the flocculation tank in the first time period; Based on the time sequence, each of the image frame data and the time sequence data of the several sensors are respectively spliced ​​to obtain multimodal data at several different times; Each of the multimodal data is input into a preset floc detection model, so that the floc detection model performs feature fusion on each data in the multimodal data through an attention mechanism, and generates a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. The floc detection model is constructed based on the target detection model. Each frame-level floc feature vector and the time-series data from the several sensors are input into a preset wastewater treatment status assessment model. The wastewater treatment status assessment model generates hidden state vectors at each time step based on the input data through a gating mechanism, and outputs the current wastewater treatment status based on each hidden state vector. The hidden state vectors are used to capture the dependency relationship between floc morphology features and sensor data features. The wastewater treatment status assessment model is constructed based on a recurrent neural network model.

2. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, The acquisition of several image frame data and several sensor time-series data of the flocculation tank within a first time period includes: The flocculation tank is acquired using a pre-set industrial camera within a first time period, along with several initial image frames. The flocculation tank acquires several initial sensor time-series data within a first time period through a preset sensor cluster. The several initial sensor time-series data include influent flow rate time-series data, pH value time-series data, conductivity time-series data, raw water turbidity time-series data, and temperature time-series data. Denoising, enhancement, and color space conversion operations are performed sequentially on each of the initial image frame data to obtain each image frame data; The initial sensor time series data are sequentially filtered, outlier removal, missing value interpolation, and normalization to obtain the time series data for each sensor.

3. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, Based on time sequence, each of the image frame data and the time-series data from the plurality of sensors are concatenated to obtain multimodal data at several different times, including: The image frame data is traversed sequentially. For any image frame data, according to the timestamp of the image frame data, the corresponding sensor sub-data is extracted from the time series data of each sensor, and the sensor sub-data is concatenated with the image frame data to obtain the multimodal data with the corresponding timestamp. After the traversal is completed, multimodal data at several different times are obtained.

4. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, The floc detection model uses an attention mechanism to fuse features from various data points in the multimodal data, and generates a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result, including: Semantic features are extracted from the image frame data in the multimodal data through the visual backbone network inside the model to obtain the corresponding semantic image feature map. The model uses a fully connected network to encode the features of each sensor sub-data in the multimodal data, thereby obtaining the corresponding sensor feature vector. The sensor feature vector is expanded into a sensor feature map with the same spatial dimension as the semantic image feature map by spatial copying or broadcasting. The semantic image feature map and the sensor feature map are fused using an attention mechanism to generate a multimodal joint feature map; The model uses an internal detection head to perform target detection on the multimodal joint feature map, thereby determining the position of each floc in the multimodal joint feature map. Based on the position of each floc, the corresponding region features are cropped from the semantic image feature map, and pooling is performed on each region feature to obtain the morphological features of each floc. The frame-level flocculent feature vector is obtained by aggregating and statistically analyzing each of the morphological features according to a number of preset index types.

5. The wastewater treatment status monitoring method based on image feature analysis as described in claim 4, characterized in that, The wastewater treatment status monitoring method also includes: When the first semantic image feature map corresponding to the last moment in the first time period is obtained, several sensor data corresponding to the last moment in the first time period and the current dosage of the flocculation tank are input into a preset physical information neural network, so that the physical information neural network calculates the theoretical number of flocs and the theoretical average volume of flocs based on the several sensor data and the current dosage. Based on the first semantic image feature map, determine the corresponding number of observed flocs and the average volume of observed flocs; Calculate the first deviation between the theoretical number of flocs and the observed number of flocs; calculate the second deviation between the theoretical average volume of flocs and the observed average volume of flocs; Based on the first deviation and the second deviation, determine whether there is any abnormal data acquisition of each sensor in the flocculation tank; If an abnormal data acquisition is detected, the time-series data of the several sensors are corrected based on several historical sensor data to obtain several corrected sensor time-series data. Based on each of the image frame data and each of the corrected sensor time-series data, the floc detection model is used to regenerate each frame-level floc feature vector. Sensor abnormality information is generated and sent to the designated device.

6. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, The wastewater treatment status assessment model generates hidden state vectors at various time points based on the input data through a gating mechanism, and outputs the current wastewater treatment status based on each hidden state vector, including: Each frame-level floc feature vector is concatenated with the time-series data from the several sensors to obtain a time-series joint feature vector. Based on the time sequence, the hidden state vectors are extracted sequentially from the joint feature vectors at different times in the time sequence joint feature vectors through the gating unit inside the model, so as to obtain the corresponding hidden state vectors. During the extraction of the hidden state vector, for any current joint feature vector, corresponding forgotten information, newly added information, output information, and candidate memory cells are generated through the forget gate, input gate, output gate, and memory cell, respectively; the memory cells of the previous time step are updated according to the forgotten information, the newly added information, and the candidate memory cells to obtain the current memory cell; and the hidden state vector of the current time step is generated according to the current memory cell and the hidden state vector of the previous time step. Calculate the attention weights of each hidden state vector, and perform weighted fusion of each hidden state vector according to each attention weight to obtain the context vector; The model uses a fully connected classification head to generate probability distribution data of the wastewater treatment status based on the context vector, and then outputs the current wastewater treatment status based on the probability distribution data.

7. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, The process of constructing the floc detection model based on the target detection model includes: An initial floc labeling model is constructed based on the YOLOv8 model. The initial floc labeling model includes a visual backbone network, a fully connected network, a dimension synchronization module, a feature fusion module, and a detection head. Acquire several historical multimodal data sets and manually mark the floc positions in each historical multimodal data set to construct the first training dataset; The initial floc labeling model is trained using the first training dataset, and the various model parameters in the initial floc labeling model are updated to obtain the floc labeling model; The floc labeling model is combined with a pre-trained feature extraction module to construct the floc detection model, wherein the feature extraction module is used to extract frame-level floc feature vectors from the corresponding semantic image feature map according to each floc position.

8. The wastewater treatment status monitoring method based on image feature analysis as described in claim 1, characterized in that, The wastewater treatment status assessment model is constructed based on a recurrent neural network model, including: An initial evaluation model is constructed based on the LSTM model, which includes a gating unit, a feature fusion unit, and a fully connected classification head. Historical multimodal datasets from different time periods are acquired, and the joint feature vectors of historical time series from different time periods are obtained through the floc detection model. Based on the wastewater quality indicators of the wastewater treatment effluent in each corresponding time period, the wastewater treatment status corresponding to each of the historical time series joint feature vectors is manually marked, and then a second training dataset is constructed. The initial evaluation model is trained using the second training dataset, and the various model parameters in the initial evaluation model are updated to obtain the wastewater treatment status evaluation model.

9. A wastewater treatment status monitoring method based on image feature analysis as described in any one of claims 1-8, characterized in that, The wastewater treatment status monitoring method further includes controlling the dosage of chemicals in the flocculation tank during a second time period based on the current wastewater treatment status, specifically: Based on the linear extrapolation method, the predicted influent flow rate change trend and the predicted raw water turbidity change trend of the flocculation tank in the second time period are predicted by using the influent flow rate time series data and raw water turbidity time series data from the time series data of the aforementioned sensors. Based on the current wastewater treatment status, the estimated influent flow rate change trend, and the estimated raw water turbidity change trend, a query is performed in a preset rule base to determine the corresponding dosing rule, and the target dosing amount is determined according to the dosing rule. The dosing rule includes antecedent conditions and a conclusion. The antecedent conditions include the wastewater treatment status, the influent flow rate change trend, and the raw water turbidity change trend, and the conclusion is the dosing amount. Based on the target dosage and the current dosage, the dosage of the flocculation tank during the second time period is controlled by a preset control strategy.

10. A wastewater treatment status monitoring system based on image feature analysis, characterized in that, It includes an acquisition module, a data fusion module, a floc detection module, and a state assessment module; The acquisition module is used to acquire several image frame data and several sensor time series data of the flocculation tank in the first time period. The first data fusion module is used to concatenate each of the image frame data with the time-series data of the several sensors based on time order to obtain multimodal data at several different times. The floc detection module is used to input each of the multimodal data into a preset floc detection model, so that the floc detection model can perform feature fusion on each data in the multimodal data through an attention mechanism, and generate a frame-level floc feature vector corresponding to the multimodal data based on the feature fusion result. The floc detection model is constructed based on the target detection model. The state assessment module is used to input the frame-level floc feature vectors and the time-series data of the several sensors into a preset wastewater treatment state assessment model, so that the wastewater treatment state assessment model generates hidden state vectors at each time step according to the input data through a gating mechanism, and outputs the current wastewater treatment state according to each hidden state vector. The hidden state vectors are used to capture the dependency relationship between floc morphological features and sensor data features. The wastewater treatment state assessment model is constructed based on a recurrent neural network model.

Citation Information

Cited By

  • Flocculant residue influence assessment method and system based on multi-modal fusion

    CN121919820A

  • Method for dynamically detecting flocculation state in sewage

    CN122079332A