Activated sludge concentration detection method and system based on multi-modal modeling and atlas analysis

By combining multimodal modeling and graph analysis with three-dimensional sedimentation graphs and neural networks, the problems of lag and susceptibility to interference in existing MLSS detection are solved, enabling real-time, accurate, and low-cost detection of activated sludge concentration.

CN121661419APending Publication Date: 2026-03-13XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing MLSS detection technology suffers from problems such as lag, high cost, and susceptibility to interference, making it difficult to achieve real-time and accurate detection of activated sludge concentration.

Method used

A method based on multimodal modeling and graph analysis was adopted to obtain a three-dimensional sedimentation graph of the sludge mixed liquor sedimentation process. By combining the ResNeXt model and a fully connected neural network, visual features and sedimentation performance indicators were integrated to construct a fully connected neural network regression model for MLSS prediction.

Benefits of technology

It achieves real-time, accurate, and interference-resistant MLSS detection, reduces maintenance costs, improves detection accuracy and system stability, and adapts to different water quality and process conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661419A_ABST
    Figure CN121661419A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sewage treatment, and discloses an activated sludge concentration detection method and system based on multi-modal modeling and atlas analysis. The method comprises the following steps: continuously collecting activated sludge sedimentation process images through a digital camera, constructing a sedimentation map, and synchronously measuring SV5, SV30 and real MLSS values; a pre-trained network is utilized to extract visual features of three settlement sub-maps equally divided into five minutes, SV5 and SV30 are standardized and dimensionally raised into feature vectors, and a comprehensive feature vector is formed after cross-modal fusion; and finally, real-time accurate prediction of the MLSS is realized through a full-connection neural network regression model. The method effectively overcomes the problems that a traditional gravimetric method lags behind and an online sensor is prone to interference, has the advantages of being high in interference resistance, low in cost and convenient to deploy, and can be used for in-situ real-time intelligent detection of the sludge concentration of a sewage plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent detection technology for wastewater treatment, specifically relating to a method and system for detecting activated sludge concentration based on multimodal modeling and spectral analysis. Background Technology

[0002] As the core of modern wastewater treatment processes, the stable and efficient operation of the activated sludge process relies on the precise monitoring and control of key process parameters. Mixed Liquor Suspended Solids (MLSS) concentration is one of the most fundamental parameters, directly affecting microbial metabolic activity, sludge settling performance, and final effluent quality. Therefore, accurate and rapid detection of MLSS is crucial for process control.

[0003] Currently, MLSS measurement mainly relies on two types of technologies: offline laboratory testing and online automated monitoring. Laboratory testing typically uses the standard gravimetric method, which involves sampling, filtering, drying (to constant weight at 105°C), cooling, and then weighing. While this method is considered the benchmark, the process is cumbersome and time-consuming, with a significant lag of at least two hours, making it impossible to provide real-time feedback for process control. Furthermore, sample loss is easily caused during sampling, transfer, and drying, introducing unavoidable human error, which fails to meet the real-time data requirements of modern wastewater treatment plants.

[0004] To overcome the hysteresis problem of offline measurements, online sensors based on optical (e.g., turbidity, laser scattering) or electrochemical principles have emerged and been applied. These devices can be directly installed in reactors to achieve continuous measurement. However, their measurement performance heavily depends on the stability of the sensing element, posing significant challenges in practical applications. Optical sensors are susceptible to interference from water color, bubbles, suspended impurities, and probe surface contamination (scaling, biofilm adhesion); electrochemical sensors are sensitive to changes in water quality components such as coexisting ions and pH value. These factors cause measurement signal drift, requiring frequent manual on-site calibration to maintain accuracy. This not only incurs high maintenance costs but also generally lacks reliability and generalizability, exhibiting poor stability in wastewater treatment plants with different water qualities or processes.

[0005] In recent years, with the development of machine vision technology, some studies have emerged that attempt to use image processing techniques to assess sludge characteristics, such as estimating sludge concentration or settling performance through a single static image. However, most of these methods rely on image information at a single moment, failing to fully utilize the dynamic sequence characteristics of the settling process. They have a limited information dimension, making it difficult to cope with the complex and ever-changing actual aquatic environment, and their prediction accuracy and stability need to be improved.

[0006] Therefore, existing technologies are either severely outdated, susceptible to interference, have high maintenance costs, or fail to fully utilize available information. Developing a new MLSS detection method that is real-time, accurate, interference-resistant, and low-cost has become a key technical challenge that urgently needs to be addressed to improve the intelligent operation level of wastewater treatment plants. Summary of the Invention

[0007] The purpose of this invention is to solve the problem that the visual method of MLSS detection in the prior art has limited information and is difficult to detect accurately in real time, and to provide a method and system for detecting activated sludge concentration based on multimodal modeling and spectrum analysis.

[0008] To achieve the above objectives, the present invention employs the following technical solution: This invention proposes a method for detecting activated sludge concentration based on multimodal modeling and spectral analysis, comprising the following steps: The sludge-liquid mixture settling ratio was processed to obtain a three-dimensional settling map reflecting the sludge interface settling process; The three-dimensional settlement map is divided into sub-maps according to the time dimension. The sub-maps are then input into the ResNeXt model to extract visual feature vectors. A comprehensive feature vector is obtained based on the sludge mixed liquor settling ratio and visual feature vector; A fully connected neural network regression model is constructed based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

[0009] Preferably, the step of processing the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process specifically involves: Three-dimensional sedimentation maps were obtained, and the 5-minute sedimentation ratio SV5 and 30-minute sedimentation ratio SV5 of the sludge mixture sample at the sampling time were determined. 30 The true MLSS concentration of the sludge mixed liquor sample was determined using the standard gravimetric method. The three-dimensional sedimentation map was preprocessed, including background subtraction and grayscale conversion, and then stitched together in chronological order to form a three-dimensional sedimentation map reflecting the sludge interface sedimentation process.

[0010] Preferably, the step of dividing the three-dimensional settlement map into sub-maps according to the time dimension, and then inputting the sub-maps into the ResNeXt model to extract visual feature vectors, specifically involves: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

[0011] Preferably, the step of obtaining the comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector specifically involves: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

[0012] Preferably, the step of constructing a fully connected neural network regression model based on the comprehensive feature vector specifically involves: The fully connected neural network regression model consists of one input layer, two hidden layers, and one output layer. The hidden layers use the ReLU activation function to introduce a nonlinear transformation, while the output layer uses a linear activation function for regression prediction. A Dropout layer is added to the network to suppress overfitting. The nonlinear transformation is performed through the hidden layers to output the predicted value of MLSS.

[0013] This invention proposes an activated sludge concentration detection system based on multimodal modeling and spectral analysis, comprising: The data preprocessing module is used to process the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process. The feature extraction module is used to divide the three-dimensional settlement map into sub-maps according to the time dimension, and input the sub-maps into the ResNeXt model to extract visual feature vectors. A feature fusion module is used to obtain a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector; The sludge concentration detection module is used to construct a fully connected neural network regression model based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

[0014] Preferably, the step of dividing the three-dimensional settlement map into sub-maps according to the time dimension, and then inputting the sub-maps into the ResNeXt model to extract visual feature vectors, specifically involves: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

[0015] Preferably, the step of obtaining the comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector specifically involves: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

[0016] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of an activated sludge concentration detection method based on multimodal modeling and spectral analysis.

[0017] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of an activated sludge concentration detection method based on multimodal modeling and spectral analysis.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention proposes a method for detecting activated sludge concentration based on multimodal modeling and spectral analysis. By acquiring a three-dimensional sedimentation map of the sludge mixed liquor settling process, it first integrates spatial features and temporal dynamic information to overcome the limitations of traditional visual methods that only capture information at a single moment, enriching the data dimensions to adapt to complex aquatic environments. Second, the three-dimensional map is divided into sub-maps according to time and input into the ResNeXt model. Its grouped convolutional structure enhances the anti-interference ability of feature extraction, accurately capturing subtle changes during the sedimentation process and avoiding interference from water impurities and bubbles. Third, a comprehensive feature vector is constructed by fusing visual feature vectors with key sedimentation performance indicators (SV5, SV30), achieving complementarity between visual information and core process parameters, further improving the comprehensiveness and accuracy of feature representation. Finally, a fully connected neural network regression model is used to analyze the comprehensive features, eliminating the need for cumbersome offline sampling and drying steps, and outputting MLSS prediction values ​​in real time. This solves the problem of insufficient accuracy in traditional visual methods, meets the requirements of real-time detection, and eliminates the need for frequent manual calibration, reducing maintenance costs and achieving accurate, real-time, and interference-resistant MLSS detection.

[0019] This invention proposes an activated sludge concentration detection system based on multimodal modeling and spectral analysis. By dividing the system into a data preprocessing module, a feature extraction module, a feature fusion module, and a sludge concentration detection module, it obtains MLSS predicted values ​​to achieve activated sludge concentration detection. The modular approach ensures that each module is independent, facilitating unified management of all modules. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of the activated sludge concentration detection method based on multimodal modeling and spectral analysis of the present invention.

[0022] Figure 2 This is a detailed flowchart of the activated sludge concentration detection method based on multimodal modeling and spectral analysis of the present invention.

[0023] Figure 3 This is a schematic diagram of the MLSS intelligent detection model of the present invention.

[0024] Figure 4 This refers to the mean square error (MSE) and loss changes during the training process of the intelligent detection model of this invention.

[0025] Figure 5 The test set of this invention represents the model performance evaluation results under the prediction of this model.

[0026] Figure 6 This is a diagram of the activated sludge concentration detection system based on multimodal modeling and spectral analysis of the present invention.

[0027] Figure 7 This is a schematic diagram of the structure of an electronic device according to the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0029] The present invention will now be described in further detail with reference to the accompanying drawings: The purpose of this invention is to overcome the shortcomings of existing activated sludge concentration (MLSS) detection technologies, such as lag, high cost, susceptibility to interference, and reliance on manual calibration, and to provide an intelligent activated sludge concentration detection method based on multimodal modeling and spectral analysis. This method integrates visual sequence information of the sludge settling process (settling map) and key settling performance indicators (SV5, SV...). 30 This invention proposes a method for MLSS detection based on multimodal modeling and settlement mapping, which constructs a deep learning model to achieve fast, accurate, stable, and low-cost real-time online prediction of MLSS. Figure 1 As shown, it includes the following steps: S1. Process the sludge mixture sedimentation ratio to obtain a three-dimensional sedimentation map reflecting the sludge interface sedimentation process; The process of processing the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process is as follows: Three-dimensional sedimentation maps were obtained, and the 5-minute sedimentation ratio SV5 and 30-minute sedimentation ratio SV5 of the sludge mixture sample at the sampling time were determined. 30 The true MLSS concentration of the sludge mixed liquor sample was determined using the standard gravimetric method. The three-dimensional sedimentation map was preprocessed, including background subtraction and grayscale conversion, and then stitched together in chronological order to form a three-dimensional sedimentation map reflecting the sludge interface sedimentation process.

[0030] S2. Divide the three-dimensional settlement map into sub-maps according to the time dimension, and input the sub-maps into the ResNeXt model to extract visual feature vectors; The process involves dividing the three-dimensional settlement map into sub-maps based on the time dimension, and then inputting these sub-maps into the ResNeXt model to extract visual feature vectors. Specifically: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

[0031] S3. Obtain the comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector; The method for obtaining a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vectors is as follows: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

[0032] S4. Construct a fully connected neural network regression model based on the comprehensive feature vector, input the sludge data to be analyzed into the fully connected neural network regression model, and output the MLSS prediction value to realize the detection of activated sludge concentration.

[0033] The construction of a fully connected neural network regression model based on the comprehensive feature vector is specifically as follows: The fully connected neural network regression model consists of one input layer, two hidden layers, and one output layer. The hidden layers use the ReLU activation function to introduce a nonlinear transformation, while the output layer uses a linear activation function for regression prediction. A Dropout layer is added to the network to suppress overfitting. The nonlinear transformation is performed through the hidden layers to output the predicted value of MLSS.

[0034] Detailed flowchart as follows Figure 2 As shown, the model structure of the MLSS intelligent detection model is as follows: Figure 3 As shown, the training process and test results of the model are as follows: Figure 4 and Figure 5As shown, the specific steps are as follows: Step 1: Synchronous acquisition of multimodal data and construction of settlement maps This step aims to provide a high-quality data foundation for model training and prediction. The specific process is as follows: A digital camera is fixedly deployed at a specific location in the activated sludge reactor, and controlled to continuously photograph the settling sludge at fixed time intervals for N minutes to obtain X sequential images, forming a three-dimensional settling map. Simultaneously, at the sampling time, the 5-minute settling ratio (SV5) and 30-minute settling ratio (SV) of the mixed liquor sample are measured. 30 The true MLSS concentration was determined using the standard gravimetric method and used as a label for supervised learning. The images were then preprocessed, including background subtraction and grayscale conversion, and stitched together in chronological order to form a three-dimensional sedimentation map that fully reflects the sludge interface sedimentation process.

[0035] The digital camera takes one photo every 10 seconds for 15 minutes to obtain a sequence of 90 images.

[0036] The method for determining SV is as follows: During the aeration stage of the reactor, take 100 mL of the sludge-water mixture from the reactor and place it in a 100 mL graduated cylinder. Let it stand and settle. After standing for 5 minutes, read the sludge volume as SV5. After standing for 30 minutes, read the sludge volume as SV. 30 .

[0037] Determination of mixed liquor suspended solids (MLSS): Preheat quantitative filter paper to a constant weight in an oven at 103-105 °C for 2 hours. Record this weight as m0 (mg). Filter 100 mL of the sludge-water mixture sample using a Buchner funnel, ensuring the funnel walls are rinsed with pure water to transfer the sludge completely onto the filter paper. Place the filter paper and the activated sludge solids on it in an oven for 2 hours. After drying, cool the filter paper in a desiccator and weigh it using an analytical balance, recording the weight as m1 (mg). Calculate the MLSS (mg / L) based on m0 - m1.

[0038] During the data acquisition phase, this study prioritized three laboratory-scale sequencing batch reactors (SBRs), numbered R1 to R3, and set different aeration rates to simulate the operation of activated sludge in an actual wastewater treatment plant. The activated sludge samples used were taken from the biological reactor of a wastewater treatment plant in Xi'an.

[0039] The aeration rates for R1 to R3 are set to 0.5, 1, and 2 L / S, respectively.

[0040] The three reactors are identical in specifications, each with an effective volume of 3 liters, a height of 1000 mm, and an inner diameter of 60 mm. In terms of design, the inlet is cleverly positioned on the side wall, the aeration head is installed at the bottom, and the outlet is strategically located in the middle.

[0041] The operating cycle of this invention is set to 6 hours, and each cycle includes the following five stages in sequence: static influent, anaerobic digestion, aeration, sedimentation, and drainage. The time allocation for each stage is shown in Table 1.

[0042] During influent intake, a peristaltic pump injects simulated wastewater into the reactor at a uniform rate through the bottom inlet, and constant-speed stirring ensures even water distribution. During aeration, an air compressor and aeration sand heads provide the air supply, with the aeration rate adjusted in real-time by a glass rotor flow meter. During effluent discharge, the supernatant is discharged at a 50% replacement rate through the outlet located 500 mm above the bottom of the reactor, controlled by the opening and closing time of the drain valve, and collected in a waste liquid tank.

[0043] The aforementioned equipment, including the peristaltic pump, drain valve, air compressor, and agitator, is all operated automatically via a time controller, thereby ensuring the precision and stability of the entire operation process.

[0044] Table 1 Reactor operating parameters / min

[0045] Preferably, synthetic wastewater was used as the reaction substrate in this study. The carbon, nitrogen, and phosphorus sources were provided by sodium acetate, ammonium sulfate, and potassium dihydrogen phosphate, respectively, to ensure a sufficient supply of nutrients required for microbial metabolism. The concentrations of carbon (C), nitrogen (N), and phosphorus (P) in the wastewater were precisely maintained within the set ranges; the specific ratios are shown in Table 2.

[0046] Table 2. Influent Input Parameter Range for the Reactor

[0047] Specifically, the ranges of MLSS and SV indices measured in this experiment are shown in Table 3.

[0048] Table 3 Measurement range of MLSS and SV indices

[0049] Step 2: Cross-modal feature extraction and fusion This step is crucial for the model to perceive and understand multi-source information. First, the aforementioned settlement map is divided into three 5-minute sub-maps (before, during, and after the initial settlement phase) along the time dimension. Each sub-map is then input into a pre-trained ResNeXt network (with the top-level classifier removed) to extract three 512-dimensional high-order visual feature vectors, capturing the pattern features of different settlement stages. Simultaneously, numerical indices SV5 and SV... 30 After Z-score standardization, the feature vectors were upscaled to 512 dimensions using a multilayer perceptron (MLP) to align with the dimensions of the visual feature vectors. Finally, a feature concatenation strategy was employed to fuse the three visual feature vectors with the two numerical feature vectors into a single 2560-dimensional comprehensive feature vector, thereby comprehensively characterizing the dynamic and static properties of sludge settling.

[0050] Step 21. Divide the 3D settlement map generated in Step 1 into three independent sub-maps along the time axis, corresponding to the first 5 minutes (0-5 min), the middle 5 minutes (5-10 min), and the last 5 minutes (10-15 min) of the settlement process, respectively. Adjust each sub-map to a fixed size of 224×224 pixels and perform normalization processing to make it meet the input requirements of the pre-trained model.

[0051] Step 22. Visual feature extraction uses a ResNeXt-34 model pre-trained on the ImageNet dataset. The top-level global average pooling layer and classifier are removed, leaving only the last convolutional layer (conv5_x). Each sub-map is forward-propagated through this network to extract its output feature map, which is then subjected to global average pooling (GAP) to obtain three 512-dimensional high-order visual feature vectors V_f1, V_f2, and V_f3, representing the visual dynamic characteristics of the early, middle, and late stages of settlement, respectively.

[0052] The ResNeXt-34 model used consists of 34 trainable layers, mainly comprising four core parts: an initial layer, residual block stacking, skip connections, and a regression layer. The model first processes the three input sedimentation maps through the initial layer. Then, during the four stages of residual block stacking, the number of channels is gradually increased while the size of the feature maps is reduced. The skip connection design effectively alleviates the vanishing and exploding gradient problems.

[0053] Step 23. Numerical feature processing stage: The measured settlement indices SV5 and SV... 30 Z-score standardization is performed, and the calculation formula is as follows:

[0054] Specifically, μ and σ are SV5 and SV6 in the dataset, respectively. 30The mean and standard deviation of the values ​​are then input into a two-layer fully connected network (MLP) with a hidden layer dimension of 256 and a ReLU activation function. The output layer dimension is 512. Finally, each numerical index is upscaled into a 512-dimensional feature vector V_s1 and V_s2 to align with the dimension of the visual feature vector.

[0055] S24. A concatenation method is used for cross-modal feature fusion, merging the five 512-dimensional feature vectors [V_f1, V_f2, V_f3, V_s1, V_s2] into a single 2560-dimensional comprehensive feature vector. This vector integrates multi-time-period settlement visual information and key process indicators, forming a high-representation input for the subsequent regression model.

[0056] The feature fusion strategy of this invention adopts a feature-level fusion method, that is, after extracting features of different modalities, they are immediately concatenated into a high-dimensional feature vector.

[0057] It should be noted that, in order to improve feature quality, a fully connected layer can be introduced after fusion for dimensionality reduction and feature compression, and a Dropout layer (e.g., with a dropout rate set to 0.5) can be used to suppress overfitting and further enhance the model's generalization ability.

[0058] Step 3: Construction, training, and deployment of regression models based on MLSS real-time monitoring This step completes the accurate mapping from features to concentration and enables practical application. Using the 2560-dimensional comprehensive feature vector obtained in the second step as input, a fully connected neural network regression model is constructed. Nonlinear transformations are performed through hidden layers, ultimately outputting the predicted MLSS value. Model training employs an optimizer and a mean squared error loss function. The dataset is divided into training, validation, and test sets, and an early stopping strategy is used to prevent overfitting and achieve optimal generalization performance. Finally, the trained model is deployed to the control system of an actual wastewater treatment plant. The same preprocessing and feature fusion process is performed on newly acquired data in real time, enabling real-time, online, and intelligent detection of activated sludge concentration.

[0059] The learning rate adjustment strategy of this invention is as follows: 1. Monitoring metric: Validation loss is used as the monitoring metric; 2. Adjustment condition: When the validation loss does not improve within 10 consecutive epochs (patience=10); 3. Adjustment method: Multiply the learning rate by 0.5 (factor=0.5), i.e., halve it; 4. Adjustment mode: When the monitoring metric no longer decreases (mode='min').

[0060] Step 31. Regression Model Structure Design: This fully connected neural network consists of an input layer (2560 neurons), two hidden layers (containing 1024 and 512 neurons respectively), and an output layer (1 neuron). The hidden layers use the ReLU activation function to introduce a nonlinear transformation, while the output layer uses a linear activation function for regression prediction. A Dropout layer (with a dropout rate set to 0.3) is added to the network to further suppress overfitting.

[0061] Step 32. Model Training Strategy: The AdamW optimizer is used, with an initial learning rate of 0.001 and mean squared error (MSE) as the loss function. The dataset is divided into training, validation, and test sets in a 70%:15%:15% ratio. The batch size is set to 8, and the number of training epochs is set to 1000. The patience value of the early stopping strategy is set to 30, meaning that training automatically terminates and the model parameters are restored to the minimum level when the validation set loss does not decrease for 30 consecutive training epochs.

[0062] Step 33. Model Evaluation and Optimization: The mean squared error (MSE) and coefficient of determination (R²) are used as evaluation metrics on the test set to comprehensively assess model performance. Hyperparameter tuning is performed based on the validation set performance, and a learning rate decay strategy is employed if necessary to improve model convergence stability.

[0063] Specifically, the final trained MLSS multimodal detection model achieved the following performance: MSE = 0.126, R² = 0.916. The training process and test results are as follows: Figure 4 and Figure 5 As shown.

[0064] Specifically, Figure 4 In the diagram, (a) shows the change in the loss function value during the training process of the MLSS prediction model. Figure 4 (b) shows the change in MSE value during the training process of the MLSS prediction model.

[0065] Step 34. System Deployment and Real-time Monitoring: Save the trained optimal model in TensorFlowSavedModel format and deploy it on the computer of the central control system of the wastewater treatment pilot plant. In practical applications, the system automatically executes the data acquisition, preprocessing, and feature fusion processes in steps S1 and S2, inputting the generated 2560-dimensional feature vector into the loaded model and outputting the MLSS prediction value in real time. This system can be integrated with existing monitoring systems to achieve continuous automatic monitoring of MLSS, historical data query, out-of-limit alarms, and process control feedback.

[0066] Step 35. Model Update and Maintenance: The system is designed with a regular model update mechanism. Every certain period, the model is fine-tuned using newly collected data to ensure that the model can adapt to changes in water quality and process adjustments, and maintain long-term prediction accuracy.

[0067] Every six months, the model is fine-tuned using newly collected data.

[0068] In summary, compared with other technologies, this invention overcomes the sludge loss problem during the measurement process, unlike traditional methods (based on sludge weight after two hours in an oven). Unlike optical and electrochemical instruments that rely on calibration lines, this invention primarily uses sludge settling process maps for measurement, unaffected by external environmental factors (such as water color and charged impurities). Therefore, this invention has significant advantages in generalization ability and accuracy. Furthermore, this method mainly utilizes ordinary digital cameras to acquire settling maps, resulting in lower costs and convenient application to in-situ real-time automatic detection of sludge concentration in actual wastewater treatment plants, contributing to improved plant intelligence.

[0069] Beneficial effects include: 1) Real-time and fast MLSS detection was achieved, overcoming the serious lag of the traditional gravimetric method.

[0070] This invention can instantly obtain results through model calculation after a 15-minute settling experiment. Compared with the national standard gravimetric method, which requires at least 2 hours of drying process, it greatly shortens the detection time and provides data support for real-time process control of wastewater treatment plants.

[0071] 2) It has excellent anti-interference and generalization capabilities, solving the problem that online sensors are susceptible to environmental interference.

[0072] The core of this invention is to analyze the dynamic behavior spectrum of sludge settling, rather than the optical or electrochemical properties of water. Therefore, it fundamentally avoids measurement errors caused by common factors such as water color, bubbles, probe contamination, and charged ions in the water, and can maintain stable and reliable measurement performance in wastewater treatment plants with different water qualities and process conditions.

[0073] 3) Significantly reduced hardware costs and maintenance complexity.

[0074] The main data acquisition device of this invention is a common digital camera, which does not rely on expensive and delicate dedicated optical or electrochemical sensors, resulting in extremely low hardware costs. At the same time, the method avoids frequent manual cleaning and calibration operations, reducing subsequent maintenance costs and manpower input.

[0075] 4) Improved measurement accuracy and reliability.

[0076] By integrating dynamic sequence image information (settlement atlas) of the entire settlement process and key settlement performance indicators (SV5, SV... 30 The multimodal model constructed in this invention utilizes richer and more representative information dimensions. Compared with methods based on a single static image or a single sensor signal, it can more comprehensively and accurately characterize sludge characteristics, thereby achieving higher-precision MLSS prediction.

[0077] 5) Easy to deploy and implement, effectively improving the intelligence level of sewage treatment plants.

[0078] This method requires no large-scale modification to existing wastewater treatment structures; it can be deployed in situ simply by adding cameras and edge computing devices. It provides an efficient and low-cost solution for continuous, automated, and online monitoring of MLSS, helping water plants transform from a traditional model relying on manual experience to a data-driven intelligent operation model.

[0079] Example 2 This invention proposes an activated sludge concentration detection system based on multimodal modeling and spectral analysis, such as... Figure 6 As shown, it includes: The data preprocessing module is used to process the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process. The feature extraction module is used to divide the three-dimensional settlement map into sub-maps according to the time dimension, and input the sub-maps into the ResNeXt model to extract visual feature vectors. The process involves dividing the three-dimensional settlement map into sub-maps based on the time dimension, and then inputting these sub-maps into the ResNeXt model to extract visual feature vectors. Specifically: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

[0080] A feature fusion module is used to obtain a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector; The method for obtaining a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vectors is as follows: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

[0081] The sludge concentration detection module is used to construct a fully connected neural network regression model based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

[0082] The construction of a fully connected neural network regression model based on the comprehensive feature vector is specifically as follows: The fully connected neural network regression model consists of one input layer, two hidden layers, and one output layer. The hidden layers use the ReLU activation function to introduce a nonlinear transformation, while the output layer uses a linear activation function for regression prediction. A Dropout layer is added to the network to suppress overfitting. The nonlinear transformation is performed through the hidden layers to output the predicted value of MLSS.

[0083] Example 3 Please see Figure 7 As shown, the present invention also provides an electronic device 100 for a method of detecting activated sludge concentration based on multimodal modeling and spectral analysis; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0084] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the activated sludge concentration detection method based on multimodal modeling and spectral analysis described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0085] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.

[0086] The memory 101 in the electronic device 100 stores multiple instructions to implement a method for detecting activated sludge concentration based on multimodal modeling and spectral analysis. The processor 102 can execute the multiple instructions to achieve the following: The sludge-liquid mixture settling ratio was processed to obtain a three-dimensional settling map reflecting the sludge interface settling process; The three-dimensional settlement map is divided into sub-maps according to the time dimension. The sub-maps are then input into the ResNeXt model to extract visual feature vectors. A comprehensive feature vector is obtained based on the sludge mixed liquor settling ratio and visual feature vector; A fully connected neural network regression model is constructed based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

[0087] Example 4 If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0088] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0089] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0090] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0091] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting activated sludge concentration based on multimodal modeling and spectral analysis, characterized in that, Includes the following steps: The sludge-liquid mixture settling ratio was processed to obtain a three-dimensional settling map reflecting the sludge interface settling process; The three-dimensional settlement map is divided into sub-maps according to the time dimension. The sub-maps are then input into the ResNeXt model to extract visual feature vectors. A comprehensive feature vector is obtained based on the sludge mixed liquor settling ratio and visual feature vector; A fully connected neural network regression model is constructed based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

2. The activated sludge concentration detection method based on multimodal modeling and spectral analysis according to claim 1, characterized in that, The process of processing the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process is as follows: Three-dimensional sedimentation maps were obtained, and the 5-minute sedimentation ratio SV5 and 30-minute sedimentation ratio SV5 of the sludge mixture sample at the sampling time were determined. 30 The true MLSS concentration of the sludge mixed liquor sample was determined using the standard gravimetric method. The three-dimensional sedimentation map was preprocessed, including background subtraction and grayscale conversion, and then stitched together in chronological order to form a three-dimensional sedimentation map reflecting the sludge interface sedimentation process.

3. The activated sludge concentration detection method based on multimodal modeling and spectral analysis according to claim 1, characterized in that, The process involves dividing the three-dimensional settlement map into sub-maps based on the time dimension, and then inputting these sub-maps into the ResNeXt model to extract visual feature vectors. Specifically: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

4. The activated sludge concentration detection method based on multimodal modeling and spectral analysis according to claim 1, characterized in that, The method for obtaining a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vectors is as follows: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

5. The activated sludge concentration detection method based on multimodal modeling and spectral analysis according to claim 1, characterized in that, The construction of a fully connected neural network regression model based on the comprehensive feature vector is specifically as follows: The fully connected neural network regression model consists of one input layer, two hidden layers, and one output layer. The hidden layers use the ReLU activation function to introduce a nonlinear transformation, while the output layer uses a linear activation function for regression prediction. A Dropout layer is added to the network to suppress overfitting. The nonlinear transformation is performed through the hidden layers to output the predicted value of MLSS.

6. An activated sludge concentration detection system based on multimodal modeling and spectral analysis, characterized in that, include: The data preprocessing module is used to process the sludge mixture settling ratio to obtain a three-dimensional settling map reflecting the sludge interface settling process. The feature extraction module is used to divide the three-dimensional settlement map into sub-maps according to the time dimension, and input the sub-maps into the ResNeXt model to extract visual feature vectors. A feature fusion module is used to obtain a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vector; The sludge concentration detection module is used to construct a fully connected neural network regression model based on the comprehensive feature vector. The sludge data to be analyzed is input into the fully connected neural network regression model, and the MLSS prediction value is output to realize the detection of activated sludge concentration.

7. The activated sludge concentration detection system based on multimodal modeling and spectral analysis according to claim 6, characterized in that, The process involves dividing the three-dimensional settlement map into sub-maps based on the time dimension, and then inputting these sub-maps into the ResNeXt model to extract visual feature vectors. Specifically: The three-dimensional settlement map was evenly divided into three independent sub-maps along the time axis, corresponding to the first 5 minutes, the middle 5 minutes and the last 5 minutes of the settlement process, respectively, and then normalized. The ResNeXt-34 model was adopted. The top-level global average pooling layer and classifier were removed, and the remaining layer was the last convolutional layer. Each sub-map was forward-propagated through the network to extract the output feature map and perform global average pooling to obtain the visual feature vector. The visual feature vector includes three visual feature vectors V_f1, V_f2, and V_f3, which represent the visual dynamic characteristics of the early, middle and late stages of settlement, respectively.

8. The activated sludge concentration detection system based on multimodal modeling and spectral analysis according to claim 6, characterized in that, The method for obtaining a comprehensive feature vector based on the sludge mixed liquor settling ratio and visual feature vectors is as follows: Sludge mixture sedimentation ratio includes the 5-minute sedimentation ratio (SV5) and the 30-minute sedimentation ratio (SV) of the sludge mixture sample at the sampling time. 30 , The sludge-liquid mixture settling ratio was standardized by Z-score. The two standardized scalars were then input into a two-layer fully connected network to upgrade each numerical index into a high-dimensional feature vector V_s1 and V_s2. Visual feature vectors and high-dimensional feature vectors are concatenated to perform cross-modal feature fusion and obtain a comprehensive feature vector.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the activated sludge concentration detection method based on multimodal modeling and spectral analysis as described in any one of claims 1 to 5.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the activated sludge concentration detection method based on multimodal modeling and spectral analysis as described in any one of claims 1 to 5.