Sewage detection and analysis method based on the Internet of Things
By deploying sensors and cloud platforms in the sewage treatment system and combining feature engineering and anomaly detection algorithms to optimize traditional models, the problem of traditional models' insufficient capture of temporal and spatial changes has been solved, and the intelligent control capabilities and prediction accuracy of the sewage treatment system have been improved.
Patent Information
- Application Number
- CN202510698078.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-28
AI Technical Summary
When faced with pollutants with strong temporal and spatial variations, existing technologies and traditional models are unable to effectively capture dynamic changes, resulting in large deviations in pollutant concentration predictions, affecting the real-time adjustment and optimization of sewage treatment.
A variety of sensors are deployed at key locations in the sewage treatment system to collect pollutant concentration and environmental data in real time. The data is then transmitted to the cloud platform via wireless communication technology for preprocessing and feature engineering. The polynomial regression model and OCSVM anomaly detection algorithm are combined to optimize traditional machine learning models and evaluate and improve their accuracy in capturing dynamic changes in the spatiotemporal dimensions.
It improves the model's ability to identify and predict complex pollution behaviors, enhances the intelligent regulation and control capabilities of the sewage treatment system, and achieves accurate and efficient sewage quality management.
Smart Images

Figure CN120217125B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sewage detection, and in particular to a sewage detection and analysis method based on the Internet of Things. Background Art
[0002] Wastewater testing and analysis involves the detection and analysis of wastewater composition, pollutants, and their concentrations through various physical, chemical, and biological methods. The goal is to assess wastewater quality, determine whether it contains harmful substances such as organic matter, heavy metals, and bacteria, and determine the potential impacts of these pollutants on the environment and public health. Wastewater testing and analysis not only helps governments and businesses develop pollution control measures but also provides a basis for improvements in wastewater treatment technology, ensuring that wastewater discharges meet environmental standards.
[0003] The existing technology has the following shortcomings:
[0004] Existing technologies use big data analysis, machine learning, and artificial intelligence algorithms to conduct in-depth analysis of sewage testing data, identify the main pollutants in sewage, and predict changing trends in pollutants. However, although machine learning algorithms can predict pollutant concentrations using historical data, when faced with pollutants that have strong spatiotemporal variations, the model may not be able to effectively capture this complex spatiotemporal dependency. For example, the concentration of certain pollutants may change more slowly at night due to temperature drops and biodegradation, while increasing rapidly during the day due to increased industrial activity. Traditional models may not be able to capture this dynamic change well, resulting in large deviations in the prediction of pollutant concentrations, which in turn affects the real-time adjustment and optimization of sewage treatment. Summary of the Invention
[0005] The purpose of the present invention is to provide a sewage detection and analysis method based on the Internet of Things to address the shortcomings of the background technology.
[0006] In order to achieve the above objectives, the present invention provides the following technical solutions: a sewage detection and analysis method based on the Internet of Things, comprising:
[0007] Deploy a variety of sensors at key locations in the sewage treatment system to collect real-time data on pollutant concentrations and environmental conditions in sewage, and transmit the data to the cloud platform in real time via wireless communication technology;
[0008] The cloud platform performs preprocessing and feature engineering on the collected raw data, including extracting the difference characteristics of pollutant concentrations that vary day and night, as well as the interaction characteristics between pollutants;
[0009] Conduct a comprehensive analysis of pollutant concentration difference characteristics and interaction characteristics to evaluate the accuracy of traditional machine learning models in capturing dynamic changes in pollutant concentrations in time and space;
[0010] Based on the evaluation results, the capture accuracy levels are divided into accurate capture, incomplete accuracy capture, and inaccuracy capture. The performance of the incomplete accuracy capture traditional machine learning model is optimized through anomaly detection algorithms.
[0011] Based on the optimized traditional machine learning model, the pollutant concentration in sewage is predicted.
[0012] Preferably, the sensors include: chemical oxygen demand sensor, ammonia nitrogen sensor, heavy metal sensor, pH sensor, dissolved oxygen sensor, and temperature and humidity sensors; key positions include water inlet, discharge port, different treatment tanks and water outlet.
[0013] Preferably, the concentration change rate fluctuation value is generated after analyzing the concentration change rate in the extracted pollutant concentration difference characteristics of day and night changes. The generation method is:
[0014] Calculate the concentration change rate , which represents the relative change in the average concentration of a pollutant between daytime and nighttime, and is expressed as: ; is the average concentration of pollutants during the daytime period, is the average concentration of pollutants during the night time period; ϵ is the minimum constant;
[0015] The continuous monitoring time is divided into N consecutive natural days, and a concentration change rate is extracted every day. , where i=1,2,...,N; a simplified fluctuation model is used to define the concentration change rate fluctuation value as the normalized average deviation of the change rate within N days, and the formula is: ; is the concentration change rate fluctuation value.
[0016] Preferably, after analyzing the ratio terms in the extracted interaction characteristics between pollutants, relative ratio anomalies between pollutants are generated, and the generation method is:
[0017] Suppose that M ratio features are monitored and S historical data samples are provided to form an S×M feature matrix X, where each row is the ratio vector of the i-th data;
[0018] Calculate the overall mean and covariance matrix of the ratio vector, the mean vector : ; Covariance matrix Σ: ; T is the matrix transpose, for any new pollutant ratio sample vector , its Mahalanobis distance The calculation formula is: ; Set the abnormality judgment threshold W, if >W, the sample is a proportion outlier and is recorded as the relative proportion outlier among pollutants.
[0019] Preferably, the concentration change rate fluctuation values and the relative proportion anomalies between pollutants are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as inputs of the machine learning model. The machine learning model uses each set of comprehensive feature vectors to predict the accuracy value labels of the traditional machine learning model for capturing dynamic changes in pollutant concentrations in the time and space dimensions as the prediction target, and takes minimizing the sum of the prediction errors of all traditional machine learning models for capturing accuracy value labels for dynamic changes in pollutant concentrations in the time and space dimensions as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence and the model training is stopped. The accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentrations in the time and space dimensions is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0020] Preferably, the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension is compared with a gradient accuracy threshold, the gradient accuracy threshold includes a first accuracy threshold and a second accuracy threshold, and the first accuracy threshold is less than the second accuracy threshold, and the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension is compared with the first accuracy threshold and the second accuracy threshold respectively;
[0021] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than the second accuracy threshold, it is marked as accurate capture and can be directly used for prediction;
[0022] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than or equal to the first accuracy threshold and less than or equal to the second accuracy threshold, it is marked as incomplete accuracy capture and the traditional machine learning model is optimized;
[0023] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentrations in the temporal and spatial dimensions is less than the first accuracy threshold, it will be marked as inaccurate capture and the prediction result will not be used.
[0024] Preferably, samples whose accuracy of the traditional machine learning model is between the first accuracy threshold and the second accuracy threshold are extracted to form a sample set to be optimized. , which includes the concentration change rate fluctuation value and the relative proportion abnormal value between pollutants Two characteristic dimensions;
[0025] use Training the OCSVM model:
[0026] Using the RBF kernel, specify the abnormal sample ratio ν = 0.05, that is, a maximum of 5% of the samples are allowed to be abnormal; the input dimension is: ;
[0027] Use the trained OCSVM model to Classify the samples in and output labels: ;
[0028] Construct the optimized sample set: ;
[0029] After removing the outliers Re-evaluate the performance of traditional machine learning models in capturing the dynamics of pollutant concentrations:
[0030] Calculate the error score. If the error decreases, it means that OCSVM effectively removes noise samples.
[0031] If the capture accuracy value of the optimized sample is greater than the second accuracy threshold, the traditional model is upgraded from incomplete accuracy to accurate;
[0032] If there is only a slight improvement but it is still between the first accuracy threshold and the second accuracy threshold, it is retained in the state of awaiting auxiliary optimization.
[0033] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0034] 1. This invention deploys multiple types of sensors at key locations within the sewage treatment system, combining wireless communications with cloud-based data processing to achieve real-time collection and efficient analysis of pollutant concentrations and environmental parameters. Leveraging feature engineering techniques, the system extracts differences in pollutant concentrations across the diurnal cycle and the interactions between pollutants. It then constructs concentration rate fluctuations and relative proportional anomalies, which are then integrated into a comprehensive feature vector to evaluate the accuracy of traditional machine learning models in addressing the spatiotemporal dynamics of pollutant concentrations.
[0035] 2. By introducing a model evaluation mechanism based on polynomial regression and an OCSVM anomaly detection optimization strategy, this invention enables hierarchical assessment and adaptive optimization of the accuracy of traditional models, thereby enhancing the model's ability to identify and predict complex pollution behaviors. This approach not only improves the model's stability and prediction accuracy in dynamic environments, but also enhances the intelligent control capabilities of sewage treatment systems, providing a new technical path for accurate, efficient, and interpretable sewage quality prediction and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0037] Figure 1 This is a mind map of the method of the present invention. DETAILED DESCRIPTION
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] For examples, see Figure 1 As shown, the sewage detection and analysis method based on the Internet of Things described in this embodiment includes:
[0040] Deploy a variety of sensors at key locations in the sewage treatment system to collect real-time data on pollutant concentrations and environmental conditions in sewage, and transmit the data to the cloud platform in real time via wireless communication technology;
[0041] The cloud platform performs preprocessing and feature engineering on the collected raw data, including extracting the difference characteristics of pollutant concentrations that vary day and night, as well as the interaction characteristics between pollutants;
[0042] Conduct a comprehensive analysis of pollutant concentration difference characteristics and interaction characteristics to evaluate the accuracy of traditional machine learning models in capturing dynamic changes in pollutant concentrations in time and space;
[0043] Based on the evaluation results, the capture accuracy levels are divided into accurate capture, incomplete accuracy capture, and inaccuracy capture. The performance of the incomplete accuracy capture traditional machine learning model is optimized through anomaly detection algorithms.
[0044] Based on the optimized traditional machine learning model, the pollutant concentration in sewage is predicted.
[0045] According to the different links of sewage treatment and target pollutants, appropriate sensors are selected for installation. Common sensor types include:
[0046] Chemical oxygen demand sensor: used to measure the concentration of organic pollutants in water and assess the degree of water pollution.
[0047] Ammonia nitrogen sensor: detects the ammonia nitrogen concentration in sewage, commonly used in industrial wastewater treatment and sewage treatment plants.
[0048] Heavy metal sensor: detects harmful heavy metal elements in water, such as lead, mercury, arsenic, etc.
[0049] pH sensor: used to measure the pH value of sewage. Too high or too low pH value will affect the efficiency of sewage treatment.
[0050] Dissolved oxygen sensor: used to measure the dissolved oxygen concentration in water, which is especially important in the biological treatment stage.
[0051] Temperature and humidity sensors: provide environmental parameters to help analyze the relationship between changes in sewage and the external environment.
[0052] Depending on the sewage treatment process, sensors are typically installed at key locations such as the water inlet, discharge outlet, various treatment tanks (such as sedimentation tanks and aeration tanks), and outlet. A rational layout ensures comprehensive collection of key sewage data, reflecting dynamic changes during the treatment process.
[0053] Sensors continuously monitor various indicators in the sewage in real time during system operation. The sensors regularly collect data and convert it into digital signals for subsequent processing and analysis.
[0054] The data collection frequency for different sensors should be set based on the characteristics of the sewage treatment process. For example, COD and ammonia nitrogen concentrations change slowly, so a lower collection frequency can be set; whereas indicators such as dissolved oxygen and pH change rapidly, requiring a higher collection frequency.
[0055] In order to ensure that data can be stably transmitted to the cloud platform, wireless communication technology suitable for sewage treatment environments is adopted. Common communication technologies include:
[0056] LoRa (Long Range Low Power Wireless Communication): Suitable for scenarios with long-distance transmission and low power consumption requirements, and is suitable for large-scale sewage treatment facilities.
[0057] NB-IoT (Narrowband Internet of Things): has a wide coverage range and low power consumption, suitable for widely distributed sensor nodes.
[0058] Wi-Fi: Suitable for short-distance data transmission and suitable for small sewage treatment facilities with centralized equipment.
[0059] Sensors transmit collected data to a data acquisition gateway via wireless communication modules. The gateway integrates the data from each sensor and transmits it to a cloud platform or local server via wireless networks. This ensures real-time data availability to meet the needs of wastewater treatment monitoring and analysis.
[0060] Before performing feature engineering, the raw data collected by the sensor must be preprocessed to ensure data quality and analyzability:
[0061] Denoising: Use algorithms such as moving average filtering and wavelet denoising to eliminate high-frequency fluctuations or invalid noise in sensor data.
[0062] Outlier detection and repair: Use the Isolation Forest or Z-score algorithm to detect extreme values and perform smoothing replacement or interpolation based on data from adjacent time periods.
[0063] Missing value filling: Recover missing data based on linear interpolation, K-nearest neighbor interpolation (KNN Imputation), or time series-based filling algorithms.
[0064] Normalize the data output by different types of sensors (such as COD in mg / L and pH as dimensionless) (such as Min-Max or Z-score normalization) to ensure that all features in the machine learning model have the same scale, thereby improving convergence speed and prediction performance.
[0065] The data collected by different sensors are aligned according to timestamps to ensure the consistency of multi-source data at the same time point to support the subsequent construction of spatiotemporal features.
[0066] The goal of feature engineering is to extract the most valuable information for model prediction from the cleaned raw data, with particular attention paid to the day-night variation differences of pollutants and the interaction characteristics between pollutants.
[0067] The concentration of pollutants in sewage shows regular differences during the day and at night due to factors such as human activities, biodegradation, and temperature changes. Extracting these features helps the model identify pollution trends related to time periods.
[0068] Label each data item as "daytime" or "nighttime" based on local sunrise and sunset times or unified settings (e.g., 06:00–18:00 is daytime and 18:00–06:00 is nighttime).
[0069] The following features are calculated within a sliding time window (e.g., every day or every two hours):
[0070] Daytime mean concentration (D-Mean);
[0071] Nighttime mean concentration (N-Mean);
[0072] Day-night difference (DN): D-Mean - N-Mean;
[0073] Day / night ratio (D / N Ratio): D-Mean / (N-Mean + ε);
[0074] Concentration change rate: |(D-Mean - N-Mean) / N-Mean|;
[0075] Time series decomposition methods (such as STL decomposition) are used to extract the trend, cycle and residual components from the original concentration series and identify diurnal periodicity.
[0076] Reflecting the temporal pattern of factory pollution discharge or changes in biological reaction activity helps the model make accurate judgments on abnormal daytime / nighttime changes.
[0077] The concentration change rate in the extracted pollutant concentration difference characteristics of day and night is analyzed to generate the concentration change rate fluctuation value. The generation method is as follows:
[0078] Calculate the concentration change rate , which represents the relative change in the average concentration of a pollutant between daytime and nighttime, and is expressed as: ; is the average concentration of pollutants during the daytime (unit: mg / L), is the average concentration of pollutants during the nighttime period (unit: mg / L); ϵ is a minimum constant to avoid division by zero.
[0079] The continuous monitoring time is divided into N consecutive natural days, and a concentration change rate is extracted every day. , where i=1,2,...,N; a simplified fluctuation model is used to define the concentration change rate fluctuation value as the normalized average deviation of the change rate within N days, and the formula is: ; is the concentration change rate fluctuation value.
[0080] Larger fluctuation values indicate unstable changes in pollutant concentrations between day and night, potentially leading to potential risks such as abnormal discharge and equipment interference. Smaller fluctuation values indicate more consistent concentration trends between day and night, helping the model to stably learn temporal patterns.
[0081] The pollutants in sewage do not exist independently. There may be complex chemical, biological or physical reaction relationships between them. Exploring these potential interactions will help build more accurate prediction models.
[0082] For each pollutant (such as COD, ammonia nitrogen, total phosphorus, heavy metals), combine them in pairs and calculate their synergistic change index:
[0083] Pearson correlation coefficient (Pearson);
[0084] Spearman rank correlation (Spearman);
[0085] Mutual Information
[0086] Construct the following cross features:
[0087] Product term (X×Y): such as COD×ammonia nitrogen, reflects possible common upward / downward trends.
[0088] Difference term (X−Y): Determine the concentration of the dominant pollutant.
[0089] Ratio term (X / (Y+ε)): Determines the relative proportional relationship between pollutants and avoids division by zero.
[0090] Logical items: For example, whether "high COD and low ammonia nitrogen" corresponds to a specific pollution event.
[0091] Using dimensionality reduction methods such as principal component analysis (PCA) and autoencoders, multiple interactive features are merged into nonlinear representations to extract potential complex pollution factors.
[0092] If the amount of data is sufficient, graph structure data can be constructed (pollutants are nodes and interaction intensity is edges) and input into the graph neural network (GNN) for modeling.
[0093] It helps the model identify common fluctuation patterns and abnormal reactions among pollutants (such as a sudden surge in a certain type of heavy metal causing rapid changes in COD), and enhances the model's ability to identify the complex behavior of pollution sources.
[0094] After analyzing the ratio terms in the interaction characteristics between the extracted pollutants, the relative proportion anomalies between the pollutants are generated. The generation method is:
[0095] Set to monitor M ratio characteristics (such as COD / NH3-N, COD / TP, etc.).
[0096] Suppose there are S historical data samples, forming an S×M feature matrix X, where each row is the ratio vector of the i-th data.
[0097] Calculate the overall mean and covariance matrix of the ratio vector, the mean vector (dimension is M): ; Covariance matrix Σ (dimension M×M): ; T is the matrix transpose, for any new pollutant ratio sample vector , its Mahalanobis distance The calculation formula is: ; Set the abnormality judgment threshold W. Common methods include: ; Use the 97.5% quantile of the chi-square distribution as the threshold (where M is the ratio dimension) or determine the threshold based on the distribution experience of historical data. >W, the sample is a ratio outlier and is recorded as the relative ratio outlier value between pollutants. A larger relative ratio outlier value between pollutants may indicate system abnormality or sudden emissions.
[0098] The concentration change rate fluctuation values and the relative proportion anomalies between pollutants are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as the input of the machine learning model. The machine learning model uses each set of comprehensive feature vectors to predict the accuracy value label of the traditional machine learning model for capturing the dynamic changes of pollutant concentrations in the temporal and spatial dimensions. The prediction target is to minimize the sum of the prediction errors of all traditional machine learning models for capturing the accuracy value label of the dynamic changes of pollutant concentrations in the temporal and spatial dimensions as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence and the model training is stopped. The accuracy value of the traditional machine learning model for capturing the dynamic changes of pollutant concentrations in the temporal and spatial dimensions is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0099] Comparing the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension with a gradient accuracy threshold, where the gradient accuracy threshold includes a first accuracy threshold and a second accuracy threshold, and the first accuracy threshold is less than the second accuracy threshold, and comparing the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension with the first accuracy threshold and the second accuracy threshold respectively;
[0100] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than the second accuracy threshold, it is marked as accurate capture and can be directly used for prediction;
[0101] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than or equal to the first accuracy threshold and less than or equal to the second accuracy threshold, it is marked as incomplete accuracy capture and the traditional machine learning model is optimized;
[0102] If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentrations in the temporal and spatial dimensions is less than the first accuracy threshold, it will be marked as inaccurate capture and the prediction result will not be used.
[0103] Extract samples whose accuracy of the traditional model is between the first accuracy threshold and the second accuracy threshold from the comprehensive feature evaluation system to form a sample set to be optimized , which includes the concentration change rate fluctuation value and the relative proportion abnormal value between pollutants Two feature dimensions.
[0104] use Training the OCSVM model:
[0105] It uses RBF (Radial Basis Function) kernel, which is suitable for nonlinear boundary detection.
[0106] Specify the proportion of abnormal samples (for example, ν=0.05, which means that at most 5% of the samples are allowed to be abnormal).
[0107] The input dimensions are: ;
[0108] Use the trained OCSVM model to Classify the samples in and output labels: ;
[0109] Construct the optimized sample set: .
[0110] After removing the outliers Re-evaluate the performance of traditional machine learning models in capturing the dynamics of pollutant concentrations:
[0111] Calculate the error score (such as RMSE, MAE). If the error decreases, it means that OCSVM effectively removes noise samples and improves the model stability and predictive power.
[0112] If the capture accuracy value of the optimized sample is greater than the second accuracy threshold, the traditional model can be upgraded from "incompletely accurate" to "accurate".
[0113] If there is only a slight improvement but it is still between the first accuracy threshold and the second accuracy threshold, it can be retained in the state of waiting for auxiliary optimization, and it is recommended to introduce an integrated algorithm or feature expansion.
[0114] After optimizing the traditional machine learning model (including using anomaly detection algorithms such as One-Class SVM to eliminate error-prone samples and improve model generalization), the optimized model can be used to make more accurate and robust predictions of the concentrations of various pollutants in wastewater. The following is a detailed description of the prediction process:
[0115] In the sewage treatment system, the data collected in real time include:
[0116] Various pollutant sensor data (such as COD, ammonia nitrogen, total phosphorus, heavy metals, etc.);
[0117] Environmental parameters (such as temperature, humidity, pH);
[0118] Derived characteristics (such as pollutant ratios, concentration change rates, flow fluctuations, etc.);
[0119] Feature engineering output (such as concentration change rate fluctuation value Vol_Δ, pollutant relative proportion anomaly value RPA, etc.).
[0120] After preprocessing (standardization, cleaning, and completion), these data constitute the feature vector input set of the model.
[0121] Load optimized traditional machine learning models (such as random forest, support vector machine, XGBoost, etc.);
[0122] The model has been trained on historical sewage data and cleaned sample sets, and its adaptability to actual application scenarios has been enhanced through anomaly elimination.
[0123] The model has been integrated with the ability to handle key features, especially the dynamic response capability when pollutant concentrations change with time, weather, process and other conditions.
[0124] At each moment or sampling period, the system uses the real-time collected data as input and feeds it into the optimized model;
[0125] The model outputs the predicted value of the target pollutant concentration at the corresponding time point. For example:
[0126] COD predicted concentration: 48.2 mg / L; ammonia nitrogen predicted concentration: 8.6 mg / L; total phosphorus predicted concentration: 1.25 mg / L;
[0127] The prediction result can be a single point prediction or the concentration trend for a certain time period (such as the prediction curve for the next 1 hour, 6 hours or 24 hours), depending on the model design (such as whether a time series modeling structure is introduced).
[0128] Based on the prediction results, determine whether it is necessary to adjust the sewage treatment process parameters, such as aeration time, dosage, coagulant type, etc.
[0129] If the predicted concentration of a certain type of pollutant is about to exceed the emission standard, the system will automatically trigger an early warning signal to alert the on-duty personnel or automatically activate the emergency treatment device.
[0130] Compare the predicted results with historical concentrations to analyze concentration change trends, assist in operational decision-making, and automatically generate daily and weekly analysis reports.
[0131] In the intelligent sewage management system, the prediction results can be used as one of the input parameters to trigger subsequent control logic (such as linkage valve control system, booster pump start-stop system, etc.).
[0132] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0133] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0134] It should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. In addition, the character " / " herein generally indicates that the objects associated with each other are in an "or" relationship, but it may also indicate an "and / or" relationship, which can be understood by referring to the context. A person of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0135] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A sewage detection and analysis method based on the Internet of Things, characterized by: include: Deploy a variety of sensors at key locations in the sewage treatment system to collect real-time data on pollutant concentrations and environmental conditions in sewage, and transmit the data to the cloud platform in real time via wireless communication technology; The cloud platform performs preprocessing and feature engineering on the collected raw data, including extracting the difference characteristics of pollutant concentrations that vary day and night, as well as the interaction characteristics between pollutants; The pollutant concentration difference characteristic of the day and night changes is the concentration change rate fluctuation value, and the interaction characteristic between pollutants is the relative proportion abnormal value between pollutants; The method for generating abnormal values of relative proportions between pollutants is as follows: set M ratio features to be monitored, set S historical data samples, and form an S×M feature matrix X, where each row is the ratio vector of the i-th data; calculate the overall mean and covariance matrix of the ratio vector, the mean vector : ; Covariance matrix Σ: ; T is the matrix transpose, for any new pollutant ratio sample vector , its Mahalanobis distance The calculation formula is: ; Set the abnormality judgment threshold W, if >W, the sample is a ratio outlier, recorded as the relative ratio outlier between pollutants; The concentration change rate in the extracted pollutant concentration difference characteristics of day and night is analyzed to generate the concentration change rate fluctuation value. The generation method is as follows: Calculate the concentration change rate , which represents the relative change in the average concentration of a pollutant between daytime and nighttime, and is expressed as: ; is the average concentration of pollutants during the daytime. is the average concentration of pollutants during the night time period; ϵ is the minimum constant; the continuous monitoring time is divided into N consecutive natural days, and a concentration change rate is extracted every day , where i=1,2,...,N; a simplified fluctuation model is used to define the concentration change rate fluctuation value as the normalized average deviation of the change rate within N days, and the formula is: ; is the concentration change rate fluctuation value; Conduct a comprehensive analysis of pollutant concentration difference characteristics and interaction characteristics to evaluate the accuracy of traditional machine learning models in capturing dynamic changes in pollutant concentrations in time and space; Based on the evaluation results, the capture accuracy levels are divided into accurate capture, incomplete accuracy capture, and inaccuracy capture. The performance of the incomplete accuracy capture traditional machine learning model is optimized through anomaly detection algorithms. Based on the optimized traditional machine learning model, the pollutant concentration in sewage is predicted.
2. The sewage detection and analysis method based on the Internet of Things according to claim 1, characterized in that: The sensors include: chemical oxygen demand sensor, ammonia nitrogen sensor, heavy metal sensor, pH sensor, dissolved oxygen sensor, and temperature and humidity sensors; key locations include water inlet, discharge outlet, different treatment pools and water outlet.
3. The sewage detection and analysis method based on the Internet of Things according to claim 1, characterized in that: The concentration change rate fluctuation values and the relative proportion anomalies between pollutants are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as the input of the machine learning model. The machine learning model uses each set of comprehensive feature vectors to predict the traditional machine learning model's capture accuracy value label of the dynamic changes in pollutant concentration in the temporal and spatial dimensions as the prediction target, and takes minimizing the sum of prediction errors of all capture accuracy value labels as the training target. The machine learning model is trained until the sum of prediction errors reaches convergence, and the model training is stopped. The capture accuracy value of the traditional machine learning model of the dynamic changes in pollutant concentration in the temporal and spatial dimensions is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
4. The sewage detection and analysis method based on the Internet of Things according to claim 3 is characterized in that: Comparing the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension with a gradient accuracy threshold, where the gradient accuracy threshold includes a first accuracy threshold and a second accuracy threshold, and the first accuracy threshold is less than the second accuracy threshold, and comparing the acquired accuracy value of the traditional machine learning model for capturing dynamic changes in pollutant concentration in the spatiotemporal dimension with the first accuracy threshold and the second accuracy threshold respectively; If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than the second accuracy threshold, it is marked as accurate capture and can be directly used for prediction; If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentration in the spatiotemporal dimension is greater than or equal to the first accuracy threshold and less than or equal to the second accuracy threshold, it is marked as incomplete accuracy capture and the traditional machine learning model is optimized; If the accuracy value of the traditional machine learning model in capturing the dynamic changes of pollutant concentrations in the temporal and spatial dimensions is less than the first accuracy threshold, it will be marked as inaccurate capture and the prediction result will not be used.
5. The sewage detection and analysis method based on the Internet of Things according to claim 4 is characterized in that: Extract samples whose accuracy of the traditional machine learning model is between the first accuracy threshold and the second accuracy threshold to form a sample set to be optimized , which includes the concentration change rate fluctuation value and the relative proportion abnormal value between pollutants Two characteristic dimensions; use Training the OCSVM model: Using the RBF kernel, specify the abnormal sample ratio ν = 0.05, that is, a maximum of 5% of the samples are allowed to be abnormal; the input dimension is: ; Use the trained OCSVM model to Classify the samples in and output labels: ; Construct the optimized sample set: ; After removing the outliers Re-evaluate the performance of traditional machine learning models in capturing the dynamics of pollutant concentrations: Calculate the error score. If the error decreases, it means that OCSVM effectively removes noise samples. If the capture accuracy value of the optimized sample is greater than the second accuracy threshold, the traditional model is upgraded from incomplete accuracy to accurate; If there is only a slight improvement but it is still between the first accuracy threshold and the second accuracy threshold, it is retained in the state of awaiting auxiliary optimization.
Citation Information
Patent Citations
Pollutant emission prediction method and system based on machine learning
CN119940632A
Method for Detecting Validity of Human Body Movement
US20240197207A1