Water quality intelligent prediction method based on rank correlation and XGBoost model

By analyzing multi-dimensional data from industrial parks using Spearman rank correlation coefficient and XGBoost model, key characteristic variables were selected, and an intelligent prediction model for influent water quality was constructed. This solved the problem of lag in traditional water quality monitoring methods and enabled efficient and accurate water quality prediction and management.

CN119558481BActive Publication Date: 2026-05-29CHONGQING UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional water quality monitoring methods are outdated in industrial park wastewater treatment plants, making it difficult to respond effectively to sudden changes in water quality in real time, which affects the compliance rate of effluent water quality and the ecological environment quality of the park.

Method used

Spearman rank correlation coefficient analysis was used to analyze the correlation between multi-dimensional data and influent water quality, key feature variables were screened, and an intelligent prediction model for influent water quality was constructed based on the XGBoost model to achieve accurate and timely prediction.

Benefits of technology

It improves the accuracy of influent water quality prediction, supports scientific decision-making in smart industrial parks, optimizes resource allocation, reduces operating costs, and promotes green development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558481B_ABST
    Figure CN119558481B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent industrial park management, and discloses a water quality intelligent prediction method based on rank correlation and an XGBoost model, which comprises the following steps: a data collection step, which collects multi-dimensional data of an industrial park and pre-processes the data; a correlation analysis step, which analyzes the correlation between the multi-dimensional data and influent water quality by using a Spearman rank correlation coefficient; a key feature variable screening step, which screens multi-dimensional data with correlation meeting preset standards as key feature variables based on the analysis result of the correlation; an intelligent prediction model construction step, which normalizes the key feature variables and constructs an influent water quality intelligent prediction model based on an XGBoost method; and a prediction step, which performs real-time prediction based on the influent water quality intelligent prediction model. The application realizes accurate and timely prediction of influent water quality of an industrial park sewage treatment plant by analyzing multi-dimensional correlation data of the industrial park sewage treatment plant and establishing an intelligent prediction model, improves sewage treatment efficiency, and reduces operation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart industrial park management technology, specifically to a water quality intelligent prediction method based on rank correlation and XGBoost model. Background Technology

[0002] With the deep integration of industrialization and informatization, smart industrial parks have become an important vehicle for promoting industrial transformation and upgrading. By integrating advanced information technology, the Internet of Things, big data analytics, and artificial intelligence, smart industrial parks optimize resource allocation, improve operational efficiency, reduce environmental pollution, and promote sustainable development. As a key facility for ensuring the environmental quality of the industrial park, the intelligent prediction and management of influent water quality is particularly important for wastewater treatment plants within the park.

[0003] The industrial park houses numerous enterprises spanning multiple industries, including chemicals, electronics, and textiles. The complex and ever-changing wastewater pollution and discharge patterns pose significant challenges to the efficient and stable operation of wastewater treatment plants. Wastewater treatment is a dynamic system, and traditional water quality monitoring methods often suffer from lag, making it difficult to respond effectively to sudden changes in water quality in real time, thus affecting the effluent quality compliance rate and the ecological environment quality of the park. The influent quality of wastewater treatment plants is influenced by various factors, including weather, industrial production, population, and changes in economic activity. These influencing factors exhibit complex, multi-dimensional correlations with water quality, which traditional methods struggle to accurately capture and effectively predict using. With the rapid development of big data technology and the widespread application of artificial intelligence algorithms, it is possible to deeply mine multi-dimensional correlated data and establish accurate influent water quality prediction models, providing strong support for the stable operation and efficient management of wastewater treatment plants. Summary of the Invention

[0004] This invention aims to provide a water quality intelligent prediction method based on rank correlation and XGBoost model. By collecting and analyzing multi-dimensional correlation data of industrial park wastewater treatment plants, an intelligent prediction model is established based on the correlation between multi-dimensional data and influent pollutants. This enables accurate and timely prediction of influent water quality of industrial park wastewater treatment plants, thereby improving wastewater treatment efficiency and reducing operating costs.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] Intelligent water quality prediction methods based on rank correlation and XGBoost models include:

[0007] The data collection process involves collecting multi-dimensional data from the industrial park and preprocessing the data.

[0008] The correlation analysis step uses Spearman's rank correlation coefficient to analyze the correlation between multi-dimensional data and influent water quality;

[0009] The key feature variable screening step involves selecting multi-dimensional data that meet preset criteria based on the correlation analysis results as key feature variables.

[0010] The steps for building an intelligent prediction model include normalizing key feature variables and constructing an intelligent prediction model for influent water quality based on the XGBoost method.

[0011] The prediction process involves real-time prediction based on an intelligent influent water quality prediction model.

[0012] The principle and advantages of this solution are as follows: In practical applications, multi-dimensional data of industrial parks are collected, and the correlation between the correlation data and the influent water quality is determined by the Spearman rank correlation coefficient. Based on this, highly correlated data are selected as input features to construct a distributed gradient enhancement library influent water quality prediction model, thereby deeply mining multi-dimensional correlation data, improving prediction accuracy, and realizing accurate and forward-looking prediction of influent water quality of sewage treatment plants, providing a scientific basis for environmental management and sewage treatment optimization in smart industrial parks.

[0013] Beneficial effects:

[0014] Improved prediction accuracy: By analyzing multi-dimensional correlation data and comprehensively considering various influencing factors, the accuracy of influent water quality prediction has been significantly improved.

[0015] Enhance decision support: Provide forward-looking water quality forecasting information for smart industrial parks, supporting managers to make more scientific and reasonable wastewater treatment decisions.

[0016] Optimize resource allocation: Dynamically adjust wastewater treatment processes based on forecast results to effectively avoid resource waste and reduce operating costs.

[0017] Promoting green development: It helps to achieve precise management of wastewater treatment, reduce pollutant emissions, and promote the smart industrial park towards green, low-carbon, and sustainable development.

[0018] Preferably, as an improvement, the correlation analysis step includes:

[0019] The sorting sub-step aligns and expands the multi-dimensional data and influent water quality according to time order, and obtains the rank for each.

[0020] The rank difference acquisition sub-step obtains the rank difference between each pair of multidimensional data and influent water quality in the observations;

[0021] The rank correlation coefficient calculation sub-step calculates the rank correlation coefficient by running a preset rank correlation coefficient calculation model based on the rank difference.

[0022] Technical benefits: Facilitates accurate acquisition of the correlation between multi-dimensional data and influent water quality.

[0023] Preferably, as an improvement, the rank correlation coefficient calculation model is as follows:

[0024]

[0025] in, Let d be the rank correlation coefficient. ij denoted as rank difference, and m as the number of observations.

[0026] Technical effect: Facilitates the quantification of rank correlation.

[0027] Preferably, as an improvement, in the sorting sub-step, when there are equal observations, the rank is the average of the corresponding sorting positions.

[0028] Technical effect: Avoids the inability to obtain the rank when the observed values ​​are equal.

[0029] Preferably, as an improvement, in the sorting sub-step, when there are equal observations, the rank correlation coefficient calculation model is as follows:

[0030]

[0031] in, T is the rank correlation coefficient. x and T y These are the number of stalemate level observations in the variables, where m is the number of observations and d is the number of observations. ij It represents the difference in rank.

[0032] Technical effect: It enables the quantification of rank correlation even when the observed values ​​are equal.

[0033] Preferably, as an improvement, the correlation includes positive correlation, negative correlation and no correlation. When the rank correlation coefficient is close to 1, it indicates a strong positive correlation. When the rank correlation coefficient is close to -1, it indicates a strong negative correlation. When the rank correlation coefficient is close to 0, it indicates a weak correlation.

[0034] Technical benefits: Facilitates accurate acquisition of key feature variables.

[0035] Preferably, as an improvement, in the intelligent prediction model construction step, the objective function for constructing the intelligent prediction model for influent water quality based on the XGBoost method is:

[0036]

[0037] Among them, y i This represents the true value of the influent water quality. The model is based on the input (x1, x2, ... x MThe predicted value is given by (M≤m), where T represents the number of leaf nodes in a tree, λ is the adjustment parameter for this term, ω represents the vector composed of the output values ​​of the leaf nodes, and β is the regularization parameter that controls the complexity of the output values ​​of the leaf nodes.

[0038] Technical benefits: It facilitates improvements in the accuracy, computational efficiency, robustness, and interpretability of influent water quality prediction.

[0039] Preferably, as an improvement, it also includes a model verification and training step, which verifies the accuracy of the model prediction based on the real-time prediction results of the intelligent prediction model for influent water quality and adjusts the wastewater treatment process parameters.

[0040] Technical benefits: It facilitates ensuring stable and compliant effluent quality while reducing energy consumption and operating costs.

[0041] Preferably, as an improvement, it also includes an integration and optimization step, which integrates the intelligent prediction model of influent water quality into the management system of the industrial park wastewater treatment plant.

[0042] Technical benefits: Facilitates online optimization of model parameters to meet the needs of high-accuracy real-time monitoring. Attached Figure Description

[0043] Figure 1 This is a flowchart illustrating a smart water quality prediction method based on rank correlation and XGBoost models.

[0044] Figure 2 A schematic diagram illustrating multi-dimensional correlation data analysis and model building. Detailed Implementation

[0045] The following detailed description illustrates the specific implementation method:

[0046] The basic implementation examples are as follows: Figure 1 Appendix Figure 2 As shown:

[0047] Intelligent water quality prediction methods based on rank correlation and XGBoost models include:

[0048] The data acquisition process involves collecting multi-dimensional data from the industrial park and preprocessing the data. This multi-dimensional data includes, but is not limited to, historical influent water quality monitoring data from wastewater treatment plants (such as COD, BOD, ammonia nitrogen, total phosphorus, and total nitrogen), economic data (such as industrial production index and park GDP), enterprise production activity data (such as production volume and raw material consumption), park population data, and meteorological data (such as temperature, humidity, and rainfall). Preprocessing techniques such as data cleaning, missing value imputation, and outlier detection are employed to preprocess the multi-dimensional data from the industrial park to ensure data quality.

[0049] The correlation analysis step uses Spearman's rank correlation coefficient to analyze multidimensional data (x1, x2, ... x). m ) and influent water quality (y1, y2, ..., y n The correlation between the data and the influent water quality is determined. Specifically, the sorting sub-step aligns and expands the multi-dimensional data and influent water quality according to time order, and obtains the rank for each. and The ranking is based on the multi-dimensional data and the influent water quality, i.e., the order of importance. ij For a pair of observations, the rank difference is... and The difference, For the rank of multi-dimensional data, This refers to the rank of the incoming water quality. If... The order and the difference between orders are shown in Table 1.

[0050] Table 1

[0051]

[0052]

[0053] The rank correlation coefficient calculation model is as follows:

[0054]

[0055] in, Let d be the rank correlation coefficient. ij denoted as rank difference, and m as the number of observations.

[0056] When there are equal observations, as shown in Table 2, the observations... The rank of the two corresponding observations is the mean of their respective rankings.

[0057] Table 2

[0058]

[0059] When there are equal observations, the rank correlation coefficient calculation model is as follows:

[0060]

[0061] in, T is the rank correlation coefficient. x and T y These are the number of stalemate level observations in the variables, m and d. ij It represents the difference in rank.

[0062] The value range is between -1 and 1. When A value close to 1 indicates a strong positive correlation between the two variables. A value close to -1 indicates a strong negative correlation between the two variables. A value close to 0 indicates that there is almost no correlation between the two variables. Based on the strength of the correlation, [the following is a possible interpretation:] ... The values ​​are divided as shown in Table 3.

[0063] Table 3

[0064]

[0065] The key feature variable selection process involves using correlation analysis results to select multi-dimensional data that meet preset correlation criteria as key feature variables. Specifically, based on the multi-dimensional data x... i (i = 1, 2, ..., m) and the water quality of each influent y j The absolute value of the rank correlation coefficient among (j = 1, 2, ..., n) Selection and influent water quality j (j=1,2,…,n) has multi-dimensional data with moderate or higher correlation, that is, key characteristic variables that have a significant impact on changes in influent water quality.

[0066] The intelligent prediction model construction steps include normalizing key feature variables and constructing an intelligent prediction model for influent water quality based on the XGBoost method. The intelligent prediction model for influent water quality created using the XGBoost method features high accuracy, computational efficiency, robustness, and strong interpretability. Specifically,

[0067] Assume the multidimensional correlation data combination is (x1, x2, ... x M (M≤m), establish a relationship with a certain influent water quality y j The intelligent prediction model XGBoost for (j=1,2,…,n) is a multi-learner simultaneous learning model, which is as follows:

[0068]

[0069] Among them, y i This represents the true value of the influent water quality. The model is based on the input (x1, x2, ... x M The predicted value given by (M≤m) Let the loss function be the mean squared error for all learners, and for regression prediction problems, let mean squared error be used. Let λ represent the complexity of K learners, T represent the number of leaf nodes in a tree, λ be the adjustment parameter for this term, ω represent the vector composed of the output values ​​of the leaf nodes, and β be the regularization parameter that controls the complexity of the output values ​​of the leaf nodes.

[0070] The dataset was divided into training, validation, and test sets in a 7:2:1 ratio. The model was trained using the training set data, and the model parameters were adjusted using the validation set data to ensure the model's prediction accuracy.

[0071] It also includes model validation and training steps, applying the constructed intelligent influent water quality prediction model to actual data, solving for the optimal output value of each leaf node and the best tree structure by minimizing the objective function, verifying the prediction accuracy based on the prediction results, adjusting the sewage treatment process parameters in a timely manner, optimizing resource allocation, ensuring stable compliance of effluent water quality, and reducing energy consumption and operating costs.

[0072] It also includes integration and optimization steps, integrating the intelligent influent water quality prediction model into the management system of the industrial park's wastewater treatment plant. This enables automatic data collection, intelligent prediction, result display, and early warning functions, providing managers with an intuitive and convenient operating interface. Based on the discrepancy between the measured data at the wastewater treatment plant's influent and the predicted data from the intelligent prediction model, the model parameters are optimized online to meet the requirements of high-accuracy real-time monitoring.

[0073] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific technical solutions and / or characteristics are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the technical solutions of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A water quality intelligent prediction method based on rank correlation and XGBoost model, characterized in that, include: The data collection steps involve collecting multi-dimensional data from the industrial park, including historical influent water quality monitoring data from the industrial park's wastewater treatment plant, production activity data from enterprises in the park, population data, meteorological data, and economic data from the park, and then preprocessing the data. The correlation analysis step uses Spearman's rank correlation coefficient to analyze the correlation between multi-dimensional data and influent water quality; The correlation analysis steps include: The sorting sub-step aligns and expands the multi-dimensional data and influent water quality according to time order, and obtains the rank for each; when there are equal observations, the rank is the average of the corresponding sorting positions; The rank difference acquisition sub-step obtains the rank difference between each pair of multidimensional data and influent water quality in the observations; The rank correlation coefficient calculation sub-step calculates the rank correlation coefficient by running a preset rank correlation coefficient calculation model based on the rank difference. The key feature variable screening step involves selecting multi-dimensional data that meet preset criteria based on the correlation analysis results as key feature variables. The steps for building the intelligent prediction model include normalizing key feature variables and constructing an intelligent prediction model for influent water quality based on the XGBoost method. The objective function for constructing the intelligent prediction model for influent water quality based on the XGBoost method is: in, This represents the true value of the influent water quality. For the model based on input The given predicted value, This represents the number of leaf nodes in a tree. This is the adjustment parameter for this item. This represents a vector consisting of the output values ​​of the leaf nodes. Regularization parameters used to control the complexity of leaf node output values; The prediction process involves real-time prediction based on an intelligent influent water quality prediction model. The integration and optimization process involves integrating the intelligent influent water quality prediction model into the management system of the industrial park's wastewater treatment plant. This enables automatic data collection, intelligent prediction, result display, and early warning functions. Furthermore, the model parameters are optimized online based on the differences between the measured data at the wastewater treatment plant's influent and the model's predicted data.

2. The intelligent water quality prediction method based on rank correlation and XGBoost model according to claim 1, characterized in that: The rank correlation coefficient calculation model is as follows: in, The rank correlation coefficient, For the rank difference, This represents the number of observations.

3. The intelligent water quality prediction method based on rank correlation and XGBoost model according to claim 1, characterized in that: In the sorting sub-step, when there are equal observations, the rank correlation coefficient calculation model is as follows: in, The rank correlation coefficient, and These are the number of observations of the stalemate level in the variables. For the number of observations, It represents the difference in rank.

4. The intelligent water quality prediction method based on rank correlation and XGBoost model according to claim 1, characterized in that: The correlation includes positive correlation, negative correlation, and no correlation. When the rank correlation coefficient is close to 1, it indicates a strong positive correlation. When the rank correlation coefficient is close to -1, it indicates a strong negative correlation. When the rank correlation coefficient is close to 0, it indicates a weak correlation.

5. The intelligent water quality prediction method based on rank correlation and XGBoost model according to claim 1, characterized in that: It also includes model validation and training steps, which verify the accuracy of model prediction based on the real-time prediction results of the intelligent influent water quality prediction model and adjust the wastewater treatment process parameters.