Method for Evaluating Surface Water Pollution Level Based on Multi-Source Data
Through the surface water pollution level assessment method based on multi-source data and combined with the bat algorithm optimization evaluation model, the problems of insufficient pollution assessment accuracy and poor interpretability in the existing technology are solved, and more accurate and scientific pollution assessment and control are achieved.
Patent Information
- Application Number
- CN202411746018.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing surface water pollution assessment methods have problems of insufficient accuracy and poor interpretability in data fusion, pollution index calculation and evaluation model construction, which is difficult to provide a reliable basis for pollution control.
The surface water pollution level assessment method based on multi-source data is adopted. By obtaining groundwater data, standard data and multi-source monitoring data, comprehensive quality levels, dynamic environmental pollution limits and single-item pollution index are calculated, and the pollution level assessment model is optimized in combination with the bat algorithm to achieve more accurate pollution assessment.
It improves the accuracy and interpretability of surface water pollution assessment, enhances the timeliness and scientificity of the method, and effectively assists in the precise control and scientific control of surface water pollution.
Smart Images

Figure CN119669936B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental science, and particularly to a method for evaluating the surface water pollution level based on multi-source data. Background Art
[0002] With the rapid development of industrialization and urbanization, the problem of surface water pollution has become increasingly serious, and traditional evaluation methods have limitations. On the one hand, relying solely on a single type of monitoring data, it is difficult to comprehensively reflect the pollution status of water bodies, and it is easy to ignore the combined effects of multiple pollutants and the influence of complex environmental factors, resulting in one-sided evaluation. On the other hand, some methods have poor accuracy, cannot accurately define the pollution level, and are difficult to provide a reliable basis for pollution control. In the face of dynamic pollution situations, they have poor adaptability, cannot timely and accurately track the pollution trend, and affect the formulation of effective management and restoration decisions for the water environment. There is an urgent need for new methods to improve the scientificity and accuracy of evaluation.
[0003] In recent years, the development of multi-source data fusion technology has provided a new approach for surface water pollution assessment. However, existing methods still have problems such as insufficient accuracy and poor interpretability in data fusion, pollution index calculation, and evaluation model construction. Therefore, developing a method for evaluating the surface water pollution level based on multi-source data to improve the evaluation accuracy and interpretability has become an important research direction in the field of environmental science. Summary of the Invention
[0004] The object of the present invention is to provide a method for evaluating the surface water pollution level based on multi-source data.
[0005] To achieve the above object, the present invention is implemented according to the following technical solution:
[0006] The present invention includes the following steps:
[0007] Obtain groundwater data, standard data, and multi-source monitoring data;
[0008] Obtain the comprehensive quality level according to the standard data and the multi-source monitoring data, obtain the dynamic environmental pollution limit according to the standard data and the groundwater data, and obtain the single pollution index according to the multi-source monitoring data and the environmental pollution limit;
[0009] Obtain the comprehensive pollution index according to the single pollution index, perform clustering on the comprehensive pollution index and the comprehensive quality level to obtain the comprehensive pollution assessment level, and establish a pollution level evaluation model using the comprehensive pollution assessment level;
[0010] Optimize the pollution level evaluation model using the bat algorithm, input the data to be evaluated into the pollution level evaluation model, and output the evaluation result.
[0011] Further, the groundwater data, the standard data, and the multi-source monitoring data include:
[0012] The groundwater data are the toxicological indicators of groundwater. The standard data are the quality grading standards for multi-source monitoring data of different water areas. The multi-source monitoring data include the physical indicators, general chemical indicators, and toxicological indicators of surface water. The physical indicators include chromaticity, total dissolved solids, total hardness, and turbidity. The toxicological indicators are the concentrations of harmful elements, where the harmful elements are elements harmful to humans and elements harmful to the environment. The general chemical indicators are the concentrations of non-toxicological elements and compounds.
[0013] Further, the method for obtaining the comprehensive quality grade based on the standard data and the multi-source monitoring data includes:
[0014] Quality grading is performed on each index of the multi-source monitoring data according to the quality grading standard of the water area where the monitoring point is located to obtain the quality grade data of each index of the multi-source monitoring data. The quality grade data of the most severely polluted index is used as the comprehensive quality grade of the multi-source monitoring data.
[0015] Further, the method for obtaining the dynamic environmental pollution limit value based on the standard data and the groundwater data includes:
[0016] Perform stability analysis on the groundwater data:
[0017]
[0018] Where S t is the groundwater data at time point t, ΔS t is the first-order difference of S t S t-1 is the groundwater data at time point t - 1, S t-i is the groundwater data at time point t - i, α is the constant term, β is the coefficient of the time trend term, γ is the stationary coefficient of the autoregressive term, is the coefficient of the i-th order lag term, p is the order of the autoregressive term, ∈ t is the error term, δ K is the coefficient of the seasonal effect of the K-th season, J t,K is the seasonal effect of the K-th season where time point t is located,
[0019] If γ = 0, the groundwater data is non-stationary; if γ < 0, the groundwater data is stationary;
[0020] Calculate the average value of the toxicological indicators with stationary data in the groundwater to obtain the background value of the harmful elements. Use the background value as the environmental pollution limit value of the harmful elements. Use the third-level standard limit value of the quality grading standard of the water area where the monitoring point is located as the environmental pollution limit value of the harmful elements without a background value.
[0021] Furthermore, the method for obtaining the single pollution index according to the multi-source monitoring data and the environmental pollution limit values includes:
[0022] Single pollution index:
[0023]
[0024] In the formula, P i (j) is the single pollution index of pollutant j in sample i, C i (j) is the concentration of pollutant j in surface water sample i, C 0 (j) is the environmental pollution limit value of pollutant j in surface water sample i. The single pollution index > 1 indicates that the pollutant exceeds the environmental pollution limit value.
[0025] Furthermore, the method for obtaining the comprehensive pollution index according to the single pollution index includes:
[0026] Comprehensive pollution index:
[0027]
[0028] In the formula, P Z (i) is the comprehensive pollution index of sample i, maxP i is the maximum value of the single pollution index of sample i, is the average value of the single pollution index of sample i.
[0029] Furthermore, the method for obtaining the comprehensive pollution assessment level by clustering the comprehensive pollution index and the comprehensive quality level includes:
[0030] Data points are established using the comprehensive pollution index and the comprehensive quality level. The preset number of clusters is 5, and 5 objects are randomly selected from the data points as the initial cluster centers.
[0031] Formula for the importance degree of attributes:
[0032]
[0033] Among them, σ a (C, D) represents the importance degree of attribute a relative to the attribute set C and the decision attribute D, where represents the dependence between C and D, C - {a} represents removing attribute a from the attribute set C, X is a subset of U, U / D represents the partition of U with respect to D, U is the universe of discourse, representing a finite non-empty data set, |U| is the total number of objects in the universe of discourse, C (X) represents the lower approximation of X with respect to C, | C (X)| is CThe total number of objects in (X), σ′ a (C, D) represents the importance degree of the normalized attribute a with respect to the attribute set C and the decision attribute D.
[0034] Rough entropy formula:
[0035]
[0036] where E A (X) is the rough entropy of the attribute set A with respect to the set X. B nR (X) = R(X) - R (X), which respectively represent the upper approximation and the lower approximation of the set X with respect to the equivalence relation R, X i is the i-th equivalence class in U / IND(R), the i-th subset in the partition of U with respect to the indiscernibility relation IND(R), |·| represents the total number of objects in the set, and n is the number of equivalence classes in U / IND(R).
[0037] Attribute importance formula:
[0038] Sig(a, A) = E A-{a} (X) - E A (X),
[0039]
[0040] where Sig(a, A) is the importance of the attribute a in the attribute set A, and E A-{a} (X) represents the rough entropy of the set after removing the attribute a from the attribute set A. represents the rough entropy of the set after removing the j-th attribute a j from the attribute set A, W j is the weight vector of a j , Sig(a j , A) is the importance of the attribute a j in the attribute set A, and m is the total number of attributes in the attribute set A.
[0041] Similarity measure formula between data points:
[0042]
[0043] where dist(x i , x j ) is the distance between the data points x i and x j , is the similarity between data points, d i and d j are respectively the decision attribute values of x i and x j . The s-th attribute a after normalization s The importance degree w with respect to the attribute set C and the decision attribute D s For the attribute a s The weight vector, e is the number of conditional attributes, c is And c js Are respectively the s-th conditional attribute values of x i And x j And the s-th conditional attribute value of x
[0044] Calculate the distance between the data points and the cluster centers, assign each data point to the nearest cluster, and in each attribute of each cluster, select the mode of the attribute values as the new cluster center, and iterate and update until the cluster centers no longer change;
[0045] The five clusters respectively correspond to five comprehensive pollution assessment levels of safe, warning line, light pollution, medium pollution and heavy pollution, and the pollution levels of the five comprehensive pollution assessment levels are clean, fairly clean, slightly polluted, moderately polluted and severely polluted in turn.
[0046] Furthermore, the method for establishing a pollution level assessment model by using the comprehensive pollution assessment level includes:
[0047] Randomly sample the groundwater data, standard data, multi-source monitoring data and comprehensive pollution assessment level to obtain a training set and a validation set, use the kernel ridge regression algorithm to train the pollution level assessment model according to the training set, predict the comprehensive pollution assessment level, and use the intelligent optimization algorithm to optimize the pollution level assessment model based on the validation set.
[0048] Furthermore, the method for optimizing the pollution level assessment model by using the bat algorithm includes:
[0049] Position update formula:
[0050]
[0051] Wherein, And Are respectively the k-th dimensional velocity components of the i-th bat at the t-th and t + 1-th iterations, And Are respectively the k-th dimensional coordinates of the i-th bat at the t-th and t + 1-th iterations, rand is a random number uniformly distributed between 0 and 1, and Y(·) is the velocity threshold function,
[0052] Velocity update formula:
[0053]
[0054] Wherein, ω max Is the maximum weight, ω minis the minimum weight, is the weight of the i-th bat at the t-th iteration, is the k-th dimensional velocity component of the i-th bat at the (t - 1)-th iteration, T max is the maximum number of iterations, t is the current iteration number, f i is the frequency of the i-th bat, is the k-th dimensional coordinate of the current best position,
[0055] Local search formula:
[0056]
[0057] where, is the new value of the k-th dimensional coordinate of the i-th bat at the t-th iteration after local search, p m is the mutation probability, flip bitvalue is the value for mutating the current position, and the bit-flip mutation operator is adopted,
[0058] The optimization objective is to minimize the objective function value of the pollution level evaluation model, and the fitness is the reciprocal of the objective function value of the hyperparameter combination corresponding to the bat.
[0059] The beneficial effects of the present invention are:
[0060] The present invention is a method for evaluating the surface water pollution level based on multi-source data. Compared with the prior art, the present invention has the following technical effects:
[0061] Based on multi-source data, the present invention comprehensively considers various indicators to improve the evaluation accuracy. Dynamically calculates the environmental pollution limit value to enhance the timeliness and scientificity of the method. Through clustering and model optimization, it improves the evaluation accuracy and interpretability, effectively assisting in the precise treatment and scientific control of surface water pollution, and providing key support for water resource protection and ecological balance maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is the step flow chart of the method for evaluating the surface water pollution level based on multi-source data of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0063] The following further describes the present invention through specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but do not limit the present invention.
[0064] The method for evaluating the surface water pollution level based on multi-source data of the present invention includes the following steps:
[0065] As Figure 1 shown, in this embodiment, it includes the following steps:
[0066] Obtain groundwater data, standard data, and multi-source monitoring data;
[0067] Obtain the comprehensive quality grade based on the standard data and the multi-source monitoring data, obtain the dynamic environmental pollution limit based on the standard data and the groundwater data, and obtain the single pollution index based on the multi-source monitoring data and the environmental pollution limit;
[0068] Obtain the comprehensive pollution index based on the single pollution index, perform clustering on the comprehensive pollution index and the comprehensive quality grade to obtain the comprehensive pollution assessment grade, and establish a pollution grade assessment model using the comprehensive pollution assessment grade;
[0069] Optimize the pollution grade assessment model using the bat algorithm, input the data to be evaluated into the pollution grade assessment model, and output the evaluation result;
[0070] In actual evaluation, the data to be evaluated are 15 samples in the mining area, and the comprehensive quality grade and quality grade data of the samples are:
[0071]
[0072] The background value of Sb in the sample is 0.015 mg / L, the background value of As is 0.05 mg / L, and the environmental pollution limits of other harmful elements are the third-level standard limits of the surface water quality classification standard in the mining area,
[0073] The single pollution index and the comprehensive pollution index are:
[0074]
[0075]
[0076] In the table, P-As represents the single pollution index of As, and Pz represents the comprehensive pollution index of the sample;
[0077] The comprehensive pollution assessment grade is:
[0078]
[0079] The comprehensive pollution assessment grade of the sample is:
[0080]
[0081]
[0082] In this embodiment, the groundwater data, the standard data, and the multi-source monitoring data include:
[0083] The groundwater data are the toxicological indicators of groundwater, and the standard data are the quality grading standards for multi-source monitoring data of different water areas. The multi-source monitoring data include the physical indicators, general chemical indicators, and toxicological indicators of surface water. The physical indicators include chromaticity, total dissolved solids, total hardness, and turbidity. The toxicological indicators are the concentrations of harmful elements, and the harmful elements are elements harmful to humans and elements harmful to the environment. The general chemical indicators are the concentrations of non-toxicological elements and compounds.
[0084] In this embodiment, the method for obtaining the comprehensive quality grade according to the standard data and the multi-source monitoring data includes:
[0085] Perform quality grading on each index of the multi-source monitoring data according to the quality grading standard of the water area where the monitoring point is located to obtain the quality grade data of each index of the multi-source monitoring data, and use the quality grade data of the most severely polluted index as the comprehensive quality grade of the multi-source monitoring data.
[0086] In this embodiment, the method for obtaining the dynamic environmental pollution limit value according to the standard data and the groundwater data includes:
[0087] Conduct a stability analysis on the groundwater data:
[0088]
[0089] Where S t is the groundwater data at time point t, ΔS t is the first-order difference of S t S t-1 is the groundwater data at time point t-1, S t-i is the groundwater data at time point t-i, α is the constant term, β is the coefficient of the time trend term, γ is the stationary coefficient of the autoregressive term, is the coefficient of the i-th lag term, p is the order of the autoregressive term, ∈ t is the error term, δ K is the coefficient of the seasonal effect of the K-th season, J t,K is the seasonal effect of the K-th season where time point t is located,
[0090] If γ = 0, the groundwater data is non-stationary; if γ < 0, the groundwater data is stationary;
[0091] Calculate the average value of the toxicological indicators with stable data in the groundwater to obtain the background value of the harmful elements, use the background value as the environmental pollution limit value of the harmful elements, and use the third-level standard limit value of the quality grading standard of the water area where the monitoring point is located as the environmental pollution limit value of the harmful elements without a background value.
[0092] In this embodiment, the method for obtaining the single - pollution index based on the multi - source monitoring data and the environmental pollution limit values includes:
[0093] Single - pollution index:
[0094]
[0095] In the formula, P i (j) is the single - pollution index of pollutant j in sample i, C i (j) is the concentration of pollutant j in surface water sample i, C 0 (j) is the environmental pollution limit value of pollutant j in surface water sample i. The single - pollution index > 1 indicates that the pollutant exceeds the environmental pollution limit value.
[0096] In this embodiment, the method for obtaining the comprehensive pollution index based on the single - pollution index includes:
[0097] Comprehensive pollution index:
[0098]
[0099] In the formula, P Z (i) is the comprehensive pollution index of sample i, maxP i is the maximum value of the single - pollution index of sample i, is the average value of the single - pollution index of sample i.
[0100] In this embodiment, the method for obtaining the comprehensive pollution assessment level by clustering the comprehensive pollution index and the comprehensive quality level includes:
[0101] Use the comprehensive pollution index and the comprehensive quality level to establish data points. The preset number of clusters is 5. Randomly select 5 objects from the data points as the initial cluster centers.
[0102] Attribute importance degree formula:
[0103]
[0104]
[0105] Among them, σ a (C,D) represents the importance degree of attribute a relative to the attribute set C and the decision attribute D, where represents the dependence between C and D, C - {a} represents removing attribute a from the attribute set C, X is a subset of U, U / D represents the partition of U with respect to D, U is the universe of discourse, representing a finite non - empty data set, |U| is the total number of objects in the universe of discourse, C (X) represents the lower approximation of X with respect to C, | C (X)| isC The total number of objects in (X), σ′ a (C, D) is the importance degree of the normalized attribute a with respect to the attribute set C and the decision attribute D,
[0106] Rough entropy formula:
[0107]
[0108] where E A (X) is the rough entropy of the attribute set A with respect to the set X, B nR (X) = R(X) - R (X), which respectively represent the upper approximation and the lower approximation of the set X with respect to the equivalence relation R, X i is the i-th equivalence class in U / IND(R), the i-th subset in the partition of U with respect to the indiscernibility relation IND(R), |·| represents the total number of objects in the set, and n is the number of equivalence classes in U / IND(R),
[0109] Attribute importance formula:
[0110] Sig(a, A) = E A-{a} (X) - E A (X),
[0111]
[0112] where Sig(a, A) is the importance of the attribute a in the attribute set A, E A-{a} (X) represents the rough entropy of the set after removing the attribute a from the attribute set A, represents the rough entropy of the set after removing the j-th attribute a j from the attribute set A, W j is the weight vector of a j , Sig(a j , A) is the importance of the attribute a j in the attribute set A, m is the total number of attributes in the attribute set A,
[0113] Similarity measurement formula between data points:
[0114]
[0115] where, dist(x i , x j ) is the distance between the data points x i and x j , is the similarity between data points, d i and d j are respectively the x i and x jThe decision attribute value, is the s-th normalized attribute a s The importance degree relative to the attribute set C and the decision attribute D, w s is the attribute a s The weight vector of, e is the number of conditional attributes, c is and c js are respectively the s-th conditional attribute values of x i and x j respectively.
[0116] Calculate the distance between the data points and the cluster centers, assign each data point to the nearest cluster, and in each attribute of each cluster, select the mode of the attribute values as the new cluster center, and iterate and update until the cluster centers no longer change;
[0117] The five clusters respectively correspond to five comprehensive pollution assessment levels of safe, warning line, light pollution, medium pollution and heavy pollution, and the pollution levels of the five comprehensive pollution assessment levels are clean, fairly clean, slightly polluted, moderately polluted and severely polluted in sequence.
[0118] In this embodiment, the method for establishing a pollution level assessment model using the comprehensive pollution assessment level includes:
[0119] Randomly sample the groundwater data, standard data, multi-source monitoring data and comprehensive pollution assessment level to obtain a training set and a validation set, use the kernel ridge regression algorithm to train the pollution level assessment model according to the training set, predict the comprehensive pollution assessment level, and use the intelligent optimization algorithm to optimize the pollution level assessment model based on the validation set.
[0120] In this embodiment, the method for optimizing the pollution level assessment model using the bat algorithm includes:
[0121] Position update formula:
[0122]
[0123] Where and are respectively the k-th dimensional velocity components of the i-th bat at the t-th and t + 1-th iterations, and are respectively the k-th dimensional coordinates of the i-th bat at the t-th and t + 1-th iterations, rand is a random number uniformly distributed between 0 and 1, and Y(·) is the velocity threshold function,
[0124] Velocity update formula:
[0125]
[0126] Where ω maxis the maximum weight, ω min is the minimum weight, is the weight of the i-th bat at the t-th iteration, is the k-th dimensional velocity component of the i-th bat at the (t-1)-th iteration, T max is the maximum number of iterations, t is the current iteration number, f i is the frequency of the i-th bat, is the k-th dimensional coordinate of the current best position,
[0127] Local search formula:
[0128]
[0129] where, is the new value of the k-th dimensional coordinate of the i-th bat after local search at the t-th iteration, p m is the mutation probability, flip bitvalue is the value for mutating the current position, and the bit-flip mutation operator is adopted,
[0130] The optimization objective is to minimize the objective function value of the pollution level evaluation model, and the fitness is the reciprocal of the objective function value of the hyperparameter combination corresponding to the bat.
[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A surface water pollution level assessment method based on multi-source data, characterized in that: The following steps are involved: Obtain groundwater data, standard data and multi-source monitoring data; Obtain a comprehensive quality grade based on the standard data and the multi-source monitoring data, obtain a dynamic environmental pollution limit based on the standard data and the groundwater data, and obtain a single pollution index based on the multi-source monitoring data and the environmental pollution limit; Obtaining a comprehensive pollution index according to the single pollution index, clustering the comprehensive pollution index and the comprehensive quality grade to obtain a comprehensive pollution assessment grade, and establishing a pollution grade assessment model using the comprehensive pollution assessment grade; The pollution level assessment model is optimized by using a bat algorithm, the data to be assessed is input into the pollution level assessment model, and an assessment result is output; The groundwater data, the standard data and the multi-source monitoring data include: Groundwater data are toxicological indicators of groundwater, standard data are quality classification standards of multi-source monitoring data of different water bodies, multi-source monitoring data include physical indicators, general chemical indicators and toxicological indicators of surface water, physical indicators include chromaticity, total dissolved solids, total hardness and turbidity, toxicological indicators are the concentration of harmful elements, harmful elements are elements harmful to human body and elements harmful to environment, general chemical indicators are the concentration of non-toxic elements and compounds; The method for obtaining a comprehensive quality level according to the standard data and the multi-source monitoring data comprises: According to the quality grading standard of the water area where the monitoring point is located, the quality grade data of each indicator of the multi-source monitoring data is obtained, and the quality grade data of the indicator with the most serious pollution is used as the comprehensive quality grade of the multi-source monitoring data; The method for obtaining a dynamic environmental pollution limit value according to the standard data and the groundwater data comprises: Perform stability analysis on groundwater data: Where S t is the groundwater data at time point t, ΔS t YesS t The first-order difference of t-1 is the groundwater data at time point t-1, S t-i is the groundwater data at time point ti, α is the constant term, β is the coefficient of the time trend term, γ is the stationary coefficient of the autoregressive term, is the coefficient of the i-th order lag term, p is the order of the autoregressive term, ∈ t is the error term, δ K is the coefficient of the seasonal effect of the Kth season, J t,K is the seasonal effect of the Kth season at time point t, If γ=0, the groundwater data is non-stationary; if γ<0, the groundwater data is stationary; The average value of toxicological indicators with stable data in groundwater is calculated to obtain the background value of harmful elements, and the background value is used as the environmental pollution limit of harmful elements. The third-level standard limit of the quality grading standard of the water area where the monitoring point is located is used as the environmental pollution limit of harmful elements without background value.
2. The surface water pollution level assessment method based on multi-source data according to claim 1 is characterized in that: The method for obtaining a single pollution index according to the multi-source monitoring data and the environmental pollution limit value includes: Single pollution index: Where P i (j) is the single pollution index of pollutant j in sample i, C i (j) is the concentration of pollutant j in surface water sample i, C0(j) is the environmental pollution limit of pollutant j in surface water sample i, and a single pollution index greater than 1 indicates that the pollutant exceeds the environmental pollution limit.
3. The surface water pollution level assessment method based on multi-source data according to claim 1 is characterized in that: The method for obtaining the comprehensive pollution index according to the single pollution index includes: Comprehensive pollution index: Where P Z (i) is the comprehensive pollution index of sample i, maxP i is the maximum value of the single pollution index of sample i, is the average value of the single pollution index of sample i.
4. The surface water pollution level assessment method based on multi-source data according to claim 1 is characterized in that: The method of clustering the comprehensive pollution index and the comprehensive quality grade to obtain a comprehensive pollution assessment grade comprises: The comprehensive pollution index and comprehensive quality grade are used to establish data points. The number of clusters is preset to 5, and 5 objects are randomly selected from the data points as the initial cluster centers. Attribute importance formula: Among them, σ a (C, D) represents the importance of attribute a relative to attribute set C and decision attribute D, where represents the dependency between C and D, C-{a} represents the removal of attribute a from the attribute set C, X is a subset of U, U / D represents the partition of U with respect to D, U is the domain, representing a finite non-empty data set, |U| is the total number of objects in the domain, C (X) represents the lower approximation of X with respect to C, | C (X)| C The total number of objects in (X), σ′ a (C, D) is the normalized importance of attribute a relative to attribute set C and decision attribute D. Rough entropy formula: Where E A (X) is the rough entropy of attribute set A with respect to set X, B nR (X) = R (X) - R (X), respectively represent the upper approximation and lower approximation of set X with respect to the equivalence relation R, X i is the i-th equivalence class in U / IND(R), the i-th subset in the partition of U with respect to the indiscernibility relation IND(R), |·| represents the total number of objects in the set, n is the number of equivalence classes in U / IND(R), Attribute importance formula: Mr(a,A)=E A-{a} (X)-E A (X), Where Sig(a,A) is the importance of attribute a in attribute set A, E A-{a} (X) represents the rough entropy of the set after removing attribute a from attribute set A. Indicates removing the jth attribute a from the attribute set A j Then about the rough entropy of the set, W j for a j The weight vector, Sig(a j ,A) is attribute a j The importance of attribute set A, m is the total number of attributes in attribute set A, The similarity measurement formula between data points is: Among them, dist(x i ,x j ) is the data point x i and x j The distance between them is the similarity between data points, d i and d j x i and x j The decision attribute value of is the normalized sth attribute a s Relative to the importance of attribute set C and decision attribute D, w s For attribute a s The weight vector, e is the number of conditional attributes, c is and c js They are x i and x j The sth conditional attribute value of Calculate the distance between the data point and the cluster center, assign each data point to the nearest cluster, and in each attribute of each cluster, select the mode of the attribute value as the new cluster center, and iterate until the cluster center no longer changes; The five clusters correspond to the five comprehensive pollution assessment levels of safety, warning line, light pollution, moderate pollution and heavy pollution. The pollution levels of the five comprehensive pollution assessment levels are clean, fairly clean, light pollution, moderate pollution and severe pollution.
5. The surface water pollution level assessment method based on multi-source data according to claim 1 is characterized in that: The method of establishing a pollution level assessment model using a comprehensive pollution assessment level includes: Random sampling was performed on groundwater data, standard data, multi-source monitoring data and comprehensive pollution assessment levels to obtain training sets and validation sets. The kernel ridge regression algorithm was used to train the pollution level assessment model based on the training set to predict the comprehensive pollution assessment level. The intelligent optimization algorithm was used to optimize the pollution level assessment model based on the validation set.
6. The surface water pollution level assessment method based on multi-source data according to claim 1 is characterized in that: The method of optimizing the pollution level assessment model using a bat algorithm includes: Position update formula: in, and are the k-th velocity components of the ith bat at the t-th and t+1-th iterations, respectively. and are the k-th coordinates of the ith bat at the t-th and t+1-th iterations, rand is a random number uniformly distributed between 0 and 1, Y(·) is the speed threshold function, Speed update formula: Among them, ω max is the maximum weight, ω min is the minimum weight, is the weight of the i-th bat at the t-th iteration, is the k-th velocity component of the i-th bat at the t-1th iteration, T max is the maximum number of iterations, t is the current number of iterations, f i is the frequency of the i-th bat, is the k-th dimension coordinate of the current best position, Local search formula: in, is the new value of the k-th dimension coordinate of the i-th bat at the t-th iteration after local search, p m is the mutation probability, flip bitvalue is the value of the mutation operation on the current position, and the bit flip mutation operator is used. The optimization goal is to minimize the objective function value of the pollution level assessment model, and the fitness is the inverse of the objective function value of the hyperparameter combination corresponding to the bat.