A water pollution source tracing and detection method and system based on data mining

Through data mining and honey badger optimization algorithm combined with CNN neural network model, the traceability problem before and after the conversion of water pollutants is solved, the comprehensive and accurate traceability of water pollution is achieved, and the accuracy and efficiency of pollution source identification are improved.

CN119380850BActive Publication Date: 2025-07-25BEIJING BEIKONG YUEHUI ENVIRONMENTAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411955406.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-28
Publication Date
2025-07-25
Estimated Expiration
2044-12-28

AI Technical Summary

Technical Problem

The prior art is difficult to deal with the traceability difficulty caused by changes in the properties or morphology of pollutants in water pollution traceability, especially the differences before and after the conversion of pollutants in water bodies increase the difficulty of traceability.

Method used

Through a data mining method, multiple groups of existing polluted water data are collected, the data before conversion is merged and clustered, the classification center data is calculated, and the honey badger optimization algorithm is used for iterative optimization, and the garbage type is identified in combination with the CNN neural network model to achieve accurate matching of the pollution data before and after conversion.

Benefits of technology

It has achieved comprehensive and accurate traceability of water pollution from the two perspectives before and after the conversion, which has improved the accuracy and efficiency of pollution source identification and reduced data complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380850B_ABST
    Figure CN119380850B_ABST
Patent Text Reader

Abstract

The present invention discloses a water pollution source tracing and detection method and system based on data mining, which relates to the field of water quality detection. The present invention traces the water pollution from two perspectives, namely before and after the conversion of water pollutants, making the tracing more comprehensive and accurate; classifies the data after the conversion of pollution in the existing polluted water body data and calculates the classification center data of each category, providing a classification basis for subsequent classification of the polluted data after the conversion in the water area to be traced and detected; for the polluted data before the conversion in the polluted data in the water area to be traced and detected, in this solution, it is directly matched with the types of sewage chemical components and product types corresponding to the surrounding enterprises or factories, and for the polluted data after the conversion, it is first classified and mapped to be matched with the types of sewage chemical components and product types corresponding to the surrounding enterprises or factories for the polluted data before the conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of water quality detection. Specifically, it particularly relates to a water pollution source tracing and detection method and system based on data mining. Background Art

[0002] Chinese Patent with the publication number CN102661939B discloses a method for quickly realizing water pollution source tracing. By constructing a chemical fingerprint information database of pollution sources in the upper reaches and surrounding areas of a water area, it helps to quickly realize the pollution source tracing of this water area. By analyzing the pollution discharge information of upstream polluting enterprises and constructing the sewage chemical fingerprint databases of each enterprise in advance, it helps to quickly realize water pollution source tracing.

[0003] When conducting water body pollution source tracing, as time goes by, the properties or forms of pollutants in the water body are constantly changing, such as degradation or oxidation-reduction reactions, etc. Therefore, there are pollutant forms before and after the conversion in the water body. After the properties or forms of pollutants are converted, it is also necessary to analyze their original forms, which increases the difficulty of tracing the sources of pollutants. Summary of the Invention

[0004] Aiming at the problems in the related technologies, the present invention proposes a water pollution source tracing and detection method and system based on data mining to overcome the above-mentioned technical problems existing in the existing related technologies.

[0005] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0006] The present invention is a water pollution source tracing and detection method based on data mining, including the following steps:

[0007] S1. Collect multiple groups of existing polluted water body data;

[0008] S2. Merge and then cluster the data before the pollution conversion in the existing polluted water body data to obtain a classification matrix set of comprehensive water pollution data before conversion and a data matrix of comprehensive classification centers of water pollution before conversion;

[0009] S3. Classify the data after the pollution conversion in the existing polluted water body data according to the classification matrix set of comprehensive water pollution data before conversion and calculate the classification center data of each category to obtain a data matrix of comprehensive classification centers of water pollution after conversion;

[0010] S4. Collect pollution data in the water area to be traced and detected in cooperation with the existing polluted water body data;

[0011] S5. Match the pollution data before the conversion in the pollution data of the water area to be traced and detected with the types of chemical components and product types discharged by the surrounding enterprises or factories to obtain the first pollution source enterprise and the second pollution source enterprise;

[0012] S6. Classify the pollution data after the conversion in the pollution data of the water area to be traced and detected according to the converted water pollution comprehensive classification center data matrix to obtain the water pollution comprehensive data set before the conversion to be traced; Match the water pollution comprehensive data set before the conversion to be traced with the types of chemical components and product types discharged by the surrounding enterprises or factories to obtain the third pollution source enterprise and the fourth pollution source enterprise;

[0013] By collecting multiple groups of existing polluted water body data, it provides data support for subsequent clustering of the chemical component content data and garbage quality data in the water pollution before the conversion; By merging and then clustering the data before the pollution conversion in the existing polluted water body data, the central data of the data before the pollution conversion in the existing polluted water body data of each classification can be obtained, providing a classification basis for subsequent classification of the data after the pollution conversion in the corresponding existing polluted water body data; By classifying the data after the pollution conversion in the existing polluted water body data and calculating the classification center data of each category, it provides a classification basis for subsequent classification of the pollution data after the conversion in the water area to be traced and detected, and then the corresponding pollution data before the conversion can be obtained for matching with the types of chemical components and product types discharged by the surrounding enterprises or factories; For the pollution data before the conversion in the pollution data of the water area to be traced and detected, in this solution, it is directly matched with the types of chemical components and product types discharged by the surrounding enterprises or factories, while for the pollution data after the conversion, it is first classified and mapped to the pollution data before the conversion for matching with the types of chemical components and product types discharged by the surrounding enterprises or factories; Tracing the pollution data in the water area to be traced and detected from two perspectives before and after the conversion makes the tracing process more comprehensive and accurate.

[0014] Preferably, S1 includes the following steps:

[0015] S11. Set the water pollution type set; According to the water pollution type set, set multiple chemical component types and garbage types of the polluted water body before the water pollution conversion to obtain the chemical component type set before the conversion and the garbage type set before the conversion;

[0016] Then, according to the water pollution type set, set multiple chemical component types and garbage types of the polluted water body after the water pollution conversion to obtain the chemical component type set after the conversion and the garbage type set after the conversion;

[0017] S12. Set various types of water pollution conversion influencing factors and various types of pollutant distribution parameters to obtain a set of conversion influencing factor types and a set of pollutant distribution parameter types;

[0018] S13. According to the pre-conversion chemical composition type set, pre-conversion waste type set, post-conversion chemical composition type set, post-conversion waste type set, conversion influencing factor type set, and post-conversion distribution parameter type set, collect multiple groups of data on the content of various pre-conversion chemical components and corresponding distribution parameters, data on the mass of various pre-conversion wastes and corresponding distribution parameters, data on the content of various post-conversion chemical components and corresponding distribution parameters, data on the mass of various post-conversion wastes and corresponding distribution parameters, and conversion influencing factor data in existing polluted water bodies, to obtain a historical pre-conversion chemical component content data matrix, a historical pre-conversion chemical distribution parameter data matrix, a historical pre-conversion waste mass data matrix, a historical pre-conversion waste distribution parameter data matrix, a historical post-conversion chemical component content data matrix, a historical post-conversion chemical distribution parameter data matrix, a historical post-conversion waste mass data matrix, a historical post-conversion waste distribution parameter data matrix, and a historical conversion influencing factor data matrix;

[0019] The pre-conversion chemical composition type set includes heavy metals, organic substances, nutrients, radioactive substances, dissolved and particulate substances, endocrine disruptors, acid-base substances, and salts, etc.; the pre-conversion waste type set includes plastic waste, industrial waste, agricultural waste, domestic waste, and medical waste, etc.; the post-conversion chemical composition type set includes the redox reaction products of heavy metals, the reaction products of biodegradation, photolysis, hydrolysis, etc. of organic pollutants, the metabolic products of microorganisms, the precipitation, dissolution, adsorption, etc. reaction products of inorganic compounds, and the decay products of radioactive substances, etc.; the post-conversion waste type set includes that the organic substances in the waste may decompose in water to produce organic compounds such as organic acids, alcohols, aldehydes, etc.; heavy metals in the waste such as lead, mercury, cadmium, etc. may dissolve or form precipitates in water; and plastic particles in the waste, etc.;

[0020] The conversion influencing factor type set includes the influence of the environmental conditions of the water body; the physical migration of pollutants, such as the movement and diffusion with water flow and air flow, and the sedimentation effect of gravity; the dissolution, dissociation, redox reaction, hydrolysis, complexation, chelation, chemical precipitation, and biodegradation of pollutants, etc.; the migration of pollutants through the physiological processes of organisms such as absorption, metabolism, reproduction, and death; phase change, osmosis, adsorption, and radioactive decay; the different physical, chemical, and biological properties of media such as surface water, soil, and atmosphere; pollutant characteristics;

[0021] The types of pollutant distribution parameters mainly include pH value, conductivity, turbidity, dissolved oxygen, chemical oxygen demand, biochemical oxygen demand, total nitrogen and total phosphorus, total number of bacteria and number of coliforms, types and quantities of algae and plankton, spatial distribution, temporal distribution, types and concentrations of pollutants, and physicochemical properties of pollutants such as solubility, volatility, stability, etc.

[0022] Preferably, the S2 includes the following steps:

[0023] S21. Horizontally merge the historical chemical composition content data matrix before conversion, the historical chemical distribution parameter data matrix before conversion, the historical waste mass data matrix before conversion, the historical waste distribution parameter data matrix before conversion, and the historical conversion influencing factor data matrix to obtain a comprehensive water pollution data matrix before conversion; then horizontally merge the historical chemical composition content data matrix after conversion, the historical chemical distribution parameter data matrix after conversion, the historical waste mass data matrix after conversion, and the historical waste distribution parameter data matrix after conversion to obtain a comprehensive water pollution data matrix after conversion;

[0024] S22. Perform a clustering operation on the comprehensive water pollution data matrix before conversion to obtain a set of comprehensive water pollution data classification matrices before conversion and a comprehensive water pollution classification center data matrix before conversion;

[0025] By horizontally merging the historical chemical composition content data matrix before conversion, the historical chemical distribution parameter data matrix before conversion, the historical waste mass data matrix before conversion, the historical waste distribution parameter data matrix before conversion, and the historical conversion influencing factor data matrix, it is convenient to perform clustering on the whole of them subsequently; by performing a clustering operation on the comprehensive water pollution data matrix before conversion, the complexity of the data is reduced.

[0026] Preferably, the S3 includes the following steps:

[0027] S31. Classify the comprehensive water pollution data matrix after conversion according to the set of comprehensive water pollution data classification matrices before conversion and the comprehensive water pollution data matrix before conversion to obtain a set of comprehensive water pollution data classification matrices after conversion;

[0028] S32. Calculate the central data of each classification matrix in the set of comprehensive water pollution data classification matrices after conversion to obtain a comprehensive water pollution classification center data matrix after conversion;

[0029] By calculating the central data of each classification matrix in the set of comprehensive water pollution data classification matrices after conversion, a classification basis is provided for subsequently classifying the converted pollutant data in the water area to be traced and detected, so as to facilitate corresponding the converted pollutant data in the water area to be traced and detected to its pollutant data before conversion, and then conduct water pollution tracing.

[0030] Preferably, S32 includes the following steps:

[0031] S321. Construct a honey badger population for calculating the classification center of the converted water pollution data; set the maximum number of iterations for calculating the honey badger population of the classification center of the converted water pollution data as and the current number of iterations as , which are respectively denoted as the maximum number of iterations for calculating the classification center of water pollution data and the current number of iterations for calculating the classification center of water pollution data;

[0032] S322. Randomly select one row of data from each classification matrix in the set of classification matrices of the converted comprehensive water pollution data multiple times as the initial position matrix of each honey badger in the honey badger population for calculating the classification center of the converted water pollution data, and obtain a second set of initial position matrices;

[0033] S323. Calculate the Euclidean distance between each row of data in each classification matrix of the set of classification matrices of the converted comprehensive water pollution data and the corresponding classification center data in the initial position matrix of the k th honey badger in the honey badger population for calculating the classification center of the converted water pollution data, and obtain a second Euclidean distance matrix;

[0034] Construct a fitness function for the k th honey badger in the honey badger population for calculating the classification center of the converted water pollution data according to the second Euclidean distance matrix;

[0035] S324. Start iteration. Before iteration, set the current number of iterations for calculating the classification center of water pollution data to 1; in the first round of iteration, use the fitness function of the k th honey badger in the honey badger population for calculating the classification center of the converted water pollution data to calculate the fitness values of the initial position matrices of each honey badger in the second set of initial position matrices, and obtain a third set of fitness values; take the maximum fitness value in the third set of fitness values and the corresponding initial position matrix of the honey badger as the third global best fitness and the third global best position respectively; update the initial position matrices of each honey badger in the second set of initial position matrices according to the third global best fitness and the third global best position; after the update is completed, add 1 to the current number of iterations for calculating the classification center of water pollution data and enter the next round of iteration;

[0036] In each subsequent round of iteration, use the fitness function of the kThe fitness function of each honey badger calculates the classification center of the transformed water pollution data updated in the previous iteration process, calculates the fitness value of the position matrix of each honey badger in the honey badger population, and obtains the fourth fitness value set; the maximum fitness in the fourth fitness value set and the corresponding position matrix of the honey badger are respectively used as the fourth global best fitness and the fourth global best position; according to the fourth global best fitness and the fourth global best position, update the position matrix of each honey badger in the honey badger population calculated from the transformed water pollution data classification center updated in the previous iteration process; after the update is completed, add 1 to the current iteration number of the water pollution data classification center and enter the next iteration;

[0037] S325. When it is satisfied, stop the iteration to obtain the second final global best position; otherwise, continue the iteration until it is satisfied; use the second final global best position as the transformed water pollution comprehensive classification center data matrix;

[0038] In this solution, the honey badger optimization algorithm is used to perform multiple iterative optimizations on the classification center data of each transformed water pollution comprehensive data classification matrix in the transformed water pollution comprehensive data classification matrix set. In each iteration process, calculate the sum of the Euclidean distances between the iteratively obtained classification center data and each data in the classification, and use the sum of the Euclidean distances as the fitness function; therefore, as the iteration progresses, the sum of the Euclidean distances between each data in each transformed water pollution comprehensive data classification matrix in the transformed water pollution comprehensive data classification matrix set and the iteratively obtained classification center data becomes smaller and smaller, indicating that the iteratively obtained classification center data is more accurate.

[0039] Preferably, S4 includes the following steps:

[0040] S41. Set the water area to be traced and detected; set several water quality detection points in the water area to be traced and detected and collect image data of the water area to be traced and detected above to obtain a water quality detection point set and water area image data to be traced;

[0041] According to the set of pre-conversion chemical component types, the set of post-conversion chemical component types, and the set of pollutant distribution parameter types, detect and collect the corresponding content data of each pre-conversion chemical component, the content data of each post-conversion chemical component, and the corresponding distribution parameter data at each water quality detection point in the water quality detection point set to obtain a matrix of content data of pre-conversion chemical components to be traced, a matrix of content data of post-conversion chemical components to be traced, and a data set of post-conversion chemical distribution parameters to be traced;

[0042] Calculate the average of the chemical component content before conversion and the chemical component content after conversion for each water quality detection point based on the data matrix of the chemical component content before conversion to be traced and the data matrix of the chemical component content after conversion to be traced, and obtain the average data set of the chemical component content before conversion to be traced and the average data set of the chemical component content after conversion to be traced;

[0043] S42. Use the pre-trained CNN neural network model to identify the garbage types before conversion and the garbage types after conversion for the image data of the water area to be traced, and calculate the corresponding garbage quality data and distribution parameter data according to the recognition results, to obtain the data set of the garbage quality before conversion to be traced, the data set of the garbage quality after conversion to be traced, and the data set of the garbage distribution parameters after conversion to be traced;

[0044] S43. Collect the conversion influence factor data in the water area to be traced according to the conversion influence factor type set, and obtain the data set of the conversion influence factors to be traced;

[0045] By collecting the content data of various pollutants in the water area to be traced, it provides a data basis for subsequently determining whether the pollution data of the water area to be traced exceeds the corresponding threshold, and then determines whether it is necessary to trace the water pollution.

[0046] Preferably, the S5 includes the following steps:

[0047] S51. Collect the corresponding sewage chemical component types and product types of the enterprises or factories around the water area to be traced, and obtain the sewage chemical component type matrix and the product type matrix;

[0048] S52. Set the corresponding chemical component content exceeding threshold and garbage quality exceeding threshold according to the chemical component type set before conversion, the garbage type set before conversion, the chemical component type set after conversion, and the garbage type set after conversion, and obtain the chemical component content threshold set before conversion, the garbage quality threshold set before conversion, the chemical component content threshold set after conversion, and the garbage quality threshold set after conversion;

[0049] When there is an average chemical component content data in the average data set of the chemical component content before conversion to be traced that is greater than or equal to the corresponding chemical component content threshold in the chemical component content threshold set before conversion, match the chemical component type corresponding to the average data set of the chemical component content before conversion to be traced with the sewage chemical component type matrix to obtain the first pollution source enterprise; otherwise, there is no need to match with the sewage chemical component type matrix;

[0050] When there is waste quality data in the waste quality data set before traceability conversion that is greater than or equal to the corresponding waste quality threshold in the waste quality threshold set before conversion, match the waste type before conversion corresponding to the waste quality data set before traceability conversion with the product type matrix to obtain the second pollution source enterprise; otherwise, there is no need to match with the product type matrix;

[0051] By setting the chemical composition content threshold set before conversion, the waste quality threshold set before conversion, the chemical composition content threshold set after conversion, and the waste quality threshold set after conversion, it provides a quantitative determination basis for determining whether the content of various pollutants in the water area to be traced and detected is within the normal range; since the pollutants can be directly corresponded to the substances discharged from the enterprise before conversion, in this solution, the data before conversion in the water area to be traced and detected collected is directly matched with the substance category content and quality discharged from the enterprise.

[0052] Preferably, the step S6 includes the following steps:

[0053] S61. When there is average chemical composition content data after conversion in the average chemical composition content data set after traceability conversion that is greater than or equal to the corresponding chemical composition content threshold in the chemical composition content threshold set after conversion, or there is waste quality data after conversion in the waste quality data set after traceability conversion that is greater than or equal to the corresponding waste quality threshold in the waste quality threshold set after conversion, merge the average chemical composition content data set after traceability conversion, the waste quality data set after traceability conversion, the distribution parameter data set after chemical conversion for traceability, the waste distribution parameter data set after traceability conversion, and the influencing factor data set after traceability conversion to obtain a comprehensive water pollution data set for traceability; otherwise, there is no need to merge;

[0054] S62. Calculate the Euclidean distance between each row of data in the comprehensive water pollution data set for traceability and the data matrix of the comprehensive water pollution classification center after conversion to obtain a set of Euclidean distances; record the data of the comprehensive water pollution classification center before conversion in the data matrix of the comprehensive water pollution classification center before conversion corresponding to the smallest Euclidean distance in the set of Euclidean distances as the comprehensive water pollution data set before traceability conversion;

[0055] When there is chemical composition content data in the comprehensive water pollution data set before traceability conversion that is greater than or equal to the corresponding chemical composition content threshold set in the chemical composition content threshold set before conversion, match the chemical composition type corresponding to the chemical composition content before conversion in the comprehensive water pollution data set before traceability conversion with the sewage chemical composition type matrix to obtain the third pollution source enterprise; otherwise, there is no need to match;

[0056] When the comprehensive water pollution dataset before traceability conversion has waste quality data greater than or equal to the corresponding pre-conversion waste quality threshold in the pre-conversion waste quality threshold set, match the waste type corresponding to the pre-conversion waste quality data in the comprehensive water pollution dataset before traceability conversion with the product type matrix to obtain the fourth pollution source enterprise; otherwise, there is no need to match with the product type matrix;

[0057] S63. Conduct water pollution traceability based on the first pollution source enterprise, the second pollution source enterprise, the third pollution source enterprise, and the fourth pollution source enterprise;

[0058] Since after the conversion occurs, the categories and contents or qualities of pollutants in the water body have changed, it is impossible to directly match with the data of pollutants discharged by surrounding enterprises. In this solution, the polluted data after conversion is first classified and then mapped to the polluted data before conversion, and then matched with the data of pollutants discharged by surrounding enterprises.

[0059] A water pollution traceability detection system based on data mining includes a water pollution characteristic type setting module, an existing water pollution data collection module, a pre-conversion data merging and clustering module, a post-conversion data classification center calculation module, a polluted data collection module for the water area to be traced and detected, a surrounding enterprise data collection module, a primary traceability matching module, a polluted data classification module for the water area to be traced and detected after conversion, and a secondary traceability matching module.

[0060] The present invention has the following beneficial effects:

[0061] 1. In the present invention, water pollution traceability is carried out from two perspectives before and after the conversion of water body pollutants, making the traceability more comprehensive and accurate; among them, by classifying the data after the pollution conversion in the existing polluted water body data and calculating the classification center data of each category, it provides a classification basis for subsequent classification of the polluted data after conversion in the water area to be traced and detected; for the polluted data before conversion in the polluted data of the water area to be traced and detected, in this solution, it is directly matched with the corresponding sewage chemical composition types and product types of surrounding enterprises or factories, and for the polluted data after conversion, it is first classified and mapped to the polluted data before conversion and then matched with the corresponding sewage chemical composition types and product types of surrounding enterprises or factories.

[0062] 2. In the present invention, by collecting the content data of various pollutants in the water area to be traced and detected, it provides a data basis for subsequent determination of whether the polluted data in the water area to be traced and detected exceeds the corresponding threshold, and further determines whether water pollution traceability is required.

[0063] 3. In the present invention, the honey badger optimization algorithm is adopted to perform multiple iterative optimizations on the classification center data of each converted comprehensive water pollution data classification matrix in the set of converted comprehensive water pollution data classification matrices. Therefore, as the iteration progresses, the sum of the Euclidean distances between each data in each converted comprehensive water pollution data classification matrix in the set of converted comprehensive water pollution data classification matrices and the iteratively obtained classification center data becomes smaller and smaller, indicating that the iteratively obtained classification center data is more accurate.

[0064] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following will briefly introduce the drawings required for describing the embodiments. Obviously, the drawings in the following description are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0066] Figure 1 It is a schematic flow chart of a water pollution source tracing detection method based on data mining according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the invention with reference to the drawings in the embodiments of the invention. Obviously, the described embodiments are only some embodiments of the invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the invention without creative efforts belong to the scope of protection of the invention.

[0068] In the description of the present invention, it should be understood that the terms "openings", "upper", "lower", "top", "middle", "inner", etc. indicating the orientation or positional relationship are only for the convenience of describing the invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the invention.

[0069] Embodiment 1

[0070] Please refer to Figure 1 , this embodiment is a water pollution source tracing detection method based on data mining, including the following steps:

[0071] S1. Collect multiple groups of existing polluted water body data;

[0072] The said S1 includes the following steps:

[0073] S11. Set the water pollution type set , a1. a 2 respectively represent chemical pollution and garbage pollution; according to the set of water pollution types, multiple chemical component types and garbage types of the polluted water body before water pollution conversion are set to obtain the set of chemical component types before conversion and the set of garbage types before conversion , 、 respectively represent the i th chemical component type and garbage type of the set polluted water body before water pollution conversion, 、 respectively represent the total numbers of chemical component types and garbage types of the set polluted water body before water pollution conversion;

[0074] Then, according to the set of water pollution types, multiple chemical component types and garbage types of the polluted water body after water pollution conversion are set to obtain the set of chemical component types after conversion and the set of garbage types after conversion , 、 respectively represent the i th chemical component type and garbage type of the set polluted water body after water pollution conversion, 、 respectively represent the total numbers of chemical component types and garbage types of the set polluted water body after water pollution conversion;

[0075] S12. Set multiple types of water pollution conversion influencing factors and multiple types of pollutant distribution parameters to obtain the set of conversion influencing factor types and the set of pollutant distribution parameter types , 、 respectively represent the i th type of water pollution conversion influencing factor and pollutant distribution parameter type set, b 、 respectively represent the total numbers of water pollution conversion influencing factors and pollutant distribution parameter types set;

[0076] S13. According to the set of chemical component types before conversion, the set of garbage types before conversion, the set of chemical component types after conversion, the set of garbage types after conversion, the set of conversion influencing factor types, and the set of distribution parameter types after conversion, collect multiple groups of data on the content of various chemical components before conversion and corresponding distribution parameter data, the mass of various garbage before conversion and corresponding distribution parameter data, the content of various chemical components after conversion and corresponding distribution parameter data, the mass of various garbage after conversion and corresponding distribution parameter data, and conversion influencing factor data in the existing polluted water bodies to obtain the historical data matrix of the content of chemical components before conversion , historical chemical distribution parameter data matrix before conversion , historical waste mass data matrix before conversion , historical waste distribution parameter data matrix before conversion , historical chemical component content data matrix after conversion , historical chemical distribution parameter data matrix after conversion , historical waste mass data matrix after conversion , historical waste distribution parameter data matrix after conversion and historical conversion influencing factor data matrix ; , , , , , , , , are as follows respectively,

[0077] ; ; ;

[0078] ; ; ;

[0079] ; ; ;

[0080] Among them, , , , , , , , , respectively represent the i th type of chemical component content data before conversion and the corresponding distribution parameter data, the j th type of waste mass data before conversion and the corresponding distribution parameter data, the j th type of chemical component content data after conversion and the corresponding distribution parameter data, the j th type of waste mass data after conversion and the corresponding distribution parameter data, and the j th type of conversion influencing factor data in the j th group of existing polluted water bodies collected; represents the total number of groups of existing polluted water bodies collected;

[0081] S2. Merge the data before the pollution occurrence conversion in the existing polluted water body data, and then perform clustering to obtain the classification matrix set of the comprehensive water pollution data before conversion and the data matrix of the comprehensive classification center of water pollution before conversion;

[0082] The S2 includes the following steps:

[0083] S21. Horizontally merge the historical chemical composition content data matrix before conversion, the historical chemical distribution parameter data matrix before conversion, the historical garbage quality data matrix before conversion, the historical garbage distribution parameter data matrix before conversion, and the historical conversion influencing factor data matrix to obtain the comprehensive water pollution data matrix c 1; then horizontally merge the historical chemical composition content data matrix after conversion, the historical chemical distribution parameter data matrix after conversion, the historical garbage quality data matrix after conversion, and the historical garbage distribution parameter data matrix after conversion to obtain the comprehensive water pollution data matrix c 2; c 1. c 2 are as follows,

[0084] ; ;

[0085] Among them, c 1ij and c 2ij respectively represent the i th type of comprehensive water pollution data before conversion and the water pollution data after conversion in the j th group of existing polluted water bodies collected, and respectively represent the total quantities of the comprehensive water pollution data before conversion and the water pollution data after conversion in the i th group of existing polluted water bodies collected; , ;

[0086] S22. Perform a clustering operation on the comprehensive water pollution data matrix c 1 to obtain the classification matrix set of the comprehensive water pollution data before conversion and the data matrix of the comprehensive classification center of water pollution before conversion , represents the c th classification matrix of the comprehensive water pollution data before conversion obtained by performing a clustering operation on the comprehensive water pollution data matrix i 1, represents the total quantity of the classification matrices of the comprehensive water pollution data before conversion obtained by performing a clustering operation on the comprehensive water pollution data matrix c 1; and Are as follows,

[0087] ; ;

[0088] Among them, represents in the j th group of the existing polluted water bodies, the k th type of water pollution data before comprehensive conversion, d i represents the total number of groups of the water pollution data before comprehensive conversion of the existing polluted water bodies assigned to; represents the c 1 water pollution comprehensive data matrix before conversion, the i th classification center dataset obtained by clustering operation, the j th type of water pollution data before comprehensive conversion;

[0089] In S22, clustering operation is performed on the c 1 water pollution comprehensive data matrix before conversion to obtain a water pollution comprehensive data classification matrix set before conversion and a water pollution comprehensive classification center data matrix before conversion including the following steps:

[0090] S221. Construct a water pollution data clustering honey badger population , represents the i th honey badger in the water pollution data clustering honey badger population, represents the scale of the water pollution data clustering honey badger population; set the maximum number of iterations of the water pollution data clustering honey badger population to be and the current number of iterations to be , which are respectively recorded as the maximum number of iterations of water pollution data clustering and the current number of iterations of water pollution data clustering; the search space dimension of the water pollution data clustering honey badger population is ;

[0091] S222. Randomly select several rows of data from the c 1 water pollution comprehensive data matrix before conversion as the initial position matrix of each honey badger in the water pollution data clustering honey badger population to obtain a first initial position matrix set , e 1k represents the initial position matrix of the k th honey badger in the water pollution data clustering honey badger population; as follows,

[0092] ;

[0093] Among them, e 1kij represents e 1k the component on the c first i classification center dataset of the water pollution comprehensive data matrix before conversion j in the

[0094] S223. Calculate the Euclidean distance between each row of data in the water pollution comprehensive data matrix c 1 and each row of data in the initial position matrix of the k th honey badger in the water pollution data clustering honey badger population before conversion, and obtain the first Euclidean distance matrix ; as follows,

[0095] ;

[0096] Among them, represents the Euclidean distance between the c th i row of data in the water pollution comprehensive data matrix k 1 and the j th row of data in the initial position matrix of the

[0097] According to the first Euclidean distance matrix construct the fitness function of the k th honey badger in the water pollution data clustering honey badger population before conversion ; as follows,

[0098] ;

[0099] S224. Start iteration. Before iteration, set the current iteration number of the water pollution data clustering to 1; in the first round of iteration, use the fitness function k of the th honey badger in the water pollution data clustering honey badger population before conversion to calculate the fitness value of the initial position matrix of each honey badger in the first initial position matrix set, and obtain the first fitness value set; take the maximum fitness value in the first fitness value set and the corresponding initial position matrix of the honey badger as the first global best fitness and the first global best position respectively; update the initial position matrix of each honey badger in the first initial position matrix set according to the first global best fitness and the first global best position; after the update is completed, add 1 to the current iteration number of the water pollution data clustering and enter the next round of iteration;

[0100] In each other iteration process, the fitness function of the k th honey badger in the honey badger population clustered by the water pollution data before conversion is used. Calculate the fitness values of the position matrices of each honey badger in the honey badger population clustered by the water pollution data before conversion updated in the previous iteration process to obtain a second set of fitness values; take the maximum fitness in the second set of fitness values and the corresponding position matrix of the honey badger as the second global best fitness and the second global best position respectively; update the position matrices of each honey badger in the honey badger population clustered by the water pollution data before conversion updated in the previous iteration process according to the second global best fitness and the second global best position; after the update is completed, add 1 to the current iteration number of the water pollution data clustering and enter the next iteration.

[0101] S225. When is satisfied, stop the iteration to obtain the first final global best position; otherwise, continue the iteration until is satisfied; take the first final global best position as the data matrix of the comprehensive classification center of the water pollution before conversion; classify the comprehensive data matrix of the water pollution before conversion c 1 according to the data matrix of the comprehensive classification center of the water pollution before conversion to obtain a set of comprehensive data classification matrices of the water pollution before conversion.

[0102] The honey badger optimization algorithm can perform global search in the search space by simulating the foraging behavior of honey badgers to find the global optimal solution; it can quickly converge to the optimal solution during the search process, improving the efficiency of the algorithm; and it is not sensitive to the initial value, has strong robustness, and can find the optimal solution under different initial conditions; based on the above advantages, in this solution, the honey badger optimization algorithm is used to perform multiple iterative adjustments on the multiple clustering center data of the comprehensive data matrix of the water pollution before conversion at the same time, and the overall dispersion degree of the comprehensive data matrix of the water pollution before conversion is used as the fitness function. Therefore, as the iteration progresses, the overall dispersion degree of the comprehensive data matrix of the water pollution before conversion becomes smaller and smaller, indicating that the clustering effect is getting better and better.

[0103] S3. Classify the data after the pollution occurrence conversion in the existing polluted water body data according to the set of comprehensive data classification matrices of the water pollution before conversion and calculate the classification center data of each category to obtain the data matrix of the comprehensive classification center of the water pollution after conversion.

[0104] The S3 includes the following steps:

[0105] S31. According to the set of comprehensive data classification matrices of the water pollution before conversion and the comprehensive data matrix of the water pollution before conversion c 1 for the comprehensive data matrix of the water pollution after conversion c2 Classify to obtain the classified matrix set of the converted comprehensive water pollution data , represents the c th converted comprehensive water pollution data classification matrix obtained by classifying the converted comprehensive water pollution data matrix i 2; as follows,

[0106] ;

[0107] Among them, represents the j th type of comprehensive converted water pollution data in the k rd group of existing polluted water bodies;

[0108] S32. Calculate the central data of each classification matrix in the classified matrix set of the converted comprehensive water pollution data to obtain the converted comprehensive water pollution classification central data matrix ; as follows,

[0109] ;

[0110] Among them, represents the c th type of comprehensive converted water pollution data in the i th classification center data set obtained by classifying the converted comprehensive water pollution data matrix j 2;

[0111] The S32 includes the following steps:

[0112] S321. Construct the honey badger population for calculating the classification center of the converted water pollution data , represents the i th honey badger in the honey badger population for calculating the classification center of the converted water pollution data, represents the scale of the honey badger population for calculating the classification center of the converted water pollution data; set the maximum number of iterations of the honey badger population for calculating the classification center of the converted water pollution data to be and the current number of iterations to be , which are respectively recorded as the maximum number of iterations for calculating the classification center of the water pollution data and the current number of iterations for calculating the classification center of the water pollution data; the search space dimension of the honey badger population for calculating the classification center of the converted water pollution data is ;

[0113] S322. Randomly select a row of data from each classification matrix in the classified matrix set of the converted comprehensive water pollution data multiple times as the initial position matrix of each honey badger in the honey badger population for calculating the classification center of the converted water pollution data to obtain the second initial position matrix set , e 2k represents the initial position matrix of the k th honey badger in the honey badger population calculated by the converted water pollution data classification center; as follows,

[0114] ;

[0115] Among them, e 2kij represents e 2k in the classification center dataset of the j th component on the dimension of the comprehensively converted water pollution data;

[0116] S323. Calculate the Euclidean distance between each row of data in each classification matrix of the converted water pollution comprehensive data classification matrix set and the corresponding classification center data in the initial position matrix of the k th honey badger in the honey badger population calculated by the converted water pollution data classification center to obtain the second Euclidean distance matrix ; as follows,

[0117] ;

[0118] Among them, represents the Euclidean distance between the i th row of data in the j th classification matrix of the converted water pollution comprehensive data classification matrix set and the k th row of data in the initial position matrix of the i th honey badger in the honey badger population calculated by the converted water pollution data classification center;

[0119] According to the second Euclidean distance matrix construct the fitness function of the k th honey badger in the honey badger population calculated by the converted water pollution data classification center ; as follows,

[0120] ;

[0121] S324. Start the iteration. Before the iteration, set the current iteration number of the water pollution data classification center to 1; in the first round of iteration, use the fitness function of the k th honey badger in the honey badger population calculated by the converted water pollution data classification center Calculate the fitness values of the initial position matrices of each honey badger in the second set of initial position matrices to obtain a third set of fitness values; take the maximum fitness value in the third set of fitness values and the corresponding initial position matrix of the honey badger as the third global best fitness and the third global best position respectively; update the initial position matrices of each honey badger in the second set of initial position matrices according to the third global best fitness and the third global best position; after the update is completed, the water pollution data classification center calculates the current iteration number add 1 and enter the next round of iteration;

[0122] In each other round of iteration, use the transformed water pollution data classification center to calculate the fitness function of the k th honey badger in the honey badger population Calculate the fitness values of the position matrices of each honey badger in the honey badger population calculated by the transformed water pollution data classification center updated in the previous round of iteration to obtain a fourth set of fitness values; take the maximum fitness value in the fourth set of fitness values and the corresponding position matrix of the honey badger as the fourth global best fitness and the fourth global best position respectively; update the position matrices of each honey badger in the honey badger population calculated by the transformed water pollution data classification center updated in the previous round of iteration according to the fourth global best fitness and the fourth global best position; after the update is completed, the water pollution data classification center calculates the current iteration number add 1 and enter the next round of iteration;

[0123] S325. When is satisfied, stop the iteration to obtain the second final global best position; otherwise, continue the iteration until is satisfied; take the second final global best position as the transformed water pollution comprehensive classification center data matrix;

[0124] S4. Cooperate with the existing polluted water body data to collect the pollution data in the water area to be traced and detected;

[0125] S4 includes the following steps:

[0126] S41. Set the water area to be traced and detected; set a number of water quality detection points in the water area to be traced and detected and collect the image data of the water area to be traced and detected above to obtain a water quality detection point set and the image data of the water area to be traced;

[0127] Detect and collect the corresponding content data of each pre-conversion chemical component, post-conversion chemical component content data, and corresponding distribution parameter data at each water quality detection point in the water quality detection point set according to the pre-conversion chemical component type set, post-conversion chemical component type set, and pollutant distribution parameter type set, to obtain the pre-conversion chemical component content data matrix to be traced, the post-conversion chemical component content data matrix to be traced, and the post-chemical conversion distribution parameter data set to be traced , respectively represent the distribution parameter data of the i th type of chemical pollutant in the water area to be traced after conversion;

[0128] Calculate the average value of the pre-conversion chemical component content and the post-conversion chemical component content at each water quality detection point according to the pre-conversion chemical component content data matrix to be traced and the post-conversion chemical component content data matrix to be traced, to obtain the pre-conversion chemical component content average data set to be traced , the post-conversion chemical component content average data set to be traced , , respectively represent the average value of the pre-conversion chemical component content and the post-conversion chemical component content of the i th type at each water quality detection point in the water quality detection point set;

[0129] S42. Use the pre-trained CNN neural network model to identify the pre-conversion garbage type and the post-conversion garbage type of the water area image data to be traced, and calculate the corresponding garbage quality data and distribution parameter data according to the recognition results, to obtain the pre-conversion garbage quality data set to be traced , the post-conversion garbage quality data set to be traced and the post-conversion garbage distribution parameter data set to be traced , , , respectively represent the calculated pre-conversion garbage quality data of the i th type, the post-conversion garbage quality data of the i th type, and the corresponding distribution parameter data;

[0130] S43. Collect the conversion influence factor data in the water area to be traced according to the conversion influence factor type set, to obtain the conversion influence factor data set to be traced , respectively represent the i th type of conversion influence factor data collected from the water area to be traced;

[0131] S5. Match the pollution data before conversion in the pollution data of the water area to be traced and detected with the types of chemical components discharged and product types corresponding to the surrounding enterprises or factories to obtain the first pollution source enterprise and the second pollution source enterprise;

[0132] S5 includes the following steps:

[0133] S51. Collect the types of chemical components discharged and product types corresponding to the enterprises or factories around the water area to be traced and detected to obtain a matrix of chemical component discharge types and a matrix of product types;

[0134] S52. Set corresponding chemical component content exceedance thresholds and waste quality exceedance thresholds according to the pre-conversion chemical component type set, pre-conversion waste type set, post-conversion chemical component type set, and post-conversion waste type set to obtain a pre-conversion chemical component content threshold set, a pre-conversion waste quality threshold set, a post-conversion chemical component content threshold set, and a post-conversion waste quality threshold set;

[0135] S53. When there is chemical component content average data in the pre-conversion chemical component content average data set of the water area to be traced that is greater than or equal to the corresponding pre-conversion chemical component content threshold in the pre-conversion chemical component content threshold set, match the chemical component type corresponding to the pre-conversion chemical component content average data set of the water area to be traced with the matrix of chemical component discharge types to obtain the first pollution source enterprise; otherwise, there is no need to match with the matrix of chemical component discharge types;

[0136] When there is waste quality data in the pre-conversion waste quality data set of the water area to be traced that is greater than or equal to the corresponding pre-conversion waste quality threshold in the pre-conversion waste quality threshold set, match the pre-conversion waste type corresponding to the pre-conversion waste quality data set of the water area to be traced with the matrix of product types to obtain the second pollution source enterprise; otherwise, there is no need to match with the matrix of product types;

[0137] S6. Classify the pollution data after conversion in the pollution data of the water area to be traced and detected according to the post-conversion water pollution comprehensive classification center data matrix to obtain a pre-conversion water pollution comprehensive data set of the water area to be traced; match the pre-conversion water pollution comprehensive data set of the water area to be traced with the types of chemical components discharged and product types corresponding to the surrounding enterprises or factories to obtain the third pollution source enterprise and the fourth pollution source enterprise;

[0138] S6 includes the following steps:

[0139] S61. When there is a converted chemical component content average data in the dataset of the chemical component content to be traced after conversion that is greater than or equal to the corresponding chemical component content threshold in the converted chemical component content threshold set, or when there is a converted waste mass data in the dataset of the waste mass to be traced after conversion that is greater than or equal to the corresponding converted waste mass threshold in the converted waste mass threshold set, merge the dataset of the chemical component content average to be traced after conversion, the dataset of the waste mass to be traced after conversion, the dataset of the distribution parameters after chemical conversion to be traced, the dataset of the waste distribution parameters to be traced after conversion, and the dataset of the influencing factors of the conversion to be traced to obtain a comprehensive dataset of water pollution to be traced; otherwise, no merging is required.

[0140] S62. Calculate the Euclidean distance between each row of data in the comprehensive dataset of water pollution to be traced and the data matrix of the converted comprehensive classification center of water pollution to obtain a set of Euclidean distances; record the converted comprehensive classification center data of the water pollution before conversion in the data matrix of the converted comprehensive classification center of water pollution before conversion corresponding to the smallest Euclidean distance in the set of Euclidean distances as the comprehensive dataset of water pollution before conversion to be traced.

[0141] When there is a chemical component content data in the comprehensive dataset of water pollution before conversion to be traced that is greater than or equal to the corresponding set of converted chemical component content thresholds in the converted chemical component content threshold set, match the chemical component types corresponding to the converted chemical component content in the comprehensive dataset of water pollution before conversion to be traced with the matrix of the chemical component types of the pollution discharge to obtain the third pollution source enterprise; otherwise, no matching is required.

[0142] When there is a waste mass data in the comprehensive dataset of water pollution before conversion to be traced that is greater than or equal to the corresponding converted waste mass threshold in the converted waste mass threshold set, match the waste types corresponding to the converted waste mass data in the comprehensive dataset of water pollution before conversion to be traced with the matrix of product types to obtain the fourth pollution source enterprise; otherwise, no matching with the matrix of product types is required.

[0143] S63. Conduct water pollution tracing based on the first pollution source enterprise, the second pollution source enterprise, the third pollution source enterprise, and the fourth pollution source enterprise.

[0144] Embodiment 2

[0145] This embodiment discloses a water pollution tracing and detection system based on data mining. The system can implement the method of the above embodiment, including a water pollution characteristic type setting module, an existing water pollution data collection module, a pre-conversion data merging and clustering module, a post-conversion data classification center calculation module, a water pollution data collection module for the water area to be traced, a surrounding enterprise data collection module, a primary tracing and matching module, a pollution data classification module after conversion to be traced, and a secondary tracing and matching module.

[0146] The water pollution characteristic type setting module is used to set multiple chemical component types, garbage types, conversion influencing factor types, and pollutant distribution parameter types before and after water pollution conversion of polluted water bodies, and obtain a set of chemical component types before conversion, a set of garbage types before conversion, a set of chemical component types after conversion, a set of garbage types after conversion, a set of conversion influencing factor types, and a set of pollutant distribution parameter types;

[0147] The existing water pollution data collection module is used to collect multiple groups of existing polluted water body data according to the set of chemical component types before conversion, the set of garbage types before conversion, the set of chemical component types after conversion, the set of garbage types after conversion, the set of conversion influencing factor types, and the set of pollutant distribution parameter types, and obtain a historical chemical component content data matrix before conversion, a historical chemical distribution parameter data matrix before conversion, a historical garbage quality data matrix before conversion, a historical garbage distribution parameter data matrix before conversion, a historical chemical component content data matrix after conversion, a historical chemical distribution parameter data matrix after conversion, a historical garbage quality data matrix after conversion, a historical garbage distribution parameter data matrix after conversion, and a historical conversion influencing factor data matrix;

[0148] The data merging and clustering module before conversion is used to merge and then cluster the historical chemical component content data matrix before conversion, the historical chemical distribution parameter data matrix before conversion, the historical garbage quality data matrix before conversion, the historical garbage distribution parameter data matrix before conversion, and the historical conversion influencing factor data matrix, and obtain a set of comprehensive water pollution data classification matrices before conversion and a comprehensive water pollution classification center data matrix before conversion;

[0149] The data classification center calculation module after conversion is used to classify and calculate the classification center for the historical chemical component content data matrix after conversion, the historical chemical distribution parameter data matrix after conversion, the historical garbage quality data matrix after conversion, and the historical garbage distribution parameter data matrix after conversion according to the set of comprehensive water pollution data classification matrices before conversion, and obtain a comprehensive water pollution classification center data matrix after conversion;

[0150] The pollution data collection module for the water area to be traced and detected is used to collect pollution data in the water area to be traced and detected according to the set of chemical component types before conversion, the set of garbage types before conversion, the set of chemical component types after conversion, the set of garbage types after conversion, the set of conversion influencing factor types, and the set of pollutant distribution parameter types, and obtain a data set of distribution parameters after chemical conversion to be traced, an average data set of chemical component contents before conversion to be traced, an average data set of chemical component contents after conversion to be traced, a data set of garbage quality before conversion to be traced, a data set of garbage quality after conversion to be traced, a data set of garbage distribution parameters after conversion to be traced, and a data set of conversion influencing factors to be traced;

[0151] The surrounding enterprise data collection module is used to collect the types of sewage chemical components and product types corresponding to the enterprises or factories around the water area to be traced and detected, and obtain a sewage chemical component type matrix and a product type matrix;

[0152] The first traceability matching module is used to match the average data set of chemical component contents before traceability conversion, the garbage mass data set before traceability conversion with the sewage chemical component type matrix and the product type matrix to obtain the first pollution source enterprise and the second pollution source enterprise;

[0153] The pollution data classification module after traceability conversion is used to classify the average data set of chemical component contents after traceability conversion, the garbage mass data set after traceability conversion, the distribution parameter data set after chemical conversion of traceability, the garbage distribution parameter data set after traceability conversion and the influencing factor data set after traceability conversion according to the data matrix of the comprehensive water pollution classification center after conversion to obtain the comprehensive water pollution data set before traceability conversion;

[0154] The second traceability matching module is used to match the comprehensive water pollution data set before traceability conversion with the sewage chemical component type matrix and the product type matrix to obtain the third pollution source enterprise and the fourth pollution source enterprise.

[0155] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0156] The preferred embodiments of the invention disclosed above are only used to help explain the invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principle and practical application of the invention, so that those skilled in the relevant technical field can understand and utilize the invention well.

Claims

1. A water pollution source tracing and detection method based on data mining, characterized in that It includes the following steps: S1. Collect multiple groups of existing polluted water body data; S2. Merge the data before the pollution occurrence conversion in the existing polluted water body data and then perform clustering to obtain the pre-conversion water pollution comprehensive data classification matrix set and the pre-conversion water pollution comprehensive classification center data matrix; S3. Classify the post-conversion water pollution comprehensive data matrix according to the pre-conversion water pollution comprehensive data classification matrix set to obtain the post-conversion water pollution comprehensive data classification matrix set; calculate the central data of each classification matrix in the post-conversion water pollution comprehensive data classification matrix set to obtain the post-conversion water pollution comprehensive classification center data matrix. It includes the following steps: S31. Construct a honey badger population by calculating the converted water pollution data classification center; set the maximum number of iterations for calculating the honey badger population by the converted water pollution data classification center as and the current number of iterations as which are respectively denoted as the maximum number of iterations for calculating the water pollution data classification center and the current number of iterations for calculating the water pollution data classification center; S32. Randomly select a row of data from each classification matrix in the post-conversion water pollution comprehensive data classification matrix set multiple times as the initial position matrix of each honey badger in the honey badger population for calculating the post-conversion water pollution data classification center, and obtain the second initial position matrix set; S33. Calculate the Euclidean distance between each row of data in each classification matrix of the post-conversion water pollution comprehensive data classification matrix set and the corresponding classification center data in the initial position matrix of the k-th honey badger in the honey badger population for calculating the post-conversion water pollution data classification center, and construct the fitness function of the k-th honey badger in the honey badger population for calculating the post-conversion water pollution data classification center; S34. Start iteration; in each iteration process, use the fitness function of the k-th honey badger in the honey badger population for calculating the post-conversion water pollution data classification center to calculate the fitness value of the position matrix of each honey badger in the honey badger population for calculating the post-conversion water pollution data classification center updated in the previous iteration process, and update the position matrix of each honey badger in the honey badger population for calculating the post-conversion water pollution data classification center updated in the previous iteration process; S35. When occurs, stop the iteration to obtain the second final global optimal position; otherwise, continue the iteration until is reached; take the second final global optimal position as the data matrix of the converted comprehensive classification center for water pollution. S4. Collect the pollution data in the water area to be traced and detected in cooperation with the existing polluted water body data; S5. Match the pollution data before the conversion in the pollution data in the water area to be traced and detected with the corresponding sewage chemical component types and product types of surrounding enterprises or factories to obtain the first pollution source enterprise and the second pollution source enterprise; S6. Classify the pollution data after the conversion in the pollution data in the water area to be traced and detected according to the post-conversion water pollution comprehensive classification center data matrix to obtain the pre-conversion water pollution comprehensive data set to be traced; match the pre-conversion water pollution comprehensive data set to be traced with the corresponding sewage chemical component types and product types of surrounding enterprises or factories to obtain the third pollution source enterprise and the fourth pollution source enterprise.

2. The water pollution source tracing and detection method based on data mining according to claim 1, wherein The S1 includes the following steps: S11. Collect the content data of various pre-conversion chemical components and their corresponding distribution parameter data, the mass data of various pre-conversion garbage and their corresponding distribution parameter data, the content data of various post-conversion chemical components and their corresponding distribution parameter data, the mass data of various post-conversion garbage and their corresponding distribution parameter data, and the conversion influencing factor data in multiple groups of existing polluted water bodies, to obtain the historical pre-conversion chemical component content data matrix, the historical pre-conversion chemical distribution parameter data matrix, the historical pre-conversion garbage mass data matrix, the historical pre-conversion garbage distribution parameter data matrix, the historical post-conversion chemical component content data matrix, the historical post-conversion chemical distribution parameter data matrix, the historical post-conversion garbage mass data matrix, the historical post-conversion garbage distribution parameter data matrix, and the historical conversion influencing factor data matrix.

3. The water pollution source tracing and detection method based on data mining according to claim 2, characterized in that, The above S2 includes the following steps: S21. Merge the historical pre-conversion chemical component content data matrix, the historical pre-conversion chemical distribution parameter data matrix, the historical pre-conversion garbage mass data matrix, the historical pre-conversion garbage distribution parameter data matrix, and the historical conversion influencing factor data matrix to obtain the pre-conversion water pollution comprehensive data matrix; Horizontally merge the historical post-conversion chemical component content data matrix, the historical post-conversion chemical distribution parameter data matrix, the historical post-conversion garbage mass data matrix, and the historical post-conversion garbage distribution parameter data matrix to obtain the post-conversion water pollution comprehensive data matrix; S22. Perform a clustering operation on the pre-conversion water pollution comprehensive data matrix to obtain the pre-conversion water pollution comprehensive data classification matrix set and the pre-conversion water pollution comprehensive classification center data matrix.

4. The water pollution source tracing and detection method based on data mining according to claim 3, characterized in that: In S22, the honey badger optimization algorithm is used to perform the clustering operation on the pre-conversion water pollution comprehensive data matrix.

5. The water pollution source tracing and detection method based on data mining according to claim 4, characterized in that The above S4 includes the following steps: S41. Set the water area to be traced and detected; at each water quality detection point, detect and collect the mean value of the content data of each pre-conversion chemical component, the mean value of the content data of the post-conversion chemical component, and the corresponding distribution parameter data, to obtain the dataset of the post-conversion distribution parameters of the chemical components to be traced, the dataset of the average content of the pre-conversion chemical components to be traced, and the dataset of the average content of the post-conversion chemical components to be traced; S42. Identify the garbage categories before and after conversion in the water area to be traced and detected, and calculate the corresponding garbage mass data and distribution parameter data, to obtain the dataset of the pre-conversion garbage mass to be traced, the dataset of the post-conversion garbage mass to be traced, and the dataset of the post-conversion garbage distribution parameters to be traced; S43. Collect the conversion influencing factor data in the water area to be traced and detected to obtain the dataset of the conversion influencing factors to be traced.

6. The water pollution source tracing and detection method based on data mining according to claim 5, characterized in that, The above S5 includes the following steps: S51. Collect the types of sewage chemical components and product types corresponding to the enterprises or factories around the water area to be traced and detected, to obtain the sewage chemical component type matrix and the product type matrix; S52. Set the threshold sets of the pre-conversion chemical component content, the pre-conversion garbage mass, the post-conversion chemical component content, and the post-conversion garbage mass; S53. Compare the average dataset of chemical component contents before conversion to be traced with the threshold set of chemical component contents before conversion, and match the chemical component types corresponding to the average dataset of chemical component contents before conversion to be traced with the matrix of sewage chemical component types to obtain the first polluting source enterprise; Compare the average dataset of garbage masses before conversion to be traced with the threshold set of garbage masses before conversion, and match the garbage types before conversion corresponding to the average dataset of garbage masses before conversion to be traced with the product type matrix to obtain the second polluting source enterprise.

7. A water pollution source tracing and detection method based on data mining according to claim 6, characterized in that, S6 includes the following steps: S61. Merge the average dataset of chemical component contents after conversion to be traced, the average dataset of garbage masses after conversion to be traced, the dataset of distribution parameters after chemical conversion to be traced, the dataset of distribution parameters of garbage after conversion to be traced, and the dataset of influencing factors of conversion to be traced according to the threshold set of chemical component contents after conversion, the threshold set of garbage masses after conversion, the average dataset of chemical component contents after conversion to be traced, and the average dataset of garbage masses after conversion to be traced to obtain the comprehensive dataset of water pollution to be traced; S62. Calculate the Euclidean distances between each row of data in the comprehensive dataset of water pollution to be traced and the data matrix of the comprehensive classification center of water pollution after conversion to obtain a set of Euclidean distances; Denote the data of the comprehensive classification center of water pollution before conversion in the data matrix of the comprehensive classification center of water pollution before conversion corresponding to the smallest Euclidean distance in the set of Euclidean distances as the comprehensive dataset of water pollution before conversion to be traced; Match the comprehensive dataset of water pollution before conversion to be traced with the matrix of sewage chemical component types and the product type matrix in combination with the threshold set of chemical component contents before conversion and the threshold set of garbage masses before conversion to obtain the third polluting source enterprise and the fourth polluting source enterprise; S63. Conduct water pollution tracing based on the first polluting source enterprise, the second polluting source enterprise, the third polluting source enterprise, and the fourth polluting source enterprise.

Citation Information

Patent Citations

  • Method for rapidly tracing to water pollution source

    CN102661939B

  • Pesticide effect improvement air detection system based on pollution tracing source

    CN108508149A

  • A system for identifying noise-polluted regions using data mining approaches and clustering techniques.

    DE202022101216U1