A method, device and medium for tracing the source of water pollution in a river basin

By dividing the river basin into assessment sections, conducting data preprocessing, cluster analysis and correlation analysis, and combining the one-dimensional hydrological and water quality model with the random forest algorithm, the difficult problem of analyzing the sources of water pollutants in the river basin was solved, and accurate tracing of pollution sources and effective governance recommendations were achieved.

CN119477643BActive Publication Date: 2025-09-19重庆市生态环境大数据应用中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411610437.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-09-19
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively analyze the sources of water pollutants in river basins, making it difficult to provide scientific basis to support decision makers in the supervision and maintenance of river basins.

Method used

By dividing the river basin into assessment sections, obtaining data from each section and preprocessing it, cluster analysis and correlation analysis are used to identify the characteristics of water quality changes. A one-dimensional hydrological and water quality model is constructed to calculate the pollution contribution, and the random forest algorithm is used to calculate the pollution importance and determine the spatial location of the pollution source.

Benefits of technology

It has achieved accurate tracing of water pollutants in river basins, provided effective advice to the Environmental Bureau on water management, and determined the spatial location and importance of pollution sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477643B_ABST
    Figure CN119477643B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of water pollution treatment technology, and in particular relates to a method, equipment, and medium for tracing the source of water pollution in a river basin. First, river data of each assessment section is obtained and preprocessed; then, a preset cluster analysis method is used to cluster the preprocessed river data in the assessment section to analyze the water quality change characteristics, and the dominant factors of the water quality change characteristics are analyzed by correlation analysis; then, a one-dimensional hydrological and water quality model is constructed, and the contribution of each assessment section to the water quality pollution of its downstream assessment section is calculated through the one-dimensional hydrological and water quality model to determine the spatial location of the pollution source; finally, the pollution importance of each upstream water assessment section of the national assessment section is calculated through the constructed random forest algorithm, and the characteristics of the assessment section with the highest pollution importance to the national assessment section are extracted. The present invention can solve the problem of the inability to effectively analyze the water quality pollutant problem in river basins in the existing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of water pollution treatment, and in particular relates to a method, equipment and medium for tracing the source of water pollution in a river basin. Background Art

[0002] Water pollution refers to the phenomenon that harmful substances enter natural water bodies, causing water quality to deteriorate, thereby affecting human health and the ecological environment; the sources of water pollution are diverse, mainly including industrial wastewater, agricultural non-point source pollution, domestic sewage and accidental leakage; water pollution has a profound impact on the environment and social economy. On the one hand, it increases the threat to aquatic life and destroys the ecological balance; on the other hand, water pollution affects human health.

[0003] With the development and continuous improvement of sensor technology and data sharing technology, the acquisition of river basin water environment data has become more convenient and the amount of data has exploded. Therefore, there is an urgent need for water environment big data analysis to achieve accurate analysis of river basin water quality health, especially for the source analysis of river basin water pollution, which can provide decision makers with a scientific basis for supervision and maintenance within the river basin on the source of pollutants that affect water quality standards.

[0004] However, in actual operations, it is still difficult to effectively analyze the sources of pollutants that affect the water quality in river basins. Therefore, there is an urgent need to provide a river basin water pollution source tracing technology to provide effective recommendations for achieving water quality standards. Summary of the Invention

[0005] The technical problem solved by the present invention is to provide a method, equipment and medium for tracing the source of water pollution in a river basin, so as to solve the problem in the prior art of being unable to effectively analyze the water quality pollutants in the river basin.

[0006] The basic solution provided by the present invention is a method for tracing the source of water pollution in a river basin, comprising:

[0007] S1: Divide the river basin into assessment sections and obtain river data for each assessment section; the assessment sections include national assessment sections;

[0008] S2: preprocessing the acquired river data;

[0009] S3: Perform cluster analysis on the water quality change characteristics of the pre-processed river data in the assessment section using the preset cluster analysis method, and analyze the dominant factors of the water quality change characteristics using the correlation analysis method;

[0010] S4: Construct a one-dimensional hydrological and water quality model, calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source;

[0011] S5: The constructed random forest algorithm is used to calculate the pollution importance of each upstream water assessment section of the national assessment section, and the characteristics of the assessment section with the highest pollution importance to the national assessment section are extracted.

[0012] Further, the S4 includes:

[0013] S4-1: Extract the locations of lateral inflows and pollution sources from the preprocessed river data, divide the target river basin into several segments, with the locations of lateral inflows and pollution sources as segment heads. The flow velocity and flow rate of each river basin are constant, and a one-dimensional water quality model is generated.

[0014] S4-2: Based on the one-dimensional water quality model, calculate the water pollution concentration at the downstream assessment section where the upstream water flows. The calculation formula is:

[0015]

[0016] Among them, X i-1 is the length of the river basin in the section, u i-1 is the average flow velocity of the river basin in the section, K is the pollutant attenuation coefficient, C 1,i represents the pollutant concentration of water flowing from the upstream assessment section to the downstream assessment section i, C 2,i-1 represents the pollutant concentration of the river water flowing downstream from the previous section of section i;

[0017] S4-3: Obtain the pollutant concentration distribution map of each assessment section based on the calculation results and determine the spatial location of the pollution source.

[0018] Further, the S5 includes:

[0019] S5-1: Extract the water quality change characteristics X of each assessment section after correlation analysis of water quality changes j , generate the feature set M of each assessment section;

[0020] S5-2: Call the random forest algorithm to calculate the water quality change characteristics X in the characteristic concentration of each assessment section j The Gini index is calculated as follows:

[0021]

[0022] Among them, K means there are k categories, p mk Indicates the proportion of category k in node m;

[0023] S5-3: Calculate water quality change characteristics X j The importance of node m is calculated as follows:

[0024]

[0025] Among them, GI l Indicates water quality change characteristics X j The Gini index before the node m branches, GI r Indicates water quality change characteristics X j Gini index after branching at node m;

[0026] S5-4: Water quality change characteristics X j The node that appears in the decision tree i is in the feature set M, and the water quality change feature X is calculated. j The importance of the i-th decision tree is calculated as follows:

[0027]

[0028] If there are n decision trees, then the water quality change feature X j The importance calculation formula is:

[0029]

[0030] Characteristics of water quality changes X j The importance of is normalized, specifically:

[0031]

[0032] Among them, the denominator Represents the sum of the water quality variation characteristics of the entire data set.

[0033] Furthermore, the S5 further includes:

[0034] S5-5: Sort the water quality change characteristics of each assessment section in descending order according to the calculated importance;

[0035] S5-6: Preset the elimination ratio, perform elimination operations with the preset elimination ratio according to the importance of water quality change characteristics, and generate a new feature set;

[0036] S5-7: Preset the number of retained features, and repeat S5-5 to S5-6 for the new feature set until a final feature set with the preset number of retained features remains;

[0037] S5-8: Select the feature set with the lowest Gini index in the final feature set of each assessment section as the assessment section feature set with the highest importance in affecting the water quality of the national assessment section.

[0038] Furthermore, the river data in S1 includes natural impact data and human impact data on water quality, the natural impact data includes geographical data, meteorological data and hydrological data, and the human impact data includes pollution source data, tunnel topography data and socio-economic data.

[0039] Furthermore, the preprocessing in S2 includes data missing filling processing, data spatial linking processing and data spatiotemporal consistency processing.

[0040] Furthermore, the S3 includes:

[0041] S3-1: Use K-means cluster analysis to conduct cluster analysis on the water quality change characteristics of river data at stations where each assessment section is located in the river basin;

[0042] S3-2: Based on the preprocessed river data, the correlation analysis method is used to analyze the correlation between water quality and the surrounding industries and rainfall in the river basin to obtain the dominant factors of water quality change characteristics.

[0043] A river basin water pollution source tracing device, applied to the above-mentioned river basin water pollution source tracing method, comprises:

[0044] Data acquisition module: used to obtain river data of each assessment section according to the assessment sections of the divided river basin; the assessment sections include national assessment sections;

[0045] Data preprocessing module: used to preprocess the acquired river data;

[0046] Water quality change characteristic analysis module: used to perform water quality change characteristic cluster analysis on the pre-processed river data in the assessment section according to the preset cluster analysis method to obtain water quality change characteristics;

[0047] Influencing factor analysis module: used to analyze the dominant factors of water quality change characteristics through correlation analysis;

[0048] Pollution contribution calculation module: used to calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the constructed one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source;

[0049] Pollution importance calculation module: used to calculate the pollution importance of each upstream water assessment section of the national examination section through the constructed random forest algorithm, and extract the characteristics of the assessment section with the highest pollution importance to the national examination section.

[0050] An electronic device includes a processor and a memory, wherein the memory stores programs or instructions, and the processor executes the above-mentioned method for tracing the source of water pollution in a river basin by calling the programs or instructions stored in the memory.

[0051] A computer-readable storage medium stores a program or instruction, wherein the program or instruction enables a computer to execute the above-mentioned method for tracing the source of water pollution in a river basin.

[0052] The principles and advantages of the present invention are: in response to the problem that the existing technology cannot effectively analyze water pollution, the present invention first divides the river basin into assessment sections, and based on the divided river basin, according to the spatial logical chain of assessment section-upstream water station-national assessment section, constructs a characteristic analysis model of river basin water quality anomalies and a related analysis model of the dominant factors of anomalies, so as to obtain the main factors of water quality changes in each assessment section of the river basin; at the same time, the application uses a one-dimensional water quality model to obtain the upstream water pollution contribution of each assessment section station, determines the spatial location of the pollution source affecting the water quality in the river basin, and then uses the random forest algorithm to calculate the pollution importance of each upstream water assessment section of the national assessment section, providing the Environmental Bureau with effective suggestions for water control. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of an embodiment of the present invention;

[0054] Figure 2 is a functional block diagram of an embodiment of the present invention;

[0055] Figure 3 Schematic diagram of a one-dimensional water quality model according to an embodiment of the present invention;

[0056] Figure 4 Schematic diagram of the electronic device structure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The following is further described in detail through specific implementation methods:

[0058] The symbols in the drawings of the specification include: electronic device 400 , processor 401 , memory 402 , input device 403 , and output device 404 .

[0059] The embodiment is basically as shown in the attached Figure 1 Shown: A method for tracing the source of water pollution in a river basin, comprising:

[0060] S1: Divide the assessment sections of the river basin and obtain river data of each assessment section; the assessment sections include national assessment sections; in this embodiment, the assessment section division of the target river basin is generated through the geographic data of the river basin. Specifically, the DEM data, river network vector data, basin town and street data, land use type data, automatic monitoring station vector data, etc. of the target river basin are obtained, and the water collection units of the monitoring stations in the river basin are divided into assessment sections.

[0061] The river data of the river basin obtained in this application include natural impact data and human impact data on water quality. Specifically, in addition to the geographical data obtained above, which belong to natural impact data, the natural impact data also include meteorological data and hydrological data. For example, meteorological data include historical meteorological data and forecast meteorological data, and the specific contents include the daily maximum, minimum and average temperatures, precipitation, evaporation, wind speed, relative humidity, radiation, etc. at meteorological battle points to meet the needs of water quality testing.

[0062] Hydrological data is obtained by accessing the government’s water information interface, including flow and water level data.

[0063] Human impact data include pollution source data, topographic data, and socioeconomic data. Pollution source data refers to the pollution source conditions of industrial enterprises, large-scale livestock and poultry breeding, and sewage treatment plants within the catchment area of ​​the river basin. The pollution source data of industrial enterprises include the enterprise's water intake, wastewater discharge, discharge destination, major pollutant generation, and discharge, etc. The pollution source conditions of sewage treatment plants include treatment scale, drainage destination, discharge, and online monitoring data of their major pollutants; livestock and poultry farm data include sewage generation and discharge conditions, etc.

[0064] The tunnel topographic data in this application refers to the river channel and riverbed topography DEM data in the Dadi 2000 coordinate system and the measured river width data.

[0065] Socioeconomic data include GDP, industrial structure data, population data, etc. of the river basin.

[0066] To better illustrate the details of the river data in this application, the following Table 1 is provided:

[0067] Table 1

[0068]

[0069]

[0070] S2: Preprocessing the acquired river data; in this embodiment, preprocessing includes data missing filling processing, data spatial link processing, and data spatiotemporal consistency processing. For data missing filling processing, for example, if the 1-hour cumulative rainfall data in the meteorological data is complete, but the 3-hour, 6-hour, 12-hour, and 24-hour cumulative rainfall data are null values, the missing data can be filled by adding the 1-hour cumulative rainfall data; for another example, if the water quality parameters in the water quality data are missing in a certain time period, the outlier detection algorithm program is used to fill the missing water quality parameters in the short period; the specific steps of data filling used in this application disclosed in this embodiment are:

[0071] Data filling:

[0072]

[0073]

[0074]

[0075] The spatial link processing of data is to establish a spatial logical chain between the monitoring data and the assessment section sites in the river basin, and to establish a unique identification code. For example, for site 1, use R1 code and assign longitude and latitude coordinates to the identification code. Site 2 is coded with R2, and so on.

[0076] Data spatiotemporal consistency processing is to process river data with time and space attributes in a temporal and spatial consistency manner, such as rainfall data, pollution source data, etc., and water quality data in a temporal and spatial consistency manner.

[0077] S3: Perform cluster analysis on the water quality change characteristics of the pre-processed river data in the assessment section using a preset cluster analysis method, and analyze the dominant factors of the water quality change characteristics using a correlation analysis method; S3 includes:

[0078] S3-1: Use K-means cluster analysis to conduct cluster analysis on the water quality change characteristics of river data at stations where each assessment section is located in the river basin;

[0079] S3-2: Based on the preprocessed river data, the correlation analysis method is used to analyze the correlation between water quality and the surrounding industries and rainfall in the river basin to obtain the dominant factors of water quality change characteristics.

[0080] In this embodiment, the analysis process of the K-means cluster analysis method in S3-1 is:

[0081] Step 1: Select K initial centroids. The initial centroids can be randomly selected, and each centroid represents a class.

[0082] Step 2: For each remaining sample point, calculate its Euclidean distance to each centroid, and assign it to the cluster with the centroid with the smallest distance between them, and calculate the centroid of each new cluster;

[0083] Step 3: After all sample points are divided, recalculate the location of the centroid of each cluster according to the division situation, then iteratively calculate the distance from each sample point to the centroid of each cluster, and re-divide all sample points;

[0084] Step 4: Repeat steps 2 and 3 until the centroid no longer changes or the maximum number of iterations is reached.

[0085] Therefore, when analyzing the characteristics of water quality changes, this application uses the K-means clustering analysis method to cluster similar water quality data into the same cluster and divide dissimilar water quality data into different clusters, thereby analyzing the characteristic properties of different water quality data and the relationship between them.

[0086] In specific use, the K-means cluster analysis method is used to cluster the river data of the assessment section based on the time scale to identify the main types and pollution characteristics of water pollution in each assessment section. After cluster analysis, several groups of classified data can be obtained, and then these classified data are subjected to feature analysis. In this embodiment, the feature analysis includes: 1. Identifying the main water quality characteristics, main pollution factors, etc. of each pollution type based on the overall distribution of factors in each pollution type; 2. The time changes and time distribution characteristics of the pollution type; 3. The high incidence seasons of various pollution types and their time correlation with the dry season / flood season / normal water season; 4. The frequency of occurrence of various pollution types to determine the dominant pollution type of the section (the pollution type with the highest occurrence frequency).

[0087] The water pollution types in the main water quality change characteristics identified by cluster analysis are analyzed for dominant factors using the S3-2 correlation analysis method. The correlation analysis method in this application adopts Pearson correlation analysis. Pearson correlation analysis relies on the correlation coefficient to make judgments. Specifically, when the absolute value of the correlation coefficient is above 0.8, it is considered that there is a strong correlation between the two variables; when it is between 0.6-0.8, the two variables are considered to be strongly correlated; when it is between 0.4-0.6, the two variables are considered to be moderately correlated; when it is 0.2-0.4, the two variables are considered to be weakly correlated; when it is below 0.2, the two variables can be considered to be extremely weakly correlated or have no correlation.

[0088] Specific applications include: 1. Analyzing the impact of rainfall on water quality changes at each assessment section: Pearson correlation analysis is performed on the rainfall data of each assessment section and the water quality time series data. The higher the correlation coefficient, the greater the impact of rainfall on water quality. 2. Based on the distribution of industry, population, agriculture, livestock and poultry farming within the catchment area of ​​each assessment section, as well as wastewater discharge, Pearson correlation analysis is performed on the frequency of water quality exceeding standards at each assessment section in the basin to analyze the impact of industry, population, agriculture, and livestock and poultry farming on water quality exceeding standards. The higher the correlation coefficient, the greater the impact on water quality.

[0089] S4: Construct a one-dimensional hydrological and water quality model, calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source; S4 includes:

[0090] S4-1: Extract the side inflow location and pollution source location from the pre-processed river data, divide the target river basin into several segments, with the side inflow location and pollution source location as the segment head. The flow velocity and flow rate of each river basin are constant, and a one-dimensional water quality model is generated. The schematic diagram of the one-dimensional water quality model constructed in this embodiment is shown in FIG. Figure 3 As shown, where Q i represents the sewage flow into the river from section i, C i represents the pollutant concentration of sewage flowing into the river from section i, Q1, i represents the river flow from upstream to section i, C1, i represents the pollutant concentration of the river water flowing upstream to section i, Q2, i represents the river flow from section i to downstream, C2, i represents the pollutant concentration of the river water flowing downward from section i, and u represents the river flow velocity;

[0091] According to the principle of conservation of mass, we have:

[0092] C2, i Q2, i =C1, i Q1, i +C i Q i

[0093] S4-2: Based on the one-dimensional water quality model, calculate the water pollution concentration at the downstream assessment section where the upstream water flows. The calculation formula is:

[0094]

[0095] Among them, X i-1 is the length of the river basin in the section, u i-1 is the average flow velocity of the river basin in the section, K is the pollutant attenuation coefficient, C 1,i represents the pollutant concentration of water flowing from the upstream assessment section to the downstream assessment section i, C 2,i-1 represents the pollutant concentration of the river water flowing downstream from the section above section i. In this case, the pollutant concentration of the water quality downstream from the upstream water is calculated based on the river flow velocity and the pollutant attenuation coefficient of 0.2;

[0096] S4-3: Obtain the pollutant concentration distribution map of each assessment section based on the calculation results and determine the spatial location of the pollution source.

[0097] S5: The constructed random forest algorithm is used to calculate the pollution importance of each upstream water assessment section of the national examination section, and the characteristics of the assessment section with the highest pollution importance to the national examination section are extracted; S5 includes:

[0098] S5-1: Extract the water quality change characteristics X of each assessment section after correlation analysis of water quality changes j , generate the feature set M of each assessment section;

[0099] S5-2: Call the random forest algorithm to calculate the water quality change characteristics X in the characteristic concentration of each assessment section j The Gini index is calculated as follows:

[0100]

[0101] Among them, K means there are k categories, p mk Indicates the proportion of category k in node m;

[0102] S5-3: Calculate water quality change characteristics X j The importance of node m is calculated as follows:

[0103]

[0104] Among them, GI l Indicates water quality change characteristics X j The Gini index before the node m branches, GI r Indicates water quality change characteristics X j Gini index after branching at node m;

[0105] S5-4: Water quality change characteristics X j The node that appears in the decision tree i is in the feature set M, and the water quality change feature X is calculated. j The importance of the i-th decision tree is calculated as follows:

[0106]

[0107] If there are n decision trees, then the water quality change feature X j The importance calculation formula is:

[0108]

[0109] Characteristics of water quality changes X j The importance of is normalized, specifically:

[0110]

[0111] Among them, the denominator Represents the sum of the water quality variation characteristics of the entire data set;

[0112] S5-5: Sort the water quality change characteristics of each assessment section in descending order according to the calculated importance;

[0113] S5-6: Preset the elimination ratio, perform elimination operations with the preset elimination ratio according to the importance of water quality change characteristics, and generate a new feature set;

[0114] S5-7: Preset the number of retained features, and repeat S5-5 to S5-6 for the new feature set until a final feature set with the preset number of retained features remains;

[0115] S5-8: Select the feature set with the lowest Gini index in the final feature set of each assessment section as the assessment section feature set with the highest importance in affecting the water quality of the national assessment section.

[0116] Therefore, this application uses a one-dimensional water quality model to obtain the contribution of upstream water pollution to each assessment section site, determine the spatial location of pollution sources affecting water quality in the river basin, and then use the random forest algorithm to calculate the pollution importance of each upstream water assessment section of the national assessment section, providing the Environmental Protection Bureau with effective suggestions for water management.

[0117] like Figure 2 As shown, in another embodiment of this embodiment, a river basin water pollution source tracing device is also included, which is applied to the above-mentioned river basin water pollution source tracing method, including:

[0118] Data acquisition module: used to obtain river data of each assessment section according to the assessment sections of the divided river basin; the assessment sections include national assessment sections;

[0119] Data preprocessing module: used to preprocess the acquired river data;

[0120] Water quality change characteristic analysis module: used to perform water quality change characteristic cluster analysis on the pre-processed river data in the assessment section according to the preset cluster analysis method to obtain water quality change characteristics;

[0121] Influencing factor analysis module: used to analyze the dominant factors of water quality change characteristics through correlation analysis;

[0122] Pollution contribution calculation module: used to calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the constructed one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source;

[0123] Pollution importance calculation module: used to calculate the pollution importance of each upstream water assessment section of the national examination section through the constructed random forest algorithm, and extract the characteristics of the assessment section with the highest pollution importance to the national examination section.

[0124] Also included is an electronic device such as Figure 4 As shown, the electronic device 400 includes one or more processors 401 and a memory 402 .

[0125] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0126] The memory 402 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may run the program instructions to implement a river basin water pollution tracing method and / or other desired functions of any embodiment of the present invention described above. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.

[0127] In one example, electronic device 400 may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 403 may include, for example, a keyboard, a mouse, etc. Output device 404 may output various information to the outside, including warning information, braking force, etc. Output device 404 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.

[0128] Of course, to simplify, Figure 4 Only some of the components related to the present invention in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 400 may further include any other appropriate components according to specific application scenarios.

[0129] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of a river basin water pollution tracing method provided by any embodiment of the present invention.

[0130] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of embodiments of the present invention, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0131] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of a method for tracing the source of water pollution in a river basin provided by any embodiment of the present invention.

[0132] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0133] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for tracing the source of water pollution in a river basin, characterized by: include: S1: Divide the river basin into assessment sections and obtain river data for each assessment section; the assessment sections include national assessment sections; S2: preprocessing the acquired river data; S3: Perform cluster analysis on the water quality change characteristics of the pre-processed river data in the assessment section using the preset cluster analysis method, and analyze the dominant factors of the water quality change characteristics using the correlation analysis method; S4: Construct a one-dimensional hydrological and water quality model, calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source; S5: Calculate the pollution importance of each upstream water assessment section of the national assessment section through the constructed random forest algorithm, and extract the characteristics of the assessment section with the highest pollution importance to the national assessment section; The S4 includes: S4-1: Extract the locations of lateral inflows and pollution sources from the preprocessed river data, divide the target river basin into several segments, with the locations of lateral inflows and pollution sources as segment heads. The flow velocity and flow rate of each river basin are constant, and a one-dimensional water quality model is generated. S4-2: Based on the one-dimensional water quality model, calculate the water pollution concentration at the downstream assessment section where the upstream water flows. The calculation formula is: in, is the length of the river basin in the section, is the average flow velocity of the river basin in the section, K is the pollutant attenuation coefficient, It represents the pollutant concentration of the water flowing from the upstream assessment section to the downstream assessment section i, represents the pollutant concentration of the river water flowing downstream from the previous section of section i; S4-3: Obtain the pollutant concentration distribution map of each assessment section based on the calculation results and determine the spatial location of the pollution source; The S5 includes: S5-1: Extract the water quality change characteristics of each assessment section after correlation analysis , generate the feature set M of each assessment section; S5-2: Call the random forest algorithm to calculate the water quality change characteristics of the characteristic concentration of each assessment section The Gini index is calculated as follows: Among them, K means there are k categories, Indicates the proportion of category k in node m; S5-3: Calculate water quality change characteristics The importance of node m is calculated as follows: in, Indicates the characteristics of water quality changes The Gini index before the node m branches, Indicates the characteristics of water quality changes Gini index after branching at node m; S5-4: Characteristics of water quality changes The nodes appearing in the decision tree i are in the feature set M, and the water quality change characteristics are calculated. The importance of the i-th decision tree is calculated as follows: If there are n decision trees, then the water quality change characteristics The importance calculation formula is: Characteristics of water quality changes The importance of is normalized, specifically: Among them, the denominator Represents the sum of the water quality variation characteristics of the entire data set.

2. The method for tracing the source of water pollution in a river basin according to claim 1, characterized in that: The S5 further includes: S5-5: Sort the water quality change characteristics of each assessment section in descending order according to the calculated importance; S5-6: Preset the elimination ratio, perform elimination operations with the preset elimination ratio according to the importance of water quality change characteristics, and generate a new feature set; S5-7: Preset the number of retained features, and repeat S5-5 to S5-6 for the new feature set until a final feature set with the preset number of retained features remains; S5-8: Select the feature set with the lowest Gini index in the final feature set of each assessment section as the assessment section feature set with the highest importance in affecting the water quality of the national assessment section.

3. The method for tracing the source of water pollution in a river basin according to claim 2, characterized in that: The river data in S1 includes natural impact data and human impact data on water quality. The natural impact data includes geographical data, meteorological data and hydrological data, and the human impact data includes pollution source data, tunnel topography data and socio-economic data.

4. The method for tracing the source of water pollution in a river basin according to claim 3, characterized in that: The preprocessing in S2 includes data missing filling processing, data spatial link processing and data spatiotemporal consistency processing.

5. The method for tracing the source of water pollution in a river basin according to claim 4, characterized in that: The S3 includes: S3-1: Use K-means cluster analysis to conduct cluster analysis on the water quality change characteristics of river data at stations where each assessment section is located in the river basin; S3-2: Based on the preprocessed river data, the correlation analysis method is used to analyze the correlation between water quality and the surrounding industries and rainfall in the river basin to obtain the dominant factors of water quality change characteristics.

6. A river basin water pollution source tracing device, applied to a river basin water pollution source tracing method as described in any one of claims 1 to 5 above, characterized in that: include: Data acquisition module: used to obtain river data of each assessment section according to the assessment sections of the divided river basin; The assessment sections include national examination sections; Data preprocessing module: used to preprocess the acquired river data; Water quality change characteristic analysis module: used to perform water quality change characteristic cluster analysis on the pre-processed river data in the assessment section according to the preset cluster analysis method to obtain water quality change characteristics; Influencing factor analysis module: used to analyze the dominant factors of water quality change characteristics through correlation analysis; Pollution contribution calculation module: used to calculate the contribution of each assessment section to the water quality pollution of its downstream assessment section through the constructed one-dimensional hydrological and water quality model, and determine the spatial location of the pollution source; Pollution importance calculation module: used to calculate the pollution importance of each upstream water assessment section of the national examination section through the constructed random forest algorithm, and extract the characteristics of the assessment section with the highest pollution importance to the national examination section.

7. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores programs or instructions, and the processor executes a river basin water pollution source tracing method as described in any one of claims 1 to 5 by calling the programs or instructions stored in the memory.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute a method for tracing the source of water pollution in a river basin as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Water environment pollution analysis system and method based on big data

    CN112417788A

  • Air quality monitoring system and method

    US20220397520A1