A method for accurately identifying pollution sources in a karst groundwater river system

By comprehensively applying GIS, hydrogeological models and machine learning algorithms, the problem that traditional methods are difficult to accurately identify pollution sources in karst underground river systems is solved, and the accurate identification of pollution sources is achieved, and the efficiency of environmental protection and water resource management is improved.

CN119272162BActive Publication Date: 2025-06-27INST OF KARST GEOLOGY CAGS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411325417.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-06-27
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

Traditional pollution source identification methods are difficult to accurately identify pollution sources in karst underground river systems, and are time-consuming and labor-intensive, making it difficult to achieve real-time monitoring and rapid response.

Method used

The geographical information system (GIS), hydrogeological models and machine learning algorithms are used to collect historical monitoring data, geological and hydrological data, emission source data and location data to perform regional division, conformal prediction, correction and establish pollution source identification models to achieve accurate identification of pollution sources.

Benefits of technology

It has achieved efficient and accurate identification of pollution sources in karst underground river systems, and has important application value for environmental protection and water resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272162B_ABST
    Figure CN119272162B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for accurately identifying pollution sources in a karst groundwater river system, which includes collecting historical monitoring data, geological and hydrological data, emission source data, and location data of a research area; dividing the area into influence areas based on the geological and hydrological data and the location data, and performing conformal prediction on the historical monitoring data based on the emission source data of the influence areas to obtain influence data; correcting the influence data with the location data to obtain pollution source data, and establishing a pollution source identification model based on the pollution source data; inputting the data to be identified into the pollution source identification model and outputting an identification result. This method can not only improve the identification accuracy of pollution sources in the karst groundwater river system, but also has good interpretability and can be directly applied to the accurate identification system of pollution sources in the karst groundwater river system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of environmental science and hydrology, and particularly to a method for accurately identifying pollution sources in a karst groundwater system. Background Art

[0002] Due to its unique geological structure and hydrological characteristics, the karst groundwater system is highly sensitive and complex to the transmission of pollutants. These systems are usually composed of complex underground river networks, caves, and fissures. Once pollutants enter, they will spread rapidly, making it difficult to track and control. Traditional pollution source identification methods, such as simple monitoring and chemical analysis, often cannot accurately reveal the location and characteristics of pollution sources, especially in the highly heterogeneous environment of karst areas. In addition, these methods are usually time-consuming and laborious, and it is difficult to achieve real-time monitoring and rapid response. Therefore, developing a method that can quickly and accurately identify pollution sources in a karst groundwater system is crucial for the effective management and protection of these fragile water ecosystems. It is in this context that the present invention proposes an innovative method for accurately identifying pollution sources, aiming to improve the accuracy and efficiency of identification by comprehensively applying advanced technologies such as geographic information system (GIS), hydrogeological models, and machine learning algorithms. Summary of the Invention

[0003] The object of the present invention is to provide a method for accurately identifying pollution sources in a karst groundwater system.

[0004] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0005] The present invention includes the following steps:

[0006] Collect historical monitoring data, hydrogeological data, emission source data, and location data of the research area;

[0007] Based on the hydrogeological data and the location data, conduct regional division to obtain the influence area, and conduct conformal prediction on the historical monitoring data based on the emission source data of the influence area to obtain influence data;

[0008] Use the location data to correct the influence data to obtain pollution source data, and establish a pollution source identification model based on the pollution source data;

[0009] Input the data to be identified into the pollution source identification model and output the identification result.

[0010] Further, the method for collecting historical monitoring data, hydrogeological data, emission source data, and location data of the research area includes:

[0011] The historical monitoring data are the types and concentrations of pollutants. Geological and hydrological data are obtained through a GIS system. The geological and hydrological data include the boundaries of the underground river system and the velocity of the underground river. The emission source data are the historical emission sources and the corresponding emission source lists. The location data include the locations of the emission sources and the monitoring points.

[0012] Furthermore, a method for dividing a region into an impact region based on the geological and hydrological data and the location data includes:

[0013] Dividing the region within 5 kilometers upstream of the monitoring point into an identification region according to the geological and hydrological data and the location data.

[0014] Using the location data of the emission sources within the identification region to form data points, and clustering the data points to obtain c clustering clusters:

[0015]

[0016] Where is the objective function, n is the number of data points, c is the number of clustering clusters, η ij is the membership degree of the i-th data point to the j-th clustering cluster, representing the spatial correlation of the data point to the clustering cluster, ‖·‖ represents the Euclidean distance, M is the fuzzy factor, 1.5 < M < 2.5, Λ j is the fuzzy penalty coefficient of the j-th clustering cluster, X i is the i-th data point, u j is the center point of the clustering cluster j, C j is the set of data points in the clustering cluster j, |C j | is the number of data points in the clustering cluster j, h is the bandwidth parameter, 0.4 < h < 1, is the i-th noise term, and the calculation formula for the membership degree:

[0017]

[0018] where δ is used to avoid a zero denominator, 0 < δ < 0.001, and the clustering center update formula:

[0019]

[0020] where is the center point of the clustering cluster j after the (s + 1)-th iteration, is the membership degree of the data point i to the clustering cluster j after the (s + 1)-th iteration,

[0021] Dividing the data points into the clustering cluster with the highest membership degree, and taking the c clustering clusters obtained by clustering as the impact region.

[0022] Furthermore, a method for performing conformal prediction on the historical monitoring data based on the emission source data of the influence area to obtain influence data includes:

[0023] Fitting the historical monitoring data according to the emission source data of the influence area:

[0024] Taking the emission source data as feature data and the historical monitoring data as the target value to establish a regression model, and the quantile loss function of the regression model:

[0025]

[0026] where is the quantile loss of the k-th feature data for the q-th quantile, s k and are respectively the target value and the predicted value of the k-th feature data, and q is the target quantile.

[0027] The regularization loss function of the regression model:

[0028]

[0029] where n is the number of feature data, T is the number of leaf nodes, γ is the regularization parameter, λ is the coefficient of the L2 regularization term, and w j is the weight of leaf node j.

[0030] Selecting the feature data based on the split gain, and the split gain of the decision tree:

[0031]

[0032] where I L and I R are respectively the sample sets of the left subtree and the right subtree after the decision tree split, I is the sample set of the decision tree before splitting, and ζ is the penalty coefficient. and are respectively the first-order derivatives of the quantile loss functions of the left subtree and the right subtree. and are respectively the second-order derivatives of the quantile loss functions of the left subtree and the right subtree. and are respectively the first-order derivative and the second-order derivative of the quantile loss function before splitting.

[0033] Calculating the SHAP value of the emission source data for the target value as the influence degree, and taking the emission source data with the influence degree greater than the preset influence degree threshold as the influence data.

[0034] Furthermore, a method for correcting the influence data with the position data to obtain pollution source data includes:

[0035] Using the impact data as the independent variable and the historical monitoring data as the dependent variable, calculate the distance data of the emission source relative to the nearest underground river system boundary and the monitoring point according to the location data of the emission source and the monitoring point, and establish a geostatistical regression model:

[0036]

[0037] Where, Y i is the i-th dependent variable, which is the historical monitoring data of the i-th monitoring point, and X ij is the j-th independent variable of the i-th dependent variable, j is the index of the emission source in the impact data, p is the number of independent variables, β0(u i ,v i ) is the intercept term of the i-th monitoring point, β j (u i ,v i ) is the spatial geographical location function of the j-th independent variable at the i-th monitoring point (u i ,v i ), ε i is the error term, (u i ,v i ) is the location data of the i-th monitoring point, β flow (u i ,v i ) is the regression coefficient of the underground river velocity at (u i ,v i ), V i is the underground river velocity of the i-th monitoring point,

[0038] Estimate the spatial geographical location function using local least squares:

[0039]

[0040] Where, is the local least squares estimate value of β(u o ,v o ) at the estimation point (u o ,v o ), β(u o ,v o ) is the spatial geographical location function matrix at the estimation point (u o ,v o ), β(u o ,v o ) = (β0(u o ,v o ), β1(u o ,v o ),..., β p (u o ,v o ))T , X T is the transpose of the independent variable matrix X, and W(u o , v o ) is the spatial weight matrix, where the element w i (u o , v o ) is the spatial weight of the i-th monitoring point at the estimation point (u o , v o ). Y is the dependent variable matrix,

[0041] Determine the spatial weight through the kernel function:

[0042]

[0043] where w i (u j , v j ) is the spatial weight of the i-th monitoring point at the j-th emission source (u j , v j ). d river,ij is the distance data from the j-th emission source to the boundary of the underground river system of the nearest i-th monitoring point, and d mon,ij is the distance data from the i-th monitoring point to the j-th emission source. h1 and h2 are the bandwidth parameters of d river,ij and d mon,ij selected according to the AICc information criterion,

[0044] Take the product of the spatial weight and the emission source data as the pollution index and normalize it. The emission sources with a pollution index greater than the preset pollution threshold are regarded as pollution sources.

[0045] Furthermore, a method for establishing a pollution source identification model based on the pollution source data includes:

[0046] Establish a data set using impact data, location data, historical detection data, and pollution source data, where the impact data, location data, and historical detection data are feature variables, and the pollution source data is the target variable. Learn the mapping relationship between the feature variables and the target variable in the data set through a machine learning algorithm to establish a pollution source identification model.

[0047] In a second aspect, an embodiment of the present application also provides an electronic device, including:

[0048] A processor; and a memory arranged to store computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the method steps described in the first aspect.

[0049] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, cause the electronic device to execute the method steps described in the first aspect.

[0050] The beneficial effects of the present invention are as follows:

[0051] The present invention is a method for accurately identifying pollution sources in a karst groundwater river system. Compared with the prior art, the present invention has the following technical effects:

[0052] The method for accurately identifying pollution sources in a karst groundwater river system provided by the present invention realizes the accurate identification of pollution sources in the karst groundwater river system by comprehensively using various data and advanced mathematical models, and has the characteristics of high efficiency and accuracy, and has important application value for environmental protection and water resource management. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of the steps of a method for accurately identifying pollution sources in a karst groundwater river system according to the present invention;

[0054] Figure 2 is a schematic structural diagram of an electronic device in an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention, but not to limit the present invention.

[0056] A method for accurately identifying pollution sources in a karst groundwater river system according to the present invention includes the following steps:

[0057] As Figure 1 shown, in this embodiment, the following steps are included:

[0058] Collect historical monitoring data, geological and hydrological data, emission source data, and location data of the research area;

[0059] In the actual assessment, historical monitoring data: the type of pollutant is heavy metal, and the concentration is 0.05 mg / L; geological and hydrological data: the boundary of the underground river system is formed by connecting the points (109.8°, 22.8°), (110.2°, 22.6°), (110.4°, 23.4°) and (109.6°, 23.2°), and the underground river velocity is 0.5 m / s; emission source data: historical emission sources: emission source list: the heavy metal emission of Factory A is 100 kg / year, the heavy metal emission of Factory B is 50 kg / year, the heavy metal emission of Factory C is 70 kg / year, and the heavy metal emission of Factory D is 40 kg / year; location data: emission source locations: Factory A, (110.12°, 23.12°), Factory B, (110.13°, 23.13°), Factory C, (110.14°, 23.14°), Factory D, (110.15°, 23.15°), monitoring point location: 110.1° east longitude, 23.1° north latitude;

[0060] Based on the geological and hydrological data and the location data, the affected area is obtained by regional division, and conformal prediction is performed on the historical monitoring data based on the emission source data of the affected area to obtain the impact data;

[0061] In the actual assessment, Factory A, B, and C are within the identified area. Factory A is the affected area 1, Factory B and C are the affected area 2, and the impact data is the emission source data of Factory A, B, and C;

[0062] The position data is used to correct the impact data to obtain the pollution source data, and a pollution source identification model is established based on the pollution source data;

[0063] In the actual assessment, the pollution indices of Factory A, B, and C are 1, 0.3, and 0.21 respectively, the preset pollution threshold is 0.25, and the pollution source data is Factory A and Factory B;

[0064] The data to be identified is input into the pollution source identification model, and the identification result is output.

[0065] In this embodiment, the method for collecting historical monitoring data, geological and hydrological data, emission source data, and location data of the research area includes:

[0066] The historical monitoring data is the type and concentration of pollutants. The geological and hydrological data is obtained through the GIS system. The geological and hydrological data includes the boundary of the underground river system and the underground river velocity. The emission source data is the historical emission sources and the corresponding emission source list. The location data includes the locations of the emission sources and the monitoring points.

[0067] In this embodiment, the method for obtaining the affected area by regional division based on the geological and hydrological data and the location data includes:

[0068] Divide the area within 5 kilometers upstream of the monitoring point into an identification area according to geological and hydrological data and location data.

[0069] Use the location data of the emission sources in the identification area to form data points, and perform clustering on the data points to obtain c clustering clusters:

[0070]

[0071] Among them, is the objective function, n is the number of data points, c is the number of clustering clusters, η ij is the membership degree of the i-th data point to the j-th clustering cluster, representing the spatial correlation between the data point and the clustering cluster, ‖·‖ represents the Euclidean distance, M is the fuzzy factor, 1.5 < M < 2.5, Λ j is the fuzzy penalty coefficient of the j-th clustering cluster, X i is the i-th data point, u j is the center point of the clustering cluster j, C j is the set of data points in the clustering cluster j, |C j | is the number of data points in the clustering cluster j, h is the bandwidth parameter, 0.4 < h < 1, is the i-th noise term, and the calculation formula for the membership degree:

[0072]

[0073] Among them, δ is used to avoid the denominator being zero, 0 < δ < 0.001, and the clustering center update formula:

[0074]

[0075] Among them is the center point of the clustering cluster j after s + 1 steps of iteration, is the membership degree of the data point i to the clustering cluster j after s + 1 steps of iteration,

[0076] Assign the data point to the clustering cluster with the highest membership degree, and use the c clustering clusters obtained by clustering as the impact area.

[0077] In this embodiment, a method for performing conformal prediction on the historical monitoring data based on the emission source data of the impact area to obtain impact data includes:

[0078] Fit the historical monitoring data according to the emission source data of the impact area:

[0079] Use the emission source data as feature data and the historical monitoring data as the target value to establish a regression model. The quantile loss function of the regression model:

[0080]

[0081] Among them is the quantile loss of the k-th feature data for the q-th quantile, s k and are the target value and the predicted value of the k-th feature data respectively, q is the target quantile,

[0082] Regularization loss function of the regression model:

[0083]

[0084] Among them, n is the number of feature data, T is the number of leaf nodes, γ is the regularization parameter, λ is the coefficient of the L2 regularization term, w j is the weight of leaf node j,

[0085] Select feature data based on the split gain, split gain of the decision tree:

[0086]

[0087] Among them, I L and I R are the sample sets of the left subtree and the right subtree after the decision tree split respectively, I is the sample set of the decision tree before splitting, ζ is the penalty coefficient, and are the first-order derivatives of the quantile loss functions of the left subtree and the right subtree respectively, and are the second-order derivatives of the quantile loss functions of the left subtree and the right subtree respectively, and are the first-order derivative and the second-order derivative of the quantile loss function before splitting respectively,

[0088] Calculate the SHAP value of the emission source data for the target value as the influence degree, and take the emission source data with the influence degree greater than the preset influence degree threshold as the influence data, and the preset influence degree threshold is 0.65.

[0089] In this embodiment, the method for correcting the influence data by using the position data to obtain the pollution source data includes:

[0090] Use the influence data as the independent variable and the historical monitoring data as the dependent variable, calculate the distance data of the emission source relative to the nearest underground river system boundary and the monitoring point according to the position data of the emission source and the monitoring point, and establish a geographical regression model:

[0091]

[0092] Among them, Y i is the i-th dependent variable, which is the historical monitoring data of the i-th monitoring point, X ijis the j-th independent variable of the i-th dependent variable, where j is the index of the emission source in the influencing data, p is the number of independent variables, and β0(u i ,v i ) is the intercept term of the i-th monitoring point, and β j (u i ,v i ) is the spatial geographical location function of the j-th independent variable at the i-th monitoring point (u i ,v i ), ε i is the error term, (u i ,v i ) is the location data of the i-th monitoring point, and β flow (u i ,v i ) is the regression coefficient of the underground river velocity at (u i ,v i ), V i is the underground river velocity of the i-th monitoring point.

[0093] The local least squares method is used to estimate the spatial geographical location function:

[0094]

[0095] where is the local least squares estimated value of β(u o ,v o ) at the estimation point (u o ,v o ), β(u o ,v o ) is the matrix of the spatial geographical location function at the estimation point (u o ,v o ), β(u o ,v o ) = (β0(u o ,v o ), β1(u o ,v o ),..., β p (u o ,v o )) T , X T is the transpose of the independent variable matrix X, W(u o ,v o ) is the spatial weight matrix, and the element w i (u o ,v o ) is the spatial weight of the i-th monitoring point at the estimation point (u o ,v o ), and Y is the dependent variable matrix.

[0096] Determine the spatial weight through the kernel function:

[0097]

[0098] where w i (u j , v j ) is the spatial weight of the i-th monitoring point at the j-th emission source (u j , v j ), d river,ij is the distance data from the j-th emission source to the boundary of the underground river system of the nearest i-th monitoring point, d mon,ij is the distance data from the i-th monitoring point to the j-th emission source, h1 and h2 are the bandwidth parameters of d river,ij and d mon,ij selected according to the AICc information criterion,

[0099] Take the product of the spatial weight and the emission source data as the pollution index and normalize it, and regard the emission source with the pollution index greater than the preset pollution threshold as the pollution source.

[0100] In this embodiment, the method for establishing a pollution source identification model based on the pollution source data includes:

[0101] Use the impact data, location data, historical detection data, and pollution source data to establish a data set, where the impact data, location data, and historical detection data are feature variables, and the pollution source data is the target variable. Learn the mapping relationship between the feature variables and the target variable in the data set through a machine learning algorithm to establish a pollution source identification model.

[0102] Figure 2 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 2 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0103] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, Figure 2 only a bidirectional arrow is used in Figure 2 , but it does not mean that there is only one bus or one type of bus.

[0104] Memory, used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0105] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a precise identification device for pollution sources of a karst underground river system at the logical level. The processor executes the program stored in the memory and is specifically used to execute any one of the foregoing precise identification methods for pollution sources of a karst underground river system.

[0106] The above is as in the present application Figure 1A method for accurately identifying pollution sources of a karst groundwater river system disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0107] The electronic device can also execute Figure 1 a method for accurately identifying pollution sources of a karst groundwater river system in Figure 1 and implement the functions of the illustrated embodiment, which will not be elaborated herein in the embodiments of the present application.

[0108] The embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, execute any of the above-mentioned methods for accurately identifying pollution sources of a karst groundwater river system.

[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0110] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0113] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0114] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flashRAM). The memory is an example of a computer-readable medium.

[0115] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0116] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0117] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for accurately identifying pollution sources in a karst underground river system, characterized in that: The following steps are involved: Collect historical monitoring data, geological and hydrological data, emission source data and location data for the study area; Performing regional division based on the geological and hydrological data and the location data to obtain an impact area, and performing conformal prediction on the historical monitoring data based on the emission source data of the impact area to obtain impact data; Using the position data to correct the impact data to obtain pollution source data, and establishing a pollution source identification model based on the pollution source data; Inputting the data to be identified into the pollution source identification model and outputting the identification result; The method for performing conformal prediction on the historical monitoring data based on the emission source data of the impact area to obtain the impact data comprises: Fitting historical monitoring data based on emission source data in the affected area: The emission source data is used as the characteristic data and the historical monitoring data is used as the target value to establish a regression model. The quantile loss function of the regression model is: ; in is the kth feature data for the Quantile loss of quantiles, and are the target value and predicted value of the kth feature data, respectively. is the target quantile, Regularized loss function for regression model: ; Where n is the number of feature data, T is the number of leaf nodes, γ is the regularization parameter, and λ is the coefficient of the L2 regularization term. is the weight of leaf node j, Feature data is selected based on split gain, split gain of decision tree: ; in and are the sample sets of the left subtree and the right subtree after the decision tree splits, is the sample set of the decision tree before splitting, is the penalty coefficient, and are the first-order derivatives of the quantile loss function of the left subtree and the right subtree, respectively. and are the second-order derivatives of the quantile loss function for the left and right subtrees, respectively. and are the first and second derivatives of the quantile loss function before splitting, The SHAP value of the emission source data on the target value is calculated as the influence, and the emission source data with an influence greater than a preset influence threshold is taken as the influence data.

2. According to claim 1, a method for accurately identifying pollution sources in a karst underground river system is characterized in that: Methods for collecting historical monitoring data, geological and hydrological data, emission source data, and location data for the study area, including: The historical monitoring data include the types and concentrations of pollutants. The geological and hydrological data are obtained through the GIS system. The geological and hydrological data include the boundaries of the underground river system and the underground river flow rate. The emission source data include the historical emission sources and the corresponding emission source list. The location data include the location of the emission source and the monitoring point.

3. According to claim 1, a method for accurately identifying pollution sources in a karst underground river system is characterized in that: The method for obtaining the affected area by performing regional division based on the geological and hydrological data and the location data comprises: The area within 5 km upstream of the monitoring point is divided into identification areas based on geological and hydrological data and location data. The location data of emission sources in the identified area are used to form data points, and the data points are clustered to obtain c clusters: ; in, is the objective function, n is the number of data points, is the number of clusters, is the membership of the i-th data point to the j-th cluster, indicating the spatial correlation of the data point to the cluster. represents the Euclidean distance, M is the fuzzy factor, 1.5<M<2.5, is the fuzzy penalty coefficient of the jth cluster, It is data points, is the center point of cluster j, is the set of data points in cluster j, is the number of data points in cluster j, h is the bandwidth parameter, 0.4<h<1, is the ith noise term, The calculation formula of membership degree is: ; in, Used to avoid denominator being zero, 0< <0.001, cluster center update formula: ; in is the center point of cluster j after iteration s+1 steps, is the membership of data point i to cluster j after iteration s+1 steps, The data points are divided into the clusters with the highest membership, and the c clusters obtained by clustering are taken as the influence areas.

4. According to claim 1, a method for accurately identifying pollution sources in a karst underground river system is characterized in that: The method of using the position data to correct the impact data to obtain pollution source data includes: Using impact data as independent variables and historical monitoring data as dependent variables, the distance data of emission sources relative to the nearest underground river system boundary and monitoring points are calculated based on the location data of emission sources and monitoring points, and a geographic regression model is established: ; in, It is The dependent variable is Historical monitoring data of monitoring points, It is The dependent variable independent variables, is the index of the emission source in the impact data, p is the number of independent variables, For the The intercept term of the monitoring points, It is The independent variable is Monitoring points The spatial geographic location function of is the error term, For the The location data of the monitoring points, The underground river flows at a speed of The regression coefficient at is the underground river velocity at the ith monitoring point, The spatial geographic location function is estimated using local least squares: ; in, for At the estimated point The local least squares estimate at , For estimated points The spatial geographic location function matrix at , is the independent variable matrix The transpose of is a spatial weight matrix, where the elements It is Monitoring points at the estimated point The spatial weight at is the dependent variable matrix, Determine the spatial weights through the kernel function: ; in For the The monitoring point is Emission sources The spatial weight at For the emission source to the nearest The distance data of the underground river system boundary at each monitoring point, For the Monitoring point to The distance data of each emission source, and are selected according to the AICc information criterion. and The bandwidth parameter, The product of the spatial weight and the emission source data is taken as the pollution index and normalized, and the emission source with a pollution index greater than the preset pollution threshold is regarded as the pollution source.

5. According to claim 1, a method for accurately identifying pollution sources in a karst underground river system is characterized in that: The method for establishing a pollution source identification model based on the pollution source data includes: A data set is established using impact data, location data, historical detection data and pollution source data, in which impact data, location data and historical detection data are feature variables and pollution source data is the target variable. The mapping relationship between the feature variables and the target variables in the data set is learned through a machine learning algorithm to establish a pollution source identification model.

6. An electronic device comprising: processor; as well as A memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 5.

7. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, enables the electronic device to execute the method according to any one of claims 1 to 5.