Method and apparatus for analyzing sources of pollutants in groundwater, device, medium

CN121298873BActive Publication Date: 2026-06-02TECH CENT FOR SOIL AGRI & RURAL ECOLOGY & ENVIRONMENT MINIST OF ECOLOGY & ENVIRONMENT

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TECH CENT FOR SOIL AGRI & RURAL ECOLOGY & ENVIRONMENT MINIST OF ECOLOGY & ENVIRONMENT
Filing Date
2025-12-11
Publication Date
2026-06-02

Smart Images

  • Figure CN121298873B_ABST
    Figure CN121298873B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of to groundwater's pollutant source analysis method and device, equipment, medium, it is related to environmental chemistry analysis technical field.The method comprises: water sample collection is carried out in the candidate pollution source sampling point of target groundwater, candidate pollution source water sample is acquired, and water sample collection is carried out in the downstream sampling point of target groundwater, downstream water sample is acquired;Wherein, target groundwater flows from candidate pollution source sampling point to downstream sampling point;According to candidate pollution source water sample, chemical composition feature extraction is carried out, and first chemical feature is obtained, and according to downstream water sample, chemical composition feature extraction is carried out, and second chemical feature is obtained;Similarity is calculated according to first chemical feature and second chemical feature, and target chemical feature similarity is obtained;According to target chemical feature similarity, target pollution source sampling point is determined from candidate pollution source sampling point.The embodiment of the application can improve the accuracy of pollutant source analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of environmental chemical analysis technology, and in particular to a method, apparatus, equipment, and medium for analyzing the sources of pollutants in groundwater. Background Technology

[0002] Source analysis is a method for identifying the origins and propagation pathways of pollutants in the environment. For example, in groundwater pollution control, determining the source of pollutants in groundwater is crucial for developing effective remediation strategies. However, current source analysis methods are mostly qualitative, such as using colorimetry or experience-based spectral comparison methods to analyze groundwater samples to determine the source of pollutants. These methods are highly subjective, making it difficult to accurately determine the source of pollutants and unsuitable for complex groundwater chemical environments.

[0003] Therefore, improving the accuracy of pollutant source analysis has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main objective of this application is to provide a method, apparatus, equipment, and medium for analyzing the sources of pollutants in groundwater, aiming to improve the accuracy of pollutant source analysis.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for analyzing the sources of pollutants in groundwater, the method comprising:

[0006] Water samples are collected at candidate pollution source sampling points of the target groundwater to obtain candidate pollution source water samples, and water samples are also collected at downstream sampling points of the target groundwater to obtain downstream water samples; wherein the target groundwater flows from the candidate pollution source sampling points to the downstream sampling points;

[0007] Based on the candidate pollution source water sample, chemical composition characteristics are extracted to obtain the first chemical characteristic, and based on the downstream water sample, chemical composition characteristics are extracted to obtain the second chemical characteristic;

[0008] The similarity of the target chemical feature is obtained by calculating the similarity between the first chemical feature and the second chemical feature.

[0009] Based on the similarity of the target chemical characteristics, the target pollution source sampling point is determined from the candidate pollution source sampling points.

[0010] In some embodiments, the step of calculating the similarity of the target chemical feature based on the first chemical feature and the second chemical feature includes:

[0011] The similarity measure is obtained by using the dynamic time warping module in the pre-trained similarity calculation model to calculate the similarity between the first chemical feature and the second chemical feature.

[0012] The support vector machine module in the similarity calculation model is used to calculate the matching score between the first chemical feature and the second chemical feature to obtain the feature matching score.

[0013] The target chemical feature similarity is obtained by weighted summation of the feature similarity metric and the feature matching score.

[0014] In some embodiments, the first chemical feature includes a first characteristic peak sequence, and the second chemical feature includes a second characteristic peak sequence;

[0015] The dynamic time warping module in the pre-trained similarity calculation model performs similarity measurement on the first chemical feature and the second chemical feature to obtain the feature similarity measurement, including:

[0016] The initial feature peak distance matrix is ​​obtained by calculating the distance matrix based on the first feature peak sequence and the second feature peak sequence;

[0017] Dynamic path planning is performed based on the initial feature peak distance matrix to obtain the cumulative distance matrix;

[0018] The feature similarity measure is determined based on the matrix elements of the cumulative distance matrix.

[0019] In some embodiments, the step of calculating a feature matching score by using the support vector machine module in the similarity calculation model to perform matching score calculation on the first chemical feature and the second chemical feature includes:

[0020] The target classification hyperplane is obtained by determining the hyperplane based on the first chemical feature and the second chemical feature.

[0021] The projection distance is calculated based on the first chemical feature and the target classification hyperplane to obtain the first feature projection distance, and the projection distance is calculated based on the second chemical feature and the target classification hyperplane to obtain the second feature projection distance.

[0022] The feature matching score is obtained by calculating the distance similarity based on the first feature projection distance and the second feature projection distance.

[0023] In some embodiments, the step of extracting chemical composition features from the candidate pollution source water sample to obtain a first chemical feature, and extracting chemical composition features from the downstream water sample to obtain a second chemical feature, includes:

[0024] Chemical fingerprinting was performed on the water samples from the candidate pollution sources to obtain the first chemical fingerprint spectrum;

[0025] The downstream water sample was subjected to chemical fingerprinting to obtain a second chemical fingerprint.

[0026] The first chemical fingerprint spectrum is subjected to characteristic peak sequence extraction to obtain the first characteristic peak sequence, and the first characteristic peak sequence is determined as the first chemical feature;

[0027] The characteristic peak sequence of the second chemical fingerprint is extracted to obtain the second characteristic peak sequence, and the second characteristic peak sequence is identified as the second chemical feature.

[0028] In some embodiments, the number of candidate pollution source sampling points is at least two, and each target chemical feature similarity corresponds to one candidate pollution source sampling point;

[0029] The step of determining the target pollution source sampling point from the candidate pollution source sampling points based on the target chemical feature similarity includes:

[0030] The sampling point with the highest similarity to the target chemical feature is selected from at least two candidate pollution source sampling points to obtain the maximum similarity sampling point;

[0031] If the similarity of the target chemical feature of the maximum similarity sampling point is greater than or equal to a preset similarity threshold, the maximum similarity sampling point is determined as the target pollution source sampling point.

[0032] In some embodiments, at least two of the candidate pollution source water samples include candidate pollution source water samples at at least two time points, each time point corresponding to a target chemical feature similarity;

[0033] After determining the target pollution source sampling point from the candidate pollution source sampling points based on the target chemical feature similarity, the method further includes:

[0034] Based on the similarity of the target chemical features of the sampling points of the target pollution source at at least two time points, the time point with the greatest similarity of the target chemical features is determined as the target time point;

[0035] Information on the discharge behavior of the target pollution source sampling point at the target time point is obtained, so as to perform pollutant source analysis based on the discharge behavior information.

[0036] To achieve the above objectives, a second aspect of this application provides an apparatus for analyzing the sources of pollutants in groundwater, the apparatus comprising:

[0037] The water sampling module is used to collect water samples at candidate pollution source sampling points of the target groundwater to obtain candidate pollution source water samples, and to collect water samples at downstream sampling points of the target groundwater to obtain downstream water samples; wherein the target groundwater flows from the candidate pollution source sampling points to the downstream sampling points;

[0038] The feature extraction module is used to extract chemical composition features from the candidate pollution source water sample to obtain a first chemical feature, and to extract chemical composition features from the downstream water sample to obtain a second chemical feature;

[0039] The similarity calculation module is used to calculate the similarity between the first chemical feature and the second chemical feature to obtain the similarity of the target chemical feature;

[0040] The pollution source determination module is used to determine the target pollution source sampling point from the candidate pollution source sampling points based on the similarity of the target chemical characteristics.

[0041] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0042] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0043] This application proposes a method, apparatus, equipment, and medium for analyzing the sources of pollutants in groundwater. It collects water samples from candidate pollution source sampling points in the target groundwater to obtain candidate pollution source water samples, and then collects water samples from downstream sampling points to obtain downstream water samples. This allows for the collection of spatially correlated water samples based on the direction of groundwater flow, accurately reflecting the water quality (e.g., chemical composition) at different locations, providing a sample basis for subsequent pollution source homology analysis. Chemical composition characteristics are extracted from both the candidate pollution source water samples and the downstream water samples to obtain first and second chemical characteristics. This allows for precise identification of the chemical characteristics of pollutants in the water samples, such as the unique chemical fingerprint characteristics of metal ion pollutants and organic pollutants. Similarity calculations are performed based on the first and second chemical characteristics to obtain the target chemical characteristic similarity. This enables quantitative analysis of the complex relationship between the chemical characteristics of the candidate pollution source water samples and the downstream water samples, rather than simply qualitative analysis, thus providing a more accurate resolution of the homology association of pollutants between the candidate pollution source water samples and the downstream water samples. Since the downstream sampling point is located downstream of the candidate pollution source sampling point, and the similarity of the target chemical characteristics can reflect the consistency of the chemical characteristics of the downstream pollutants with those of the pollutants in the candidate pollution source, the target pollution source sampling point that matches the downstream pollutants (e.g., the highest matching degree) can be screened from the candidate pollution sources by similarity comparison. This improves the accuracy of pollutant source analysis, reduces the risk of subjective misjudgment in complex groundwater chemical environments, and accurately locates the pollution source. Attached Figure Description

[0044] Figure 1 This is a flowchart of a method for analyzing the sources of pollutants in groundwater provided in an embodiment of this application;

[0045] Figure 2 yes Figure 1 The flowchart for step 102 in the document;

[0046] Figure 3 yes Figure 1 The flowchart for step 103 in the text;

[0047] Figure 4 yes Figure 3 The flowchart for step 301 in the document;

[0048] Figure 5 yes Figure 3 The flowchart for step 302 in the document;

[0049] Figure 6 yes Figure 1 The flowchart for step 104 in the document;

[0050] Figure 7This is a flowchart of a method for analyzing the sources of contaminants in groundwater, provided in another embodiment of this application;

[0051] Figure 8 This is a schematic diagram of the structure of the groundwater pollutant source analysis device provided in the embodiments of this application;

[0052] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0056] First, let's analyze some of the terms used in this application:

[0057] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. AI is a branch of computer science that attempts to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can be a simulation of the information processes of human consciousness and thought. AI can also be the theory, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning. This application can acquire and process relevant data based on AI technology.

[0058] Support Vector Machine (SVM) is a machine learning algorithm used for classification problems. SVM can be used to solve data classification problems in the field of pattern recognition. For example, SVM can find an optimal hyperplane in a high-dimensional space to separate data points of different classes.

[0059] Dynamic Time Warping (DTW) is an algorithm for calculating the similarity between time series. DTW can find a similarity metric between two time series by calculating the optimal matching path between them.

[0060] In groundwater pollution control, identifying the source of pollutants is crucial for developing effective remediation strategies. Pollutants in groundwater possess unique chemical compositions (such as chemical fingerprints), containing characteristic information about the various chemical components within the pollutant. However, current analyses of groundwater pollutant sources are mostly qualitative or semi-quantitative, highly subjective, and difficult to accurately determine the source of pollutants. For example, current pollutant source analysis relies on simple colorimetric methods or experience-based spectral comparisons. When faced with complex groundwater chemical environments and mixtures of multiple pollutants, it is impossible to accurately determine the source and propagation pathways of pollutants, thus impacting groundwater pollution control efforts, such as wasting resources and delaying remediation.

[0061] Based on this, embodiments of this application provide a method, apparatus, equipment, and medium for analyzing the sources of pollutants in groundwater. By accurately detecting and analyzing upstream and downstream water samples, the matching degree between the chemical fingerprints of pollutants is precisely calculated and quantitatively analyzed, thereby clearly determining the source and propagation trajectory of pollutants, improving the accuracy of pollutant source analysis, and providing a reliable basis for groundwater pollution control.

[0062] The methods, apparatus, equipment, and media for analyzing the sources of pollutants in groundwater provided in this application are specifically illustrated through the following embodiments. First, the method for analyzing the sources of pollutants in groundwater in this application is described.

[0063] The groundwater pollutant source analysis method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms; the software can be an application that implements the groundwater pollutant source analysis method, but is not limited to the above forms.

[0064] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0065] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0066] Figure 1 This is an optional flowchart of a method for analyzing the sources of contaminants in groundwater provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps 101 to 104.

[0067] Step 101: Collect water samples at the candidate pollution source sampling point of the target groundwater to obtain candidate pollution source water samples, and collect water samples at the downstream sampling point of the target groundwater to obtain downstream water samples; wherein, the target groundwater flows from the candidate pollution source sampling point to the downstream sampling point;

[0068] Step 102: Extract chemical composition features from candidate pollution source water samples to obtain the first chemical feature, and extract chemical composition features from downstream water samples to obtain the second chemical feature;

[0069] Step 103: Calculate the similarity based on the first chemical feature and the second chemical feature to obtain the similarity of the target chemical feature;

[0070] Step 104: Determine the target pollution source sampling point from the candidate pollution source sampling points based on the similarity of the target chemical characteristics.

[0071] The beneficial effects of this application's embodiments include, but are not limited to: obtaining candidate pollution source water samples by collecting water samples at candidate pollution source sampling points of the target groundwater, and obtaining downstream water samples by collecting water samples at downstream sampling points of the target groundwater. This allows for the collection of spatially correlated water samples based on the direction of groundwater flow. These water samples can accurately reflect the water quality conditions (such as chemical composition) at different locations of the groundwater, providing a sample basis for subsequent pollution homology analysis. Chemical composition characteristics are extracted from the candidate pollution source water samples and the downstream water samples respectively to obtain first and second chemical characteristics. This allows for the accurate identification of the chemical characteristics of pollutants in the water samples, such as identifying the unique chemical fingerprint characteristics of pollutants like metal ion pollutants and organic pollutants. Similarity calculations are performed based on the first and second chemical characteristics to obtain the target chemical characteristic similarity. This enables quantitative analysis of the complex relationship between the chemical characteristics of the candidate pollution source water samples and the downstream water samples, rather than simply performing qualitative analysis, thereby more accurately resolving the homology association of pollutants between the candidate pollution source water samples and the downstream water samples. Since the downstream sampling point is located downstream of the candidate pollution source sampling point, and the similarity of the target chemical characteristics can reflect the consistency of the chemical characteristics of the downstream pollutants with those of the pollutants in the candidate pollution source, the target pollution source sampling point that matches the downstream pollutants (e.g., the highest matching degree) can be screened from the candidate pollution sources by similarity comparison. This improves the accuracy of pollutant source analysis, reduces the risk of subjective misjudgment in complex groundwater chemical environments, and accurately locates the pollution source.

[0072] In step 101 of some embodiments, the target groundwater refers to the groundwater for which pollutant tracing is required. Candidate pollution source sampling points are sampling locations set within the target groundwater. These sampling points are located near candidate pollution sources, for example, the distance between the sampling point and the candidate pollution source is less than a preset distance threshold. Specifically, candidate pollution source sampling points can be set near pollution sources in the upstream portion of the target groundwater. For example, the candidate pollution sources corresponding to the sampling points can include any one or more of the following: industrial pollution sources (such as wastewater discharged from factories), agricultural pollution sources (such as fertilizers, pesticides, livestock and aquaculture waste, etc.). Candidate pollution sources can also include other types of pollution sources, and are not limited to these. In some embodiments, the number of candidate pollution source sampling points can be one or more, i.e., at least one.

[0073] In some embodiments, it should be noted that the downstream sampling point is a sampling location set in the downstream portion of the target groundwater, and the flow direction of the target groundwater is from the candidate pollution source sampling point to the downstream sampling point. The candidate pollution source water sample is a groundwater sample collected at the candidate pollution source sampling point. The downstream water sample is a groundwater sample collected at the downstream sampling point.

[0074] In some embodiments, it should be noted that agricultural pollution sources can cause agricultural non-point source pollution. Agricultural non-point source pollution refers to the pollution of the ecological environment caused by the unreasonable use of chemical inputs such as fertilizers, pesticides, and mulch films during agricultural production, as well as the untimely or improper handling of livestock and aquaculture waste, crop straw, etc., resulting in nutrients such as nitrogen, phosphorus, and organic matter entering groundwater under the combined influence of rainfall and topography.

[0075] In some embodiments, sampling points can be rationally set up in the upstream and downstream areas of the groundwater based on the hydrogeological characteristics of the target groundwater and the distribution of potential pollution sources, and the sampling frequency can be set. For example, a candidate pollution source sampling point A1 can be set up in the groundwater near an industrial emission source, and a downstream sampling point B1 can be set up downstream of sampling point A1. As another example, a candidate pollution source sampling point A2 can be set up in the groundwater in areas with concentrated agricultural non-point source pollution (such as agricultural planting areas, near fertilizer and pesticide warehouses, and near irrigation water sources), and a downstream sampling point B2 can be set up downstream of sampling point A2.

[0076] In some embodiments, specifically, water sampling equipment such as bottle-type deep water samplers and flow ratio samplers can be used to collect water samples, thereby ensuring that the collected water samples are not contaminated by the outside world and can truly reflect the water quality of the groundwater.

[0077] In step 102 of some embodiments, the first chemical feature is the chemical composition feature of the candidate pollution source water sample. The second chemical feature is the chemical composition feature of the downstream water sample. In some embodiments, the first chemical feature (or the second chemical feature) may be a feature of the chemical fingerprint spectrum of the candidate pollution source water sample (or downstream water sample), such as the characteristic parameters of the characteristic peaks of the chemical fingerprint spectrum (e.g., peak position, peak intensity, peak area, etc.). As for the meaning and function of the chemical fingerprint spectrum, please refer to the specific description of step 201 below, which will not be repeated here.

[0078] In step 103 of some embodiments, the target chemical feature similarity is the similarity between the first chemical feature and the second chemical feature, that is, the similarity between the chemical features of the candidate pollution source water sample and the chemical features of the downstream water sample. In some embodiments, specifically, the target chemical feature similarity can be the cosine similarity between the first chemical feature and the second chemical feature, or it can be other types of similarity, not limited thereto. It should be noted that the higher the target chemical feature similarity, the higher the degree of matching of the chemical features of pollutants in the upstream and downstream water samples (that is, the candidate pollution source water sample and the downstream water sample), that is, the pollutants in the downstream water sample are more likely to originate from the upstream area corresponding to the candidate pollution source water sample.

[0079] In step 104 of some embodiments, the target pollution source sampling point is a sampling point selected from the candidate pollution source sampling points. For example, a candidate pollution source sampling point whose target chemical feature similarity exceeds a preset similarity threshold can be determined as the target pollution source sampling point. As another example, when there are multiple candidate pollution source sampling points, the candidate pollution source sampling point with the highest target chemical feature similarity can be determined as the target pollution source sampling point.

[0080] Please see Figure 2 In some embodiments, step 102 may include, but is not limited to, steps 201 to 204:

[0081] Step 201: Perform chemical fingerprinting on the water sample from the candidate pollution source to obtain the first chemical fingerprint.

[0082] Step 202: Perform chemical fingerprinting on the downstream water sample to obtain a second chemical fingerprint.

[0083] Step 203: Extract the characteristic peak sequence from the first chemical fingerprint spectrum to obtain the first characteristic peak sequence, and determine the first characteristic peak sequence as the first chemical feature;

[0084] Step 204: Extract the characteristic peak sequence from the second chemical fingerprint spectrum to obtain the second characteristic peak sequence, and identify the second characteristic peak sequence as the second chemical feature.

[0085] The advantage of this embodiment lies in the fact that by performing chemical fingerprinting on candidate pollution source water samples and downstream water samples respectively, a first chemical fingerprint spectrum and a second chemical fingerprint spectrum are obtained. This allows for the precise capture of the unique chemical identification characteristics of pollutants (such as metal ions, organic pollutants, etc.). Then, characteristic peak sequences are extracted from the first and second chemical fingerprint spectra to obtain the first and second characteristic peak sequences, which are identified as chemical features. This allows for the extraction of time-series characteristic peak sequences from the spectral information, providing chemically characteristic data for subsequent similarity calculations, thereby improving the accuracy of pollutant source analysis and enabling precise location of pollution sources.

[0086] In step 201 of some embodiments, the first chemical fingerprint spectrum is a chemical fingerprint spectrum obtained from the detection of a candidate pollution source water sample. It should be noted that a chemical fingerprint refers to a stable and quantifiable combination of chemical characteristics in a substance (such as a pollutant). A chemical fingerprint spectrum is a spectrum used to characterize the chemical characteristics of a substance, such as a chromatogram or a spectrum.

[0087] In step 202 of some embodiments, the second chemical fingerprint is a chemical fingerprint obtained from the detection of a downstream water sample.

[0088] In some embodiments, specifically, chemical fingerprinting of water samples (including candidate pollution source water samples and downstream water samples) can be performed using chemical composition analysis instruments. For example, in industrial pollution analysis scenarios, inductively coupled plasma mass spectrometry (ICP-MS) can be used to analyze water samples and obtain chemical fingerprint spectra, allowing the source of heavy metal pollutants (such as lead, mercury, cadmium, etc.) in the water samples to be determined based on the characteristics of the chemical fingerprint spectra. Organic pollutants in water samples, such as benzene, toluene, xylene, etc., can also be detected using gas chromatography-mass spectrometry (GC-MS). As another example, in agricultural pollution analysis scenarios, inorganic salt pollutants such as nitrates and phosphates in water samples can be detected using ion chromatography. Pesticide residues in water samples can also be analyzed using high-performance liquid chromatography (HPLC).

[0089] It should be noted that inductively coupled plasma mass spectrometry (ICP-MS) is a mass spectrometer used to determine ultra-trace elements and isotope ratios. ICP-MS can be used to determine trace, ultra-trace, and trace metal elements in groundwater, and mass spectra can be obtained through ICP-MS.

[0090] In some embodiments, the chemical fingerprint spectrum acquired by gas chromatography-mass spectrometry (GC-MS) may include any one or more of gas chromatograms, mass spectrometers, and total ion chromatograms (TIC). In the gas chromatogram, the x-axis represents time and the y-axis represents signal intensity. In the mass spectrometer, the x-axis represents the mass-to-charge ratio and the y-axis represents ion intensity. In the total ion chromatogram, the x-axis represents time and the y-axis represents ion intensities.

[0091] In some embodiments, the raw data (i.e., chemical fingerprint spectrum) acquired by the instrument can be preprocessed to remove instrument noise and interference signals, ensuring the accuracy of the chemical fingerprint spectrum so as to accurately extract the characteristic parameters of the chemical fingerprint spectrum.

[0092] In step 203 of some embodiments, the first characteristic peak sequence is the characteristic peak sequence of the first chemical fingerprint spectrum. In some embodiments, it should be noted that characteristic peaks are absorption peaks used to identify the presence of chemical bonds or groups. Absorption peaks are characteristic spectral parameters formed by the selective absorption of light of a specific wavelength by a substance.

[0093] In step 204 of some embodiments, the second characteristic peak sequence is the characteristic peak sequence of the second chemical fingerprint spectrum. Specifically, the first characteristic peak sequence (or the second characteristic peak sequence) may have parameters such as peak position, peak intensity, and peak area.

[0094] Please see Figure 3 In some embodiments, step 103 may include, but is not limited to, steps 301 to 303:

[0095] Step 301: Through the dynamic time warping module in the pre-trained similarity calculation model, the similarity measurement of the first chemical feature and the second chemical feature is calculated to obtain the feature similarity measurement.

[0096] Step 302: The support vector machine module in the similarity calculation model is used to calculate the matching score of the first chemical feature and the second chemical feature to obtain the feature matching score.

[0097] Step 303: Perform a weighted summation based on the feature similarity measure and the feature matching score to obtain the target chemical feature similarity.

[0098] The advantage of this embodiment lies in the fact that the dynamic time warping module in the pre-trained similarity calculation model performs similarity measurement calculations on the first and second chemical features, which can accurately capture the local similarity between candidate pollution sources and downstream water samples in chemical features (such as pollutant type, concentration change trend, etc.). The support vector machine module in the similarity calculation model calculates matching scores for the first and second chemical features, which can quantify the nonlinear correlation between the chemical features of candidate pollution sources and downstream water samples. Then, the target chemical feature similarity is obtained by weighted summation based on the feature similarity measurement and feature matching scores. This combines the local temporal alignment capability of dynamic time warping with the global pattern recognition advantage of support vector machines, improving the accuracy of target chemical feature similarity and thus improving the accuracy of pollutant source analysis. It is suitable for complex groundwater chemical environments (such as multi-source mixed pollution).

[0099] In step 301 of some embodiments, the dynamic time warping module is a module based on the dynamic time warping (DTW) algorithm. The feature similarity measure is a similarity measure calculated by the dynamic time warping algorithm on the first chemical feature and the second chemical feature.

[0100] In some embodiments, it should be noted that due to the dynamic changes in groundwater flow velocity and pollutant migration processes, the chemical fingerprint spectra of upstream and downstream water samples (i.e., candidate pollution source water samples and downstream water samples) collected at different times may exhibit scaling or shifts along the time axis. The DTW algorithm can obtain a similarity measure by calculating the optimal matching path between the first and second chemical features.

[0101] In step 302 of some embodiments, the support vector machine module is a module based on the support vector machine (SVM) algorithm. The feature matching score is the matching score calculated by the support vector machine algorithm on the first chemical feature and the second chemical feature. In some embodiments, it should be noted that the SVM algorithm can effectively handle small sample, nonlinear and high-dimensional data problems, and compared with traditional multivariate statistical analysis methods, it can more accurately capture the complex relationship between chemical fingerprint features and pollutant sources.

[0102] In some embodiments, chemical fingerprint data of upstream and downstream water samples from known sources can be used as training data to adjust the parameters of the support vector machine module, so as to use the trained support vector machine module to capture the nonlinear mapping relationship between chemical fingerprints and pollutant sources.

[0103] In step 303 of some embodiments, the target chemical feature similarity is a weighted sum of the feature similarity measure and the feature matching score. By combining the calculation results of the SVM algorithm and the DTW algorithm through weighted summation, the dynamic changes of chemical features over time can be taken into account while considering the similarity of chemical features, thereby more accurately calculating the similarity of the chemical features of pollutants in upstream and downstream water samples. It should be noted that the weights of the feature similarity measure or the feature matching score can be set and adjusted according to needs, and the embodiments of this application do not limit this.

[0104] In some embodiments, before performing similarity measurement using a pre-trained similarity calculation model, the similarity calculation model can be trained and tested multiple times using cross-validation. For example, the training set data can be divided multiple times for model training and validation, and the accuracy, recall, and other metrics of the similarity calculation model on different validation sets can be calculated to determine the stability and reliability of the similarity calculation model. Furthermore, based on a preset confidence interval, the fluctuation range of the similarity calculated by the similarity calculation model at a certain confidence level can be statistically analyzed to analyze the credibility of the model's output results.

[0105] Please see Figure 4 In some embodiments, the first chemical feature includes a first characteristic peak sequence, and the second chemical feature includes a second characteristic peak sequence;

[0106] Step 301 may include, but is not limited to, steps 401 through 403:

[0107] Step 401: Calculate the distance matrix based on the first and second characteristic peak sequences to obtain the initial characteristic peak distance matrix;

[0108] Step 402: Perform dynamic path planning based on the initial feature peak distance matrix to obtain the cumulative distance matrix;

[0109] Step 403: Determine the feature similarity measure based on the matrix elements of the cumulative distance matrix.

[0110] The advantage of this embodiment lies in the fact that an initial characteristic peak distance matrix is ​​calculated based on the first and second characteristic peak sequences, which allows for precise quantification of local differences between characteristic peaks of different time series (such as pollutant concentration peaks). A cumulative distance matrix is ​​obtained through dynamic path planning based on the initial characteristic peak distance matrix, enabling the capture of local similarities between candidate pollution sources and downstream water samples in key characteristic peak positions, intensity variation trends, and other features using the Dynamic Time Warping (DTW) algorithm. The feature similarity metric is determined based on the matrix elements of the cumulative distance matrix, allowing for the extraction of the minimum cumulative distance value to obtain a quantifiable similarity index. This improves the accuracy of identifying the chemical fingerprint of pollutants in complex groundwater chemical environments, thereby enhancing the accuracy of pollutant source analysis.

[0111] In step 401 of some embodiments, the initial feature peak distance matrix is ​​a matrix composed of the distances (e.g., Euclidean distances) between each element in the first feature peak sequence and each element in the second feature peak sequence.

[0112] In step 402 of some embodiments, the cumulative distance matrix is ​​a matrix obtained by dynamic path planning based on the initial feature peak distance matrix. For example, suppose the initial feature peak distance matrix is ​​matrix D, and the cumulative distance matrix is ​​matrix D'. Then, the element D'[i][j] in the i-th row and j-th column of the cumulative distance matrix is ​​defined as:

[0113] D'[i][j]=D[i][j]+min(D[i-1][j], D[i][j-1], D[i-1][j-1]),

[0114] Where D'[i][j] represents the element in the i-th row and j-th column of the cumulative distance matrix; D[i][j] represents the element in the i-th row and j-th column of the initial feature peak distance matrix; D[i-1][j] represents the element in the (i-1)-th row and j-th column of the initial feature peak distance matrix; D[i][j-1] represents the element in the i-th row and (j-1)-th column of the initial feature peak distance matrix; and min() represents the minimum value function, used to find the minimum value.

[0115] In step 403 of some embodiments, the feature similarity measure can be the matrix element in the last row and last column of the cumulative distance matrix. For example, assuming the cumulative distance matrix D' has m rows and n columns, then the matrix element D'[m][n] is determined as the feature similarity measure.

[0116] Please see Figure 5 In some embodiments, step 302 may include, but is not limited to, steps 501 to 503:

[0117] Step 501: Determine the hyperplane based on the first chemical feature and the second chemical feature to obtain the target classification hyperplane;

[0118] Step 502: Calculate the projection distance based on the first chemical feature and the target classification hyperplane to obtain the projection distance of the first feature, and calculate the projection distance based on the second chemical feature and the target classification hyperplane to obtain the projection distance of the second feature;

[0119] Step 503: Calculate the distance similarity based on the first feature projection distance and the second feature projection distance to obtain the feature matching score.

[0120] The advantage of this embodiment lies in the following: A target classification hyperplane is determined based on the first and second chemical features. This allows for the determination of the optimal classification interface in a high-dimensional feature space using Support Vector Machines (SVM), mapping complex chemical features to a quantifiable decision space, thereby accurately distinguishing the feature distribution patterns of different pollution sources. The first feature projection distance is calculated based on the first chemical feature and the target classification hyperplane, and the second feature projection distance is calculated based on the second chemical feature and the target classification hyperplane. This allows for the quantification of the relative positional relationship between candidate pollution sources and downstream water samples in the high-dimensional feature space through projection distance. Finally, a feature matching score is obtained by calculating distance similarity based on the first and second feature projection distances. This transforms projection distance differences into feature matching scores through a distance similarity function, improving the accuracy of chemical feature similarity analysis of water samples and consequently enhancing the accuracy of pollutant source analysis.

[0121] In step 501 of some embodiments, the target classification hyperplane is the optimal classification hyperplane used to distinguish between the first chemical feature and the second chemical feature. It should be noted that the Support Vector Machine (SVM) algorithm can find an optimal classification hyperplane in a high-dimensional space to separate data of different categories.

[0122] In step 502 of some embodiments, the first feature projection distance is the projection distance of the first chemical feature on the target classification hyperplane. The second feature projection distance is the projection distance of the second chemical feature on the target classification hyperplane.

[0123] In step 503 of some embodiments, the feature matching score is the distance similarity between the first feature projection distance and the second feature projection distance. For example, the difference between the first feature projection distance and the second feature projection distance can be used as the numerator, and the sum of the first feature projection distance and the second feature projection distance can be used as the denominator. The result is obtained by dividing by the sum of the first feature projection distance and the second feature projection distance to get the distance difference percentage. Then, the distance difference percentage is subtracted from 1 to obtain the feature matching score.

[0124] Please see Figure 6In some embodiments, the number of candidate pollution source sampling points is at least two, and each target chemical feature similarity corresponds to one candidate pollution source sampling point;

[0125] Step 104 may include, but is not limited to, steps 601 to 602:

[0126] Step 601: Select the candidate pollution source sampling point with the highest similarity to the target chemical characteristics from at least two candidate pollution source sampling points to obtain the sampling point with the highest similarity.

[0127] Step 602: If the target chemical feature similarity of the maximum similarity sampling point is greater than or equal to the preset similarity threshold, the maximum similarity sampling point is determined as the target pollution source sampling point.

[0128] The advantage of this embodiment lies in that by setting at least two candidate pollution source sampling points and calculating the similarity of the target chemical features corresponding to each point, multiple potential pollution source areas can be covered, avoiding accidental misjudgments caused by a single sampling point and providing a data foundation for multi-source pollution identification. The sampling point with the highest similarity of target chemical features from multiple candidate pollution source sampling points is selected as the maximum similarity sampling point, enabling objective identification of candidate areas with the highest pollution homology based on quantitative data. A preset similarity threshold is used for determination; only when the similarity of the maximum similarity sampling point is greater than or equal to this threshold is it identified as a target pollution source sampling point. This eliminates interference from low-matching factors, improves the reliability of pollution source determination results, and thus enhances the accuracy and reliability of pollutant source analysis.

[0129] In step 601 of some embodiments, the maximum similarity sampling point is the candidate pollution source sampling point with the highest similarity to the target chemical features.

[0130] In step 602 of some embodiments, the specific value of the similarity threshold can be set or adjusted according to requirements, and this application embodiment does not limit this.

[0131] Please see Figure 7 In some embodiments, at least two candidate pollution source water samples include candidate pollution source water samples at at least two time points, each time point corresponding to a target chemical feature similarity;

[0132] Following step 104, the method for analyzing the sources of contaminants in groundwater may also include, but is not limited to, steps 701 to 702:

[0133] Step 701: Based on the similarity of target chemical features of the target pollution source sampling points at at least two time points, determine the time point with the highest similarity of target chemical features as the target time point;

[0134] Step 702: Obtain the pollution discharge behavior information of the target pollution source sampling point at the target time point, so as to conduct pollutant source analysis based on the pollution discharge behavior information.

[0135] The advantage of this embodiment lies in capturing the temporal dynamics of pollution discharge by setting candidate pollution source water samples at at least two time points and calculating the similarity of target chemical features corresponding to each time point. Based on the similarity of target chemical features of the target pollution source sampling points at different time points, the time point corresponding to the maximum value is selected as the target time point. This allows for precise location of critical times (such as peak discharge periods) of pollution events through similarity peaks. Then, information on the discharge behavior of the target pollution source sampling points at the target time point is obtained, revealing the correlation between discharge behavior at a specific time point and downstream water quality changes, thereby improving the spatiotemporal accuracy and reliability of pollutant source analysis.

[0136] In some embodiments, water samples can be collected from target pollution source sampling points at at least two time points to obtain candidate pollution source water samples at at least two time points. For example, in agricultural pollution scenarios, water samples can be collected during key periods such as before and after crop fertilization and pesticide spraying. In some embodiments, such as in industrial pollution scenarios, water samples can be collected continuously for multiple sampling cycles according to a preset sampling period (e.g., weekly), such as collecting water samples once a week for three consecutive months, thereby obtaining water sample data at different times.

[0137] In step 701 of some embodiments, the target time point is the time point where the similarity of target chemical characteristics between the target pollution source sampling point and the downstream sampling point is the greatest.

[0138] In step 702 of some embodiments, the pollution discharge behavior information at the target time point is used to characterize the pollution discharge behavior that occurs at the target time point. For example, in an industrial pollution scenario, the pollution discharge behavior may include wastewater discharge from a chemical enterprise. As another example, in an agricultural pollution scenario, the pollution discharge behavior may include fertilizing crops, spraying pesticides, etc.

[0139] In some embodiments, for example, assuming that the sewage discharge behavior information indicates that crop fertilization occurred at a target time point (i.e., the time point with the highest similarity to the target chemical characteristics), it can be determined that crop fertilization caused pollutants in downstream water samples, thereby enabling more precise analysis of pollutant sources in time.

[0140] Please see Figure 8 This application also provides a device for analyzing the sources of pollutants in groundwater, which can implement the above-mentioned method for analyzing the sources of pollutants in groundwater. The device includes:

[0141] The water sampling module 801 is used to collect water samples at the candidate pollution source sampling point of the target groundwater to obtain candidate pollution source water samples, and to collect water samples at the downstream sampling point of the target groundwater to obtain downstream water samples; wherein, the target groundwater flows from the candidate pollution source sampling point to the downstream sampling point;

[0142] The feature extraction module 802 is used to extract chemical composition features from candidate pollution source water samples to obtain a first chemical feature, and to extract chemical composition features from downstream water samples to obtain a second chemical feature.

[0143] The similarity calculation module 803 is used to calculate the similarity based on the first chemical feature and the second chemical feature to obtain the similarity of the target chemical feature;

[0144] The pollution source determination module 804 is used to determine the target pollution source sampling point from the candidate pollution source sampling points based on the similarity of the target chemical characteristics.

[0145] In one embodiment, the pollutant source analysis device for groundwater further includes a time-point pollution analysis module, used to: determine the time point with the highest similarity of target chemical features at at least two time points based on the similarity of target chemical features of the target pollution source sampling point; and acquire information on the discharge behavior of the target pollution source sampling point at the target time point, so as to perform pollutant source analysis based on the discharge behavior information.

[0146] The specific implementation of this groundwater pollutant source analysis device is basically the same as the specific embodiment of the groundwater pollutant source analysis method described above, and will not be repeated here.

[0147] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for analyzing the sources of pollutants in groundwater. This electronic device can include any smart terminal such as a tablet computer or an in-vehicle computer.

[0148] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0149] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0150] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to execute the groundwater pollutant source analysis method of the embodiments of this application.

[0151] The input / output interface 903 is used to implement information input and output;

[0152] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0153] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0154] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0155] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for analyzing the sources of pollutants in groundwater.

[0156] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0157] It should be noted that the software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0158] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0159] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0162] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0163] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0164] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0165] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0167] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0168] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for analyzing the sources of pollutants in groundwater, characterized in that, The method includes: Water samples are collected at candidate pollution source sampling points of the target groundwater to obtain candidate pollution source water samples, and water samples are also collected at downstream sampling points of the target groundwater to obtain downstream water samples; wherein the target groundwater flows from the candidate pollution source sampling points to the downstream sampling points; Chemical composition features are extracted from the candidate pollution source water sample to obtain a first chemical feature, and chemical composition features are extracted from the downstream water sample to obtain a second chemical feature; the first chemical feature includes a first characteristic peak sequence, and the second chemical feature includes a second characteristic peak sequence. The initial feature peak distance matrix is ​​obtained by calculating the distance matrix based on the first feature peak sequence and the second feature peak sequence; Dynamic path planning is performed based on the initial feature peak distance matrix to obtain the cumulative distance matrix; the element D'[i][j] in the i-th row and j-th column of the cumulative distance matrix is ​​defined as follows: D'[i][j]=D[i][j]+min(D[i-1][j], D[i][j-1], D[i-1][j-1]), Where D'[i][j] represents the element in the i-th row and j-th column of the cumulative distance matrix; D[i][j] represents the element in the i-th row and j-th column of the initial feature peak distance matrix; D[i-1][j] represents the element in the (i-1)-th row and j-th column of the initial feature peak distance matrix; D[i][j-1] represents the element in the i-th row and (j-1)-th column of the initial feature peak distance matrix; and min() represents the minimum value function, used to find the minimum value. The matrix elements in the last row and last column of the cumulative distance matrix are determined as feature similarity measures; The support vector machine module in the pre-trained similarity calculation model is used to calculate the matching score between the first chemical feature and the second chemical feature to obtain the feature matching score. The target chemical feature similarity is obtained by weighted summation of the feature similarity measure and the feature matching score. Based on the similarity of the target chemical characteristics, the target pollution source sampling point is determined from the candidate pollution source sampling points.

2. The method according to claim 1, characterized in that, The support vector machine module in the pre-trained similarity calculation model calculates matching scores for the first chemical feature and the second chemical feature to obtain feature matching scores, including: The target classification hyperplane is obtained by determining the hyperplane based on the first chemical feature and the second chemical feature. The projection distance is calculated based on the first chemical feature and the target classification hyperplane to obtain the first feature projection distance, and the projection distance is calculated based on the second chemical feature and the target classification hyperplane to obtain the second feature projection distance. The feature matching score is obtained by calculating the distance similarity based on the first feature projection distance and the second feature projection distance.

3. The method according to any one of claims 1 to 2, characterized in that, The step of extracting chemical composition features from the candidate pollution source water sample to obtain a first chemical feature, and extracting chemical composition features from the downstream water sample to obtain a second chemical feature, includes: Chemical fingerprinting was performed on the water samples from the candidate pollution sources to obtain the first chemical fingerprint spectrum; The downstream water sample was subjected to chemical fingerprinting to obtain a second chemical fingerprint. The first chemical fingerprint spectrum is subjected to characteristic peak sequence extraction to obtain the first characteristic peak sequence, and the first characteristic peak sequence is determined as the first chemical feature; The characteristic peak sequence of the second chemical fingerprint is extracted to obtain the second characteristic peak sequence, and the second characteristic peak sequence is identified as the second chemical feature.

4. The method according to any one of claims 1 to 2, characterized in that, The number of candidate pollution source sampling points is at least two, and each target chemical feature similarity corresponds to one candidate pollution source sampling point. The step of determining the target pollution source sampling point from the candidate pollution source sampling points based on the target chemical feature similarity includes: The sampling point with the highest similarity to the target chemical feature is selected from at least two candidate pollution source sampling points to obtain the maximum similarity sampling point; If the similarity of the target chemical feature of the maximum similarity sampling point is greater than or equal to a preset similarity threshold, the maximum similarity sampling point is determined as the target pollution source sampling point.

5. The method according to any one of claims 1 to 2, characterized in that, The at least two candidate pollution source water samples include the candidate pollution source water samples at at least two time points, and each time point corresponds to a target chemical feature similarity; After determining the target pollution source sampling point from the candidate pollution source sampling points based on the target chemical feature similarity, the method further includes: Based on the similarity of the target chemical features of the sampling points of the target pollution source at at least two time points, the time point with the greatest similarity of the target chemical features is determined as the target time point; Information on the discharge behavior of the target pollution source sampling point at the target time point is obtained, so as to perform pollutant source analysis based on the discharge behavior information.

6. A device for analyzing the sources of pollutants in groundwater, characterized in that, The device includes: The water sampling module is used to collect water samples at candidate pollution source sampling points of the target groundwater to obtain candidate pollution source water samples, and to collect water samples at downstream sampling points of the target groundwater to obtain downstream water samples; wherein the target groundwater flows from the candidate pollution source sampling points to the downstream sampling points; The feature extraction module is used to extract chemical component features from the candidate pollution source water sample to obtain a first chemical feature, and to extract chemical component features from the downstream water sample to obtain a second chemical feature; the first chemical feature includes a first feature peak sequence, and the second chemical feature includes a second feature peak sequence. The similarity calculation module is used to calculate the distance matrix based on the first feature peak sequence and the second feature peak sequence to obtain an initial feature peak distance matrix; and to perform dynamic path planning based on the initial feature peak distance matrix to obtain a cumulative distance matrix; the element D'[i][j] in the i-th row and j-th column of the cumulative distance matrix is ​​defined as: D'[i][j]=D[i][j]+min(D[i-1][j], D[i][j-1], D[i-1][j-1]), Where D'[i][j] represents the element in the i-th row and j-th column of the cumulative distance matrix; D[i][j] represents the element in the i-th row and j-th column of the initial feature peak distance matrix; D[i-1][j] represents the element in the (i-1)-th row and j-th column of the initial feature peak distance matrix; D[i][j-1] represents the element in the i-th row and (j-1)-th column of the initial feature peak distance matrix; and min() represents the minimum value function, used to find the minimum value. The matrix elements in the last row and last column of the cumulative distance matrix are determined as feature similarity measures; the support vector machine module in the pre-trained similarity calculation model is used to calculate the matching score of the first chemical feature and the second chemical feature to obtain the feature matching score; the feature similarity measure and the feature matching score are weighted and summed to obtain the target chemical feature similarity. The pollution source determination module is used to determine the target pollution source sampling point from the candidate pollution source sampling points based on the similarity of the target chemical characteristics.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method for analyzing the sources of pollutants in groundwater as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for analyzing the sources of pollutants in groundwater as described in any one of claims 1 to 5.