Assessment Method and System for the Influence of Soil on Groundwater Based on Big Data Analysis

By regional division and sample balance processing of soil sample data, adjusting sample number and building a prediction model, the evaluation unfair problem caused by uneven regional sample number in the existing technology is solved, and a more accurate and comprehensive assessment of the impact of soil on groundwater is achieved.

CN119886959BActive Publication Date: 2025-07-29BEIJING GEOLOGICAL PROSPECTING WATER ENVIRONMENT ENG DESIGN & RES INST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510066740.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-07-29
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In the evaluation of the impact of soil on groundwater, the deviation in the training data set results in the model's sample size in some areas too much or too little, resulting in unfair prediction results and ignoring the actual situation in other areas.

Method used

By dividing the area to be analyzed into multiple sub-regions, pre-processing of sample data and sample balance processing are performed, the number of samples is adjusted, and the number of samples in each sub-region is consistent. The local reachable density and LOF value are used to determine the deleted sample points, and the sample points of the low-density area are added through the clustering algorithm to build a comprehensive data set to establish a prediction model.

Benefits of technology

It improves the representativeness and fairness of the data, reduces the bias of the model, and ensures the accuracy and comprehensiveness of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886959B_ABST
    Figure CN119886959B_ABST
Patent Text Reader

Abstract

The present invention discloses an evaluation method and system for the impact of soil on groundwater based on big data analysis, which relates to the technical field of hydrological resources. It includes: S1: Divide the area to be analyzed, obtain multiple sub-regions and sample data within the sub-regions, and perform preprocessing and sample balancing processing on the sample data; S2: Combine the preprocessed and sample-balanced sample data with external data, and construct a prediction model according to the combined data set; S3: Use the sample sampling point data obtained in real time as the input of the prediction model, and output to obtain real-time evaluation data. The present invention takes the sample quantity as the processing standard, increases or deletes the sample quantities of different regions, so that the sample quantities of different regions are kept consistent, thereby reducing the model bias caused by too many or too few samples in some areas, and further improving the representativeness and fairness of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrological resources, and specifically to an evaluation method and system for the impact of soil on groundwater based on big data analysis. Background Art

[0002] With the acceleration of urbanization and industrialization processes, the problem of soil pollution has become increasingly serious, posing a potential threat to groundwater. Traditional soil and groundwater environmental monitoring methods often rely on fixed-point sampling and laboratory testing. This method is not only time-consuming and laborious but also difficult to capture the true situation of pollutant migration and diffusion.

[0003] Big data analysis can reveal trends and patterns hidden behind a large number of observational records by processing massive and diverse datasets. For the study of soil-groundwater interactions, it can obtain a more comprehensive understanding at different time scales (such as seasonal variations) and spatial resolutions (from local sites to the regional level). For example, with the help of big data technology, soil properties most likely to cause groundwater pollution can be identified, and targeted prevention and control measures can be formulated accordingly.

[0004] A Chinese invention patent with the publication number CN117494042A discloses a method for fusing soil water in data-deficient areas based on climate similarity transfer learning. This method collects data on soil water (including ground station data, remote sensing, and reanalysis data) and their key influencing factors in the source domain and the target domain. After preprocessing and normalization, a data fusion framework is established. The CNN-LSTM model can effectively extract the spatial and temporal features of the data. Based on climate similarity, the source domain is screened and the target domain is refined and classified. From the perspective of climate similarity, different fusion sub-models are selected for different climate types, effectively solving the problem of poor transfer effect caused by the large difference in data distribution between the source domain and the target domain in transfer learning. This method includes steps such as source domain screening, model pre-training, target domain classification, sub-model transfer, and fusion model integration. By evaluating the fusion results on the test set, the effectiveness of the method is verified. This method uses deep learning and transfer learning technologies to alleviate the data shortage in data-deficient areas and the negative transfer problem in transfer learning, and improve the accuracy of soil water data in data-deficient areas.

[0005] Most of the above-mentioned and similar evaluation methods evaluate the impact of soil on groundwater through machine learning algorithms. However, during the prediction process, they often lead to unfair results due to biases in the training dataset. For example, when the number of soil samples in some regions far exceeds that in other regions, during the prediction process, it will tend to more accurately reflect the situation in that region while ignoring the actual conditions in other regions. Although this bias can be reduced by introducing regularization terms or adjusting weights, it will pose an obstacle to practical applications as the performance decreases during the process of improving the algorithm transparency. Summary of the Invention

[0006] The object of the present invention is to provide an evaluation method and system for the impact of soil on groundwater based on big data analysis to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solutions: An evaluation method for the impact of soil on groundwater based on big data analysis, including:

[0008] S1: Adjust the sample quantity ratio: Divide the area to be analyzed to obtain multiple sub-areas and sample data within the sub-areas, and perform preprocessing and sample balancing processing on the sample data;

[0009] S2: Establish a prediction model: Combine the preprocessed and sample-balanced sample data with external data, and construct a prediction model based on the combined data set;

[0010] S3: Obtain real-time evaluation data: Use the sample sampling point data obtained in real time as the input of the prediction model, and output the obtained real-time evaluation data;

[0011] Performing sample balancing processing on the sample data includes:

[0012] M1: Determine the standard quantity of sample data: Obtain the average sample quantity according to the total amount of sample data corresponding to all sub-areas, and determine the standard quantity of sample data;

[0013] M2: Classify the sub-areas: Divide the sub-areas into sub-areas with a quantity exceeding the standard and sub-areas with a quantity not exceeding the standard according to the standard quantity, specifically:

[0014] When the total amount of sample data corresponding to the sub-area is greater than the standard quantity, the sub-area is a sub-area with a quantity exceeding the standard, otherwise, the sub-area is a sub-area with a quantity not exceeding the standard;

[0015] M3: Adjust the sample quantity: Increase or decrease the total amount of sample data corresponding to the sub-area according to the classification category corresponding to the sub-area.

[0016] Furthermore, increasing or decreasing the total amount of sample data corresponding to the sub-area includes:

[0017] M3.1: Sample data deletion: Obtain the local reachability density and abnormality degree of each sample sampling point in the sub-areas with a quantity exceeding the standard, and delete the total amount of sample data according to the local reachability density and abnormality degree;

[0018] M3.2: Increase in sample data: Obtain the local density of each sample sampling point in the sub-regions that do not exceed the standard quantity, and based on the local density, divide the regions where the sample sampling points are located into high-density regions and low-density regions. Meanwhile, obtain the newly added sample sampling point data through the sample sampling point data in the low-density regions.

[0019] Furthermore, the deletion of the total amount of sample data includes:

[0020] N1: Obtain the local reachability density: According to the positions of the sample sampling points, obtain the distances between adjacent sample sampling points, determine the reachable distances between the sample sampling points and adjacent sample sampling points, and based on the reachable distances, obtain the local reachability density corresponding to the sample sampling points. Specifically:

[0021]

[0022] Where: is the local reachability density corresponding to the i-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point, is the reachable distance between the i-th sample sampling point and the j-th sample sampling point;

[0023] N2: Obtain the degree of abnormality: Through the local reachability density corresponding to the sample sampling points and the LOF calculation formula, obtain the LOF value corresponding to each sample sampling point. Specifically:

[0024]

[0025] Where: is the LOF value corresponding to the i-th sample sampling point, is the local reachability density corresponding to the i-th sample sampling point, is the local reachability density corresponding to the j-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point;

[0026] N3: Determine the deleted sampling points: According to the local reachability density and LOF value corresponding to the sample sampling points, obtain the initially deleted sample sampling points and the secondarily deleted sample sampling points, and based on the initially deleted sample sampling points and the secondarily deleted sample sampling points, determine the final deleted sampling points.

[0027] Furthermore, determining the final deleted sampling points includes:

[0028] N3.1: Obtain the quantity to be deleted: Obtain the difference in quantity between the total sample data volume and the standard quantity for each sub-region, and this difference in quantity is the quantity to be deleted;

[0029] N3.2: Determine the sample sampling points to be deleted: According to the LOF values corresponding to the sample sampling points, sort all the sample sampling points in descending order, and based on the sorting result, obtain the same number of sample sampling points as the quantity to be deleted, and these sample sampling points are the initially deleted sample sampling points. At the same time, according to the local reachability density corresponding to the sample sampling points, sort all the sample sampling points in ascending order, and based on the sorting result, obtain the same number of sample sampling points as the quantity to be deleted, and these sample sampling points are the secondarily deleted sample sampling points;

[0030] N3.3: Determine the initially deleted data: Match the initially deleted sample sampling points and the secondarily deleted sample sampling points, and determine the overlapping deleted sample sampling points, and these overlapping deleted sample sampling points are the initially deleted sample sampling points.

[0031] Furthermore, to determine the final deleted sampling points, it further includes:

[0032] N3.4: Determine the actual quantity to be deleted: Compare the sample quantity corresponding to the initially deleted sample sampling points with the quantity to be deleted, and based on the comparison result, determine the final deleted sampling points. Specifically:

[0033] When the sample quantity corresponding to the initially deleted sample sampling points is the same as the quantity to be deleted, then the initially deleted sample sampling points are the final deleted sampling points; otherwise, perform the next step;

[0034] N3.5: Determine the reorganized deleted data: Based on the sample quantity corresponding to the initially deleted sample sampling points and the quantity to be deleted, determine the remaining quantity to be deleted. At the same time, among the remaining sample sampling points sorted according to the local reachability density, obtain the secondarily deleted sample sampling points of the remaining quantity to be deleted, and combine the secondarily deleted sample sampling points of the remaining quantity to be deleted and the unmatched secondarily deleted sample sampling points to obtain the reorganized secondarily deleted sample sampling points;

[0035] Among the remaining sample sampling points sorted according to the LOF value, obtain the initially deleted sample sampling points of the remaining quantity to be deleted, and combine the initially deleted sample sampling points of the remaining quantity to be deleted and the unmatched initially deleted sample sampling points to obtain the reorganized initially deleted sample sampling points;

[0036] N3.6: Determine the final sampled points to be deleted: Match the initially deleted recombinant sampled points and the secondarily deleted recombinant sampled points to determine the overlapping recombinant sampled points to be deleted. Then compare the number of samples at the recombinant sampled points to be deleted with the remaining number of samples to be deleted, and based on the comparison result, determine the final sampled points to be deleted, specifically as follows:

[0037] When the number of samples at the recombinant sampled points to be deleted is the same as the remaining number of samples to be deleted, both the recombinant sampled points to be deleted and the initially deleted sampled points are the final sampled points to be deleted. Otherwise, reorder the recombinant sampled points to be deleted according to the local reachability density and determine the same number of additional sampled points to be deleted as the remaining number of samples to be deleted. Both the additional sampled points to be deleted and the initially deleted sampled points are the final sampled points to be deleted.

[0038] Furthermore, obtain the data of the newly added sampled points, including:

[0039] W1: Obtain the data to be added: Obtain the difference between the total amount of sample data and the standard amount in each sub-region, and this difference is the data to be added;

[0040] W2: Divide the density regions: Through a clustering algorithm, obtain the local density corresponding to each sampled point, and compare the local density with a preset density threshold. Based on the comparison result, determine the density region corresponding to each sampled point, specifically as follows:

[0041] When the local density corresponding to the sampled point is less than the preset density threshold, the density region corresponding to the sampled point is a low-density region; otherwise, the density region corresponding to the sampled point is a high-density region;

[0042] W3: Add the data of the sampled points: Based on the data of the sampled points in the low-density region and the data to be added, obtain the data of the newly added sampled points, specifically as follows:

[0043]

[0044] where: is the data of the newly added sampled points, is a random number, is the data of the i-th sampled point in the low-density region, is the data of the j-th sampled point in the low-density region.

[0045] Furthermore, construct a prediction model, including:

[0046] S2.1: Construct a comprehensive dataset: Combine the sample data after the preprocessing and sample balancing with external data to obtain a comprehensive dataset;

[0047] S2.2: Construct a prediction model: Divide the comprehensive dataset into a training set and a test set, and construct a prediction model through the training set and the test set. Specifically:

[0048]

[0049] Where: is the pollutant concentration corresponding to time t + 1, is the pollutant concentration corresponding to time t, is the precipitation corresponding to time t, is the soil permeability, is the optimization parameter, is the non - linear mapping function.

[0050] Furthermore, the acquisition formula of the optimization parameter is specifically:

[0051]

[0052] Where: is the gradient operator, is the comprehensive loss function, is the optimization parameter, is the task - specific loss function, is the hyperparameter of the relative importance of the density penalty term, is the hyperparameter for controlling the domain adversarial loss, is the domain adversarial loss function, is the density penalty term.

[0053] The evaluation system for the impact of soil on groundwater based on big data analysis uses the above - mentioned evaluation method for the impact of soil on groundwater based on big data analysis.

[0054] Compared with the prior art, the beneficial effects of the present invention are:

[0055] Through the sensor network and remote sensing technology, the present invention collects a sufficient number of high - quality samples from more diverse geographical regions, increasing the breadth and depth of the data. At the same time, taking the number of samples as the processing standard, the number of samples in different regions is increased or decreased, so that the number of samples in different regions remains consistent, thereby reducing the model bias caused by too many or too few samples in some regions, and further improving the representativeness and fairness of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flow schematic diagram of the evaluation method in the present invention;

[0057] Figure 2 It is a schematic flowchart for obtaining the sampling points of the final trimmed samples in the present invention;

[0058] Figure 3 It is a schematic flowchart for obtaining the sampling points of the recombined trimmed samples in the present invention;

[0059] Figure 4 It is a schematic flowchart for obtaining the sampling points of the newly added samples in the present invention. Detailed implementation manners

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0061] Most of the existing evaluation methods evaluate the impact of soil on groundwater through machine learning algorithms. However, during the prediction process, it often leads to unfair results due to the bias in the training dataset. For example, when the number of soil samples in some regions far exceeds that in other regions, during the prediction process, it will tend to more accurately reflect the situation in that region while ignoring the actual conditions in other regions. Although the bias can be reduced by introducing a regularization term or adjusting the weights, during the process of improving the algorithm transparency, it will pose an obstacle to practical applications as the performance decreases. The technical solution of this application obtains the standard quantity of sample data, increases or deletes the number of samples in different regions, makes the number of samples in different regions consistent, combines the processed sample data with external data to obtain a comprehensive dataset, and constructs a prediction model based on the comprehensive dataset, thereby reducing the model bias caused by too many or too few samples in some regions, and further improving the representativeness and fairness of the data.

[0062] Refer to Figures 1-4 , this embodiment provides an evaluation method for the impact of soil on groundwater based on big data analysis. The evaluation method includes the following steps:

[0063] Step S1: Adjust the sample quantity ratio. That is, divide the area to be analyzed to obtain multiple sub-regions and the soil samples and groundwater samples within their regions, and perform preprocessing and sample balancing processing on the obtained soil samples and groundwater samples. Specifically as follows:

[0064] Step S1.1: Obtain preprocessed sample data. That is, divide the area to be analyzed to obtain multiple sub-regions. Specifically, divide the region according to administrative regions (such as counties, cities) or natural geographical features (such as mountains, rivers) so that each sub-region has homogeneity and representativeness.

[0065] Furthermore, in the process of regional division, for sub-regions with uniform geographical environment distribution, a preset number (specifically set according to actual needs, not elaborated in this embodiment) of sampling positions are randomly selected within each sub-region for sample collection. For sub-regions with non-uniform geographical environment distribution, the sub-region is divided into multiple levels according to relevant characteristics (such as soil type, land use pattern), and at the same time, a relevant sample quantity is extracted at a preset ratio (specifically set according to actual needs, not elaborated in this embodiment) in each level.

[0066] It should be noted that during the sample collection process, for soil samples, stratified sampling is required at the selected positions, and for groundwater samples, water samples are extracted through drilling and other methods. At the same time, relevant environmental information of each sampling point needs to be recorded, such as climate conditions, vegetation coverage, and human activity intensity.

[0067] In this embodiment, relevant preprocessing is performed on the obtained samples to obtain preprocessed sample data. Specifically, through data cleaning, obvious errors in the sample data are obtained, such as values outside the reasonable range or duplicate records. At the same time, through data standardization and normalization processing, the sample data after data cleaning is unified. Furthermore, missing value filling is further performed on the unified sample data, that is, the missing data values are filled through mean / median filling, K-nearest neighbor interpolation method, etc., to ensure the integrity and accuracy of the sample data.

[0068] Step S1.2: Balance the sample data. That is, according to the total amount of sample data obtained, determine the standard quantity of the sample data. At the same time, according to the determined standard quantity, divide the sub-regions into sub-regions that exceed the standard quantity and sub-regions that do not exceed the standard quantity, and according to the division result, reduce the sample quantity corresponding to the sub-regions that exceed the standard quantity and increase the sample quantity corresponding to the sub-regions that do not exceed the standard quantity, so that the sample quantity in each sub-region is consistent. Specifically as follows:

[0069] Step M1: Determine the standard quantity of the sample data. That is, according to the total amount of sample data corresponding to all sub-regions, obtain the average sample quantity corresponding to the sample data, and this average sample quantity is the standard quantity of the sample data.

[0070] During the specific implementation process, 500 samples are obtained in sub-region A, 100 samples are obtained in sub-region B, and 200 samples are obtained in sub-region C. Therefore, the standard quantity corresponding to the three sub-regions A, B, and C is 267 samples.

[0071] Step M2: Conduct sub-region classification. That is, according to the standard quantity obtained in Step M1, all sub-regions are divided into two categories: sub-regions exceeding the standard quantity and sub-regions not exceeding the standard quantity. At the same time, according to the two divided categories, the corresponding classification category of each sub-region is determined, specifically as follows:

[0072] When the total amount of sample data corresponding to a sub-region is greater than the standard quantity in Step M1, the classification category corresponding to this sub-region is a sub-region exceeding the standard quantity; otherwise, the classification category corresponding to this sub-region is a sub-region not exceeding the standard quantity.

[0073] During the specific implementation process, according to the determined standard quantity of 267 samples, the classification category of sub-region A is a sub-region exceeding the standard quantity, and the classification categories of sub-regions B and C are sub-regions not exceeding the standard quantity.

[0074] Step M3: Adjust the sample quantity. That is, according to the classification category corresponding to each sub-region determined in Step M2, the total amount of sample data in the sub-regions exceeding the standard quantity is reduced, and the total amount of sample data in the sub-regions not exceeding the standard quantity is increased, so that the total amount of sample data in each sub-region is consistent with the standard quantity in Step M1, thus facilitating the centralized processing of the sample data of all sub-regions. Specifically as follows:

[0075] Step M3.1: Reduce sample data. That is, obtain the local reachability density and anomaly degree of each sample sampling point, and according to the local reachability density and anomaly degree, reduce the total amount of sample data in the sub-regions exceeding the standard quantity. Specifically as follows:

[0076] Step N1: Obtain the local reachability density. That is, according to the location of the sample sampling point, obtain the distance between it and the adjacent sample sampling points, and determine the reachable distance between this sample sampling point and its adjacent sample sampling points. At the same time, through the reachable distances corresponding to all sample sampling points, obtain the local reachability density corresponding to this sample sampling point, specifically as follows:

[0077]

[0078] Where: is the local reachability density corresponding to the i-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point, is the reachable distance between the i-th sample sampling point and the j-th sample sampling point.

[0079] In this embodiment, the reachable distance between the i-th sample sampling point and the j-th sample sampling point is specifically:

[0080]

[0081] Where: is the reachable distance between the i-th sample sampling point and the j-th sample sampling point, is the direct distance between the i-th sample sampling point and the j-th sample sampling point, is the k-distance of the j-th sample sampling point.

[0082] Step N2: Obtain the degree of abnormality. That is, through the local reachability density and LOF calculation formula corresponding to the sample sampling points obtained in step N1, obtain the LOF value corresponding to each sample sampling point. That is to say, through the LOF value corresponding to the sample sampling point, determine the degree of abnormality of the sample sampling point. Specifically, when the LOF value corresponding to the sample sampling point is close to 1, the corresponding sample sampling point is a normal sampling point, otherwise, the sample sampling point is an abnormal sampling point. In this embodiment, the LOF calculation formula is specifically:

[0083]

[0084] Where: is the LOF value corresponding to the i-th sample sampling point, is the local reachability density corresponding to the i-th sample sampling point, is the local reachability density corresponding to the j-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point.

[0085] Step N3: Determine the deleted sampling points. That is, according to the LOF value corresponding to the sample sampling point obtained in step N2 and the local reachability density corresponding to the sample sampling point obtained in step N1, sort all the sample sampling points in the sub-region respectively, and according to the quantity difference between the total quantity of sample data in each sub-region and the standard quantity, determine the initially deleted sample sampling points and the secondarily deleted sample sampling points from all the sorted sample sampling points, and at the same time, determine the final deleted sampling points according to the initially deleted sample sampling points and the secondarily deleted sample sampling points. Specifically as follows:

[0086] Step N3.1: Obtain the quantity to be deleted. That is, based on the total amount of sample data in each sub-region and the standard quantity of the sample data obtained in Step M1, determine the quantity difference between the total amount of sample data and the standard quantity. This quantity difference is the quantity of the sample data to be deleted.

[0087] Step N3.2: Determine the sampling points of the deleted samples. That is, based on the LOF values corresponding to the sample sampling points obtained in Step N2, sort all the sample sampling points in descending order. At the same time, for all the sample sampling points sorted in descending order according to the LOF values, successively obtain the initial deleted sample sampling points that are consistent with the quantity of the sample data to be deleted. Similarly, based on the local reachability density corresponding to the sample sampling points obtained in Step N1, sort all the sample sampling points in ascending order. And for all the sample sampling points sorted in ascending order according to the local reachability density, successively obtain the secondary deleted sample sampling points that are consistent with the quantity of the sample data to be deleted.

[0088] Step N3.3: Determine the preliminary deleted data. Match the initial deleted sample sampling points and the secondary deleted sample sampling points determined in Step N3.2, and determine the overlapping deleted sample sampling points from them. These overlapping deleted sample sampling points are the preliminary deleted sample sampling points.

[0089] Step N3.4: Determine the actual deletion quantity. That is, compare the quantity of the preliminary deleted sample sampling points obtained in Step N3.3 with the quantity to be deleted obtained in Step N3.1, and based on the comparison result, determine whether the total amount of sample data in the sub-region is the same as the standard quantity of the sample data. Specifically:

[0090] When the quantity of the preliminary deleted sample sampling points is the same as the quantity to be deleted, the preliminary deleted sample sampling points are the final deleted sampling points. Otherwise, execute the next step until the quantity of the deleted sample sampling points is the same as the quantity to be deleted, that is, all the deleted sample sampling points are the final deleted sampling points.

[0091] Step N3.5: Determine the recombined and pruned data. That is, when the number of the initially pruned data obtained in Step N3.3 is less than the number of the sample data to be pruned, among the remaining sample sampling points sorted in ascending order according to the local reachability density, obtain the secondary pruned sample sampling points with the remaining number to be pruned. Similarly, among the remaining sample sampling points sorted in descending order according to the LOF value, obtain the initially pruned sample sampling points with the remaining number to be pruned. Combine the secondary pruned sample sampling points with the remaining number to be pruned and the secondary pruned sample sampling points not matched in Step N3.3 to obtain the recombined secondary pruned sample sampling points. Similarly, combine the initially pruned sample sampling points with the remaining number to be pruned and the initially pruned sample sampling points not matched in Step N3.3 to obtain the recombined initially pruned sample sampling points.

[0092] Step N3.6: Determine the final pruned sampling points. That is, match the recombined initially pruned sample sampling points and the recombined secondary pruned sample sampling points obtained in Step N3.5, and determine the coincident recombined pruned sample sampling points therefrom. Compare the number of samples of the recombined pruned sample sampling points with the remaining number to be pruned, and determine the final pruned sample sampling points according to the comparison result. Specifically:

[0093] When the number of samples of the recombined pruned sample sampling points is the same as the remaining number to be pruned, both the recombined pruned sample sampling points and the initially pruned sample sampling points are the final pruned sample sampling points. On the contrary, when the number of samples of the recombined pruned sample sampling points is greater than the remaining number to be pruned, re - sort the recombined pruned sample sampling points in ascending order according to the local reachability density, and according to the re - arranged order, determine the sample sampling points with the same number as the remaining number to be pruned, that is, the sample sampling points are the re - pruned sample sampling points, then both the re - pruned sample sampling points and the initially pruned sample sampling points are the final pruned sample sampling points.

[0094] In the process of specific implementation, a data table as shown in Table 1 below is set, specifically:

[0095] Table 1: Data Table

[0096]

[0097] Furthermore, there are 10 sample sampling points set in Table 1 above. In this embodiment, the standard quantity of sample data is set to 7, so the quantity to be deleted is 3. Specifically, sorting according to the LOF value, the sample sorting is: E, C, I, G, A, D, F, B, H, J. Sorting according to the local reachability density, the sample sorting is: H, D, B, J, G, A, F, I, C, E. That is to say, the initially deleted sample sampling points are: E, C, and I, and the secondarily deleted sample sampling points are: H, D, and B. Since there are no overlapping sample sampling points between the initially deleted sample sampling points and the secondarily deleted sample sampling points, recombinant deleted data is continued to be obtained, that is, the recombined initially deleted sample sampling points are: E, C, I, G, A, and D, and the recombined secondarily deleted sample sampling points are: H, D, B, J, G, and A. Furthermore, the overlapping sample sampling points in the recombined initially deleted sample sampling points and the recombined secondarily deleted sample sampling points are: A, D, and G. That is to say, the determined quantity of deleted samples is the same as the quantity to be deleted, so the final deleted sample sampling points are: A, D, and G.

[0098] Step M3.2: Increasing sample data. That is, obtain the local density of each sample sampling point in the sub-region that does not exceed the standard quantity, and divide the region where the sample sampling point is located according to the local density, that is, divide it into a high-density region and a low-density region, and synthesize new sample sampling point data according to the data of the sample sampling points in the low-density region. Specifically as follows:

[0099] Step W1: Obtain the data to be added. That is, obtain the total quantity of sample data of each sub-region that does not exceed the standard quantity and the standard quantity of the sample data obtained in Step M1, and determine the quantity difference between the total quantity of sample data and the standard quantity. This quantity difference is the quantity of sample data to be added.

[0100] Step W2: Divide the density region. That is, through the clustering algorithm, obtain the local density around each sample sampling point. Specifically:

[0101]

[0102] Wherein: is the local density corresponding to the sample sampling point P, is the distance between the sample sampling point P and the sample sampling point Q, is the kernel function, is the sample sampling point Q, is the sample sampling point P, is the set of all other sample sampling points within the domain of the sample sampling point P.

[0103] Further, compare the local density corresponding to each sample sampling point with a preset density threshold (which is specifically set according to the actual scenario and will not be elaborated in this embodiment), and determine the density area corresponding to each sample sampling point according to the comparison result. Specifically:

[0104] When the local density corresponding to the sample sampling point is less than the preset density threshold, the density area corresponding to this sample sampling point is a low-density area; otherwise, the density area corresponding to the sample sampling point is a high-density area.

[0105] Step W3: Add sample sampling point data. That is, according to the low-density area determined in Step W2, gather all the sample sampling points in the low-density area, and add new sample sampling point data according to the sample sampling point data in the low-density area. Specifically:

[0106]

[0107] Where: is the newly added sample sampling point data, is a random number, is the data of the i-th sample sampling point in the low-density area, is the data of the j-th sample sampling point in the low-density area.

[0108] Further, according to the data to be added obtained in Step W1, obtain the same number of newly added sample sampling point data.

[0109] In the process of specific implementation, the data of two sample sampling points in the low-density area are respectively: X (x1 = 0.5, x2 = 0.7) and Y (x1 = 0.4, x2 = 0.6), and at the same time the random number is set to 0.3, so the new sample sampling point data is: (x1 = 0.47, x2 = 0.67).

[0110] Step S2: Establish a prediction model. That is, combine the total amount of sample data after increasing and decreasing in Step M3 with external data (such as soil chemical composition, climate conditions, and geological structure) to obtain a comprehensive data set. At the same time, construct a prediction model according to the comprehensive data set. Specifically as follows:

[0111] Step S2.1: Construct a comprehensive data set. That is, combine the total amount of sample data after increasing and decreasing in Step M3 with external data to obtain a comprehensive data set. Specifically, the soil chemical composition includes pH value and organic matter content, the climate conditions include annual average precipitation and seasonal temperature changes, and the geological structure includes rock layer distribution and fault line location.

[0112] Step S2.2: Build a prediction model. That is, divide the comprehensive data set proportionally into a training set and a test set. Specifically, divide the comprehensive data set in a ratio of 6:4, where one part is used as the training set for training and establishing the prediction model, and the other part is used as the test set for testing the established prediction model. In this embodiment, the prediction model is specifically:

[0113]

[0114] Where: is the pollutant concentration corresponding to time t+1, is the pollutant concentration corresponding to time t, is the precipitation corresponding to time t, is the soil permeability, is the optimization parameter, is the non-linear mapping function.

[0115] Furthermore, the acquisition formula for the optimization parameter is specifically:

[0116]

[0117] Where: is the gradient operator, is the comprehensive loss function, is the optimization parameter, is the task-specific loss function, is the hyperparameter for the relative importance of the density penalty term, is the hyperparameter for controlling the domain adversarial loss, is the domain adversarial loss function, is the density penalty term.

[0118] Step S3: Obtain the impact assessment. That is, according to the prediction model established in Step S2.2, use the data of the sample sampling points obtained in real time as the input, and output the corresponding assessment data.

[0119] In the process of specific implementation, the soil pollutant concentration corresponding to the current time is 5.2 mg / L, the precipitation is 80 mm, and the soil permeability is 0.001 cm / s. At the same time, it is known that the precipitation at the adjacent future time is 60 mm and the soil permeability is 0.001 cm / s, then the soil pollutant concentration at the adjacent future time is 5.8 mg / L.

[0120] This embodiment also provides an assessment system for the impact of soil on groundwater based on big data analysis. This adaptive control system uses the assessment method for the impact of soil on groundwater based on big data analysis described in the above embodiment.

[0121] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An assessment method for the impact of soil on groundwater based on big data analysis, characterized in that, It includes: S1: Adjust the sample quantity ratio: Divide the area to be analyzed to obtain multiple sub - regions and the sample data within the sub - regions, and perform pre - processing and sample balancing on the sample data; S2: Establish a prediction model: Combine the pre - processed and sample - balanced sample data with external data, and construct a prediction model based on the combined data set; S3: Obtain real - time evaluation data: Use the sample sampling point data obtained in real - time as the input of the prediction model, and output the obtained real - time evaluation data; Performing sample balancing on the sample data includes: M1: Determine the standard quantity of sample data: Obtain the average sample quantity based on the total sample data volume corresponding to all sub - regions, and determine the standard quantity of sample data; M2: Classify sub - regions: Divide the sub - regions into sub - regions with a quantity exceeding the standard and sub - regions with a quantity not exceeding the standard according to the standard quantity. Specifically: When the total sample data volume corresponding to the sub - region is greater than the standard quantity, then the sub - region is a sub - region with a quantity exceeding the standard; otherwise, the sub - region is a sub - region with a quantity not exceeding the standard. M3: Adjust the sample quantity: Increase or decrease the total sample data volume corresponding to the sub - region according to the classification category corresponding to the sub - region, including: M3.1: Sample data deletion: Obtain the local reachability density and abnormality degree of each sample sampling point in the sub - region with a quantity exceeding the standard, and delete the total sample data volume according to the local reachability density and abnormality degree, including: N1: Obtain the local reachability density: According to the position of the sample sampling point, obtain the distance between adjacent sample sampling points, determine the reachable distance between the sample sampling point and adjacent sample sampling points, and obtain the local reachability density corresponding to the sample sampling point according to the reachable distance. Specifically: , where: is the local reachability density corresponding to the i-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point, is the reachability distance between the i-th sample sampling point and the j-th sample sampling point; N2: Obtain the abnormality degree: Obtain the LOF value corresponding to each sample sampling point through the local reachability density corresponding to the sample sampling point and the LOF calculation formula. Specifically: , where: is the LOF value corresponding to the i-th sample sampling point, is the local reachability density corresponding to the i-th sample sampling point, is the local reachability density corresponding to the j-th sample sampling point, is the set of k nearest neighbors of the i-th sample sampling point, is the i-th sample sampling point, is the j-th sample sampling point; N3: Determine the deleted sampling points: Obtain the initially deleted sample sampling points and secondarily deleted sample sampling points according to the local reachability density and LOF value corresponding to the sample sampling point, and determine the final deleted sampling points according to the initially deleted sample sampling points and secondarily deleted sample sampling points, including: N3.1: Obtain the quantity to be deleted: Obtain the difference between the total sample data volume and the standard quantity of each sub - region, and the difference is the quantity to be deleted; N3.2: Determine the deleted sample sampling points: Sort all the sample sampling points in descending order according to the LOF value corresponding to the sample sampling point, and obtain the same number of sample sampling points as the quantity to be deleted according to the sorting result. The sample sampling points are the initially deleted sample sampling points. At the same time, sort all the sample sampling points in ascending order according to the local reachability density corresponding to the sample sampling point, and obtain the same number of sample sampling points as the quantity to be deleted according to the sorting result. The sample sampling points are the secondarily deleted sample sampling points; N3.3: Determine the preliminary deleted data: Match the preliminary deleted sample sampling points and the secondary deleted sample sampling points to determine the coincident deleted sample sampling points, and the coincident deleted sample sampling points are the preliminary deleted sample sampling points; N3.4: Determine the actual deletion quantity: Compare the sample quantity corresponding to the preliminary deleted sample sampling points with the quantity to be deleted, and determine the final deleted sampling points according to the comparison result. Specifically: When the sample quantity corresponding to the preliminary deleted sample sampling points is the same as the quantity to be deleted, the preliminary deleted sample sampling points are the final deleted sampling points; otherwise, proceed to the next step; N3.5: Determine the reorganized deleted data: Determine the remaining quantity to be deleted according to the sample quantity corresponding to the preliminary deleted sample sampling points and the quantity to be deleted. At the same time, among the remaining sample sampling points sorted according to the local reachability density, obtain the secondary deleted sample sampling points of the remaining quantity to be deleted, and combine the secondary deleted sample sampling points of the remaining quantity to be deleted and the unmatched secondary deleted sample sampling points to obtain the reorganized secondary deleted sample sampling points; Among the remaining sample sampling points sorted according to the LOF value, obtain the preliminary deleted sample sampling points of the remaining quantity to be deleted, and combine the preliminary deleted sample sampling points of the remaining quantity to be deleted and the unmatched preliminary deleted sample sampling points to obtain the reorganized preliminary deleted sample sampling points; N3.6: Determine the final deleted sampling points: Match the reorganized preliminary deleted sample sampling points and the reorganized secondary deleted sample sampling points to determine the coincident reorganized deleted sample sampling points, and compare the sample quantity of the reorganized deleted sample sampling points with the remaining quantity to be deleted, and determine the final deleted sample sampling points according to the comparison result. Specifically: When the sample quantity of the reorganized deleted sample sampling points is the same as the remaining quantity to be deleted, the reorganized deleted sample sampling points and the preliminary deleted sample sampling points are both the final deleted sample sampling points; otherwise, re - sort the reorganized deleted sample sampling points according to the local reachability density, and determine the re - deleted sample sampling points with the same quantity as the remaining quantity to be deleted. The re - deleted sample sampling points and the preliminary deleted sample sampling points are both the final deleted sample sampling points; M3.2: Increase sample data: Obtain the local density of each sample sampling point in the sub - regions that do not exceed the standard quantity, and divide the regions where the sample sampling points are located into high - density regions and low - density regions according to the local density. At the same time, obtain the new sample sampling point data through the sample sampling point data in the low - density region.

2. The evaluation method for the impact of soil on groundwater based on big data analysis according to claim 1, characterized in that Obtaining the new sample sampling point data includes: W1: Obtain the data to be added: Obtain the difference between the total sample data quantity of each sub - region and the standard quantity, and the difference is the data to be added; W2: Divide the density regions: Through the clustering algorithm, obtain the local density corresponding to each sample sampling point, and compare the local density with the preset density threshold. According to the comparison result, determine the density region corresponding to each sample sampling point. Specifically: When the local density corresponding to the sample sampling point is less than the preset density threshold, the density area corresponding to the sample sampling point is a low-density area; otherwise, the density area corresponding to the sample sampling point is a high-density area. W3: Add sample sampling point data: Obtain the newly added sample sampling point data according to the sample sampling point data and the data to be added in the low-density area, specifically: , where: is the newly added sample sampling point data, is a random number, is the i-th sample sampling point data in the low-density area, is the j-th sample sampling point data in the low-density area.

3. The assessment method for the impact of soil on groundwater based on big data analysis according to claim 1, wherein Construct a prediction model, including: S2.1: Construct a comprehensive data set: Combine the sample data after the preprocessing and sample balancing process with the external data to obtain a comprehensive data set. S2.2: Construct a prediction model: Divide the comprehensive data set into a training set and a test set, and construct a prediction model through the training set and the test set, specifically: , where: is the pollutant concentration corresponding to time t + 1, is the pollutant concentration corresponding to time t, is the precipitation corresponding to time t, is the soil permeability, is the optimization parameter, is the non - linear mapping function.

4. The assessment method for the impact of soil on groundwater based on big data analysis according to claim 3, characterized in that The acquisition formula of the optimization parameters is specifically: , where: is the gradient operator, is the comprehensive loss function, is the optimization parameter, is the task-specific loss function, is the hyperparameter for the relative importance of the density penalty term, is the hyperparameter for controlling the domain adversarial loss, is the domain adversarial loss function, is the density penalty term.

5. An assessment system for the impact of soil on groundwater based on big data analysis, characterized in that, The evaluation method for the impact of soil on groundwater based on big data analysis according to any one of claims 1-4 is used.

Citation Information

Patent Citations

  • Climate similarity migration-based soil and water fusion method for areas lacking data

    CN117494042A

  • Wind power prediction gale data enhancement method considering extreme gale weather

    CN114077924A

  • Underground water pollution site risk assessment system

    CN118393097A