Method and device for reorganizing mass underground water monitoring data

By combining machine learning and manual recognition technology, the massive groundwater monitoring data is subjected to multiple rounds of abnormal processing, which solves the problems of low efficiency and poor quality of traditional methods, and achieves efficient and accurate data reorganization and quality assurance.

CN120104601AInactive Publication Date: 2025-06-06中国地质环境监测院(自然资源部地质灾害技术指导中心)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510163709.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional groundwater monitoring data reorganization methods are inefficient, have poor visibility, low controllability and insufficient compatibility, making it difficult to effectively process massive monitoring data.

Method used

Using a combination of machine learning and manual recognition, a large amount of groundwater monitoring data is subjected to multiple rounds of abnormal processing, including data cleaning, normal testing, noise processing and depth anomaly analysis. Through local anomaly factor algorithm, TukeyBoxplot algorithm, STL time series decomposition algorithm and isolated forest technologies, the abnormal data is determined and eliminated, and the reorganized data is finally obtained.

Benefits of technology

It improves the efficiency and results of the reorganization of groundwater monitoring data, ensures the quality and reliability of the reorganized data, and provides scientific and accurate data support for the development, utilization and protection of groundwater resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104601A_ABST
    Figure CN120104601A_ABST
Patent Text Reader

Abstract

The invention discloses a reorganization method and device for mass underground water monitoring data, and relates to the field of data processing, and the method comprises the steps: obtaining to-be-reorganized underground water monitoring data, and carrying out the data cleaning and noise processing, and obtaining the denoised monitoring data; performing deep analysis on the denoised monitoring data by adopting a local abnormal factor algorithm, a TukeyBoxplot algorithm, an STL time sequence decomposition algorithm and an isolated forest, and determining abnormal data; obtaining an artificial discrimination result of the abnormal data; and removing abnormal data in an artificial discrimination result to obtain reorganized underground water monitoring data. According to the method, data exception identification is carried out in a machine learning mode, then marked exceptional data are further identified and judged manually, and multi-round exception processing is carried out on mass monitoring data by using a means of combining machine learning and manual identification, so that the efficiency of mass data reorganization is improved, and the quality of reorganized data is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a method and device for compiling massive groundwater monitoring data. Background Art

[0002] The compilation of groundwater monitoring data is an important part of groundwater monitoring work and an indispensable basic work for collecting, integrating and processing monitoring information. Due to its discreteness, multi-source error and incompleteness, the original groundwater monitoring data cannot be directly used for production and management. Therefore, it is necessary to use certain methods and techniques to analyze, organize, review and compile the original monitoring data to make it a systematic, complete and accurate compilation result.

[0003] Groundwater monitoring has the characteristics of high frequency and long cycle. With the passage of time, the encryption of station networks and the continuous development of information technology, the amount of data will also grow rapidly to form a massive monitoring database. The traditional compilation method has problems such as singleness, poor visibility, low controllability, insufficient compatibility, and low efficiency. In order to solve the problems of efficiency and quality of groundwater monitoring data compilation, the present invention proposes a compilation method and device for massive groundwater monitoring data. Summary of the invention

[0004] The purpose of this application is to provide a method and device for compiling massive groundwater monitoring data, which can greatly improve the efficiency and quality of groundwater monitoring data compilation work, and provide more scientific and accurate data support for the development, utilization and protection of groundwater resources.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for compiling massive groundwater monitoring data, comprising:

[0007] Obtain groundwater monitoring data to be compiled;

[0008] Cleaning the groundwater monitoring data to be compiled according to the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database, marking abnormal data, and obtaining cleaned monitoring data; the cleaned monitoring data is marked with abnormal data;

[0009] According to the historical groundwater monitoring data stored in the groundwater monitoring database, the cleaned monitoring data is subjected to normality test and noise processing to obtain denoised monitoring data; the denoised monitoring data is marked with abnormal data;

[0010] The local anomaly factor algorithm, TukeyBoxplot algorithm, STL time series decomposition algorithm and isolation forest are used to perform in-depth anomaly analysis on the denoised monitoring data to determine the abnormal data in the data;

[0011] Obtaining manual identification results of the abnormal data;

[0012] In the denoised monitoring data, abnormal data in the manual identification results are eliminated to obtain the compiled groundwater monitoring data.

[0013] In a second aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for compiling massive groundwater monitoring data.

[0014] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for compiling massive groundwater monitoring data.

[0015] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for compiling massive groundwater monitoring data.

[0016] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0017] The present application provides a compilation method and device for massive groundwater monitoring data, which obtains groundwater monitoring data to be compiled; performs data cleaning on the groundwater monitoring data to be compiled according to basic information of groundwater monitoring sites and historical groundwater monitoring data stored in a groundwater monitoring database to obtain cleaned monitoring data; abnormal data is marked in the cleaned monitoring data; normality test and noise processing are performed on the cleaned monitoring data according to the historical groundwater monitoring data stored in the groundwater monitoring database to obtain denoised monitoring data; abnormal data is marked in the denoised monitoring data; a local anomaly factor algorithm, a TukeyBoxplot algorithm, an STL time series decomposition algorithm and an isolation forest are used to perform deep anomaly analysis on the denoised monitoring data to determine the abnormal data in the data; obtains manual discrimination results of the abnormal data; in the denoised monitoring data, the abnormal data in the manual discrimination results are eliminated to obtain the compiled groundwater monitoring data. The present invention first uses machine learning to identify data anomalies, and then further identifies and judges the marked abnormal data manually. It uses a combination of machine learning and manual recognition to perform multiple rounds of anomaly processing on massive monitoring data, thereby improving the efficiency of large-scale data compilation and ensuring the quality of the compiled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 This is an application environment diagram of a method for compiling massive groundwater monitoring data in one embodiment of the present application;

[0020] Figure 2 A flowchart of a method for compiling massive groundwater monitoring data provided by an embodiment of the present application;

[0021] Figure 3 A schematic diagram of the calculation of the TukeyBoxplot algorithm provided in one embodiment of the present application;

[0022] Figure 4 Schematic diagram of each sequence of the STL time series decomposition algorithm provided in one embodiment of the present application;

[0023] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0025] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] The method for compiling massive groundwater monitoring data provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the groundwater monitoring data to be compiled to the server. After the server receives the groundwater monitoring data to be compiled, the server performs data cleaning on the groundwater monitoring data to be compiled according to the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database, and marks the abnormal data at the same time to obtain the cleaned monitoring data; the cleaned monitoring data is marked with abnormal data; the cleaned monitoring data is subjected to normality test and noise processing according to the historical groundwater monitoring data stored in the groundwater monitoring database to obtain the denoised monitoring data; the denoised monitoring data is marked with abnormal data; the local anomaly factor algorithm, TukeyBoxplot algorithm, STL time series decomposition algorithm and isolation forest are used to perform deep anomaly analysis on the denoised monitoring data to determine the abnormal data in the data; the manual discrimination result of the abnormal data is obtained; in the denoised monitoring data, the abnormal data in the manual discrimination result is eliminated to obtain the compiled groundwater monitoring data. The server can feed back the compiled groundwater monitoring data to the terminal. In addition, in some embodiments, the compilation method for massive groundwater monitoring data can also be implemented by the server or the terminal alone, such as the terminal can directly compile the groundwater monitoring data to be compiled, or the server can obtain the groundwater monitoring data to be compiled from the data storage system and compile it.

[0027] The terminal may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0028] In an exemplary embodiment, Figure 2 As shown, a method for compiling massive groundwater monitoring data is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server in the example is used to illustrate, including the following steps 101 to 106. Among them:

[0029] Step 101, obtaining groundwater monitoring data to be compiled.

[0030] Step 102, data cleaning is performed on the groundwater monitoring data to be compiled based on the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database to obtain the cleaned monitoring data; the cleaned monitoring data is marked with abnormal data. The data cleaning here mainly removes the erroneous data that violates the laws of nature and marks the abnormal data.

[0031] Step 103, performing normality test and noise processing on the cleaned monitoring data according to the historical groundwater monitoring data stored in the groundwater monitoring database to obtain denoised monitoring data; the denoised monitoring data is marked with abnormal data.

[0032] Step 104, using a local anomaly factor algorithm, a Tukey Boxplot algorithm, an STL time series decomposition algorithm and an isolation forest to perform in-depth analysis on the denoised monitoring data to determine abnormal data in the data.

[0033] Step 105, obtaining the manual identification result of the abnormal data.

[0034] Step 106, in the denoised monitoring data, the abnormal data in the manual identification results are removed to obtain the compiled groundwater monitoring data.

[0035] In implementing the above-mentioned steps 101 to 106, the present application first uses machine learning to identify data anomalies, and then further manually identifies and judges the marked abnormal data. A combination of machine learning and manual identification is used to perform multiple rounds of anomaly processing on massive monitoring data, thereby improving the efficiency of large-scale data compilation and ensuring the quality of the compiled data.

[0036] In another exemplary embodiment of the present application, in step 102, the groundwater monitoring data to be compiled is cleaned according to the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database to obtain the cleaned monitoring data, which specifically includes:

[0037] (a1) Based on the basic information of the groundwater monitoring site, the data in the groundwater monitoring data to be compiled, in which the groundwater level is lower than the bottom elevation of the monitoring well, are identified and removed.

[0038] Call up the basic information of the monitoring site, load the original monitoring data (i.e. the groundwater monitoring data to be compiled), and identify the data relationship: if it is identified that the groundwater level is lower than the bottom elevation of the monitoring well, the data will be discarded.

[0039] (a2) Based on the basic information of the groundwater monitoring site, the data whose measured groundwater depth exceeds the probe range of the monitoring equipment in the groundwater monitoring data to be compiled are identified and removed.

[0040] Call up the basic information of the monitoring site, load the original monitoring data (i.e. the groundwater monitoring data to be compiled), and identify the data relationship: if it is identified that the measured value of the groundwater depth exceeds the probe range of the monitoring equipment, the data will be discarded.

[0041] (a3) Based on the basic information of the groundwater monitoring site, the data of the groundwater level in the non-artesian monitoring wells in the groundwater monitoring data to be compiled is identified as abnormal data.

[0042] Call the basic information of the monitoring site, load the original monitoring data (i.e. the groundwater monitoring data to be compiled), and identify the data relationship: if it is identified that the groundwater level of the non-self-flowing monitoring well is higher than the ground elevation, mark the data.

[0043] (a4) Based on the historical groundwater monitoring data stored in the groundwater monitoring database, data in the groundwater monitoring data to be compiled whose groundwater level is higher than the historical highest water level are identified and marked as abnormal data.

[0044] Call the historical monitoring database (already compiled), load the original monitoring data, and identify the data relationship: if it is identified that the groundwater level is higher than the historical highest water level, mark the data.

[0045] (a5) Based on the historical groundwater monitoring data stored in the groundwater monitoring database, identify the data in the groundwater monitoring data to be compiled whose groundwater level is lower than the historical lowest water level, and mark them as abnormal data. The cleaned monitoring data refers to the data obtained after data removal and abnormal data marking of the groundwater monitoring data to be compiled.

[0046] Call the historical monitoring database (already compiled), load the original monitoring data, and identify the data relationship: if it is identified that the groundwater level is lower than the historical lowest water level, mark the data.

[0047] In another exemplary embodiment of the present application, in step 103, the groundwater monitoring data is subjected to a normality test: the historical monitoring database (already compiled) is called, the original monitoring data is loaded, and a normality test is performed on the data of the same period of the historical year. If it conforms to the normal distribution, a 3σ detection is performed based on the Chebyshev inequality, and data outside the range of (μ-3σ, μ+3σ) is identified and marked. μ is the mean value of the target data set, and σ is the standard deviation of the target data set. The target data set refers to the historical groundwater monitoring data in the groundwater monitoring database that is in the same period as the groundwater monitoring data to be compiled.

[0048] De-noising the groundwater monitoring data: ① Call the historical monitoring database (already compiled), and use FFT (Fast Fourier Transform) to obtain the frequency components of the historical data. ② Load the monitoring data, integrate it into the historical monitoring data to form a data sequence, use wavelet transform to obtain the time-frequency spectrum of the data sequence, identify the high-frequency content outside the frequency components of the historical data, and determine the time point when it appears in the time domain: if the content appears in both the time domain of the historical monitoring data and the time domain of the monitoring data to be compiled, ignore the problem; if it only appears in the time domain of the monitoring data to be compiled, mark the corresponding data as noise. ③ Use the moving average method to remove the marked noise data. Therefore, in step 103, the cleaned monitoring data is subjected to normality test and noise processing according to the historical groundwater monitoring data stored in the groundwater monitoring database to obtain the denoised monitoring data, which specifically includes:

[0049] (b1) performing a normality test on historical groundwater monitoring data in the groundwater monitoring database that is contemporaneous with the groundwater monitoring data to be compiled, identifying data outside a preset normality test range in the groundwater monitoring data to be compiled, and marking the data as abnormal data.

[0050] (b2) The frequency components of the historical groundwater monitoring data stored in the groundwater monitoring database are obtained using fast Fourier transform.

[0051] (b3) Integrating the cleaned monitoring data with the historical groundwater monitoring data to obtain an integrated data sequence.

[0052] (b4) Obtaining the time-frequency spectrum of the integrated data sequence using wavelet transform.

[0053] (b5) identifying high-frequency data other than the frequency components in the integrated data sequence according to the time-frequency spectrum.

[0054] (b6) Determine the time point at which the high-frequency data appears in the time domain.

[0055] (b7) If the time point of the high-frequency data appears both within the time domain of the historical monitoring data and within the time domain of the cleaned monitoring data, the corresponding high-frequency data is not processed.

[0056] (b8) If the time point of the high-frequency data only appears in the time domain of the cleaned monitoring data, the corresponding high-frequency data is marked as noise data.

[0057] (b9) Using a moving average method to remove the marked noise data, and obtaining the denoised monitoring data.

[0058] In another exemplary embodiment of the present application, in step 104, outliers in the monitoring data are identified based on a local outlier factor algorithm (LOF):

[0059]

[0060] Among them: LOF i is the local anomaly factor of data point i; N k (i) is the k nearest neighbors of data point i, where k represents the size of the neighborhood; MD(i, j) is the distance between data points i and j; and MD(j, j) is the distance between data points j and j.

[0061] In the anomaly identification of groundwater monitoring data, k is determined by the monitoring frequency. When the number of measurements per day is n, k can be set to 2(n-1) or 10(n-1), and the corresponding LOF thresholds are set to 2 and 1.5. Monitoring data with LOF exceeding the threshold are marked as outliers.

[0062] Therefore, in step 104, a local anomaly factor algorithm is used to perform in-depth analysis on the denoised monitoring data to determine abnormal data in the data, specifically including:

[0063] (c1) For each data point in the denoised monitoring data, determining k nearest neighboring data points of the data point.

[0064] (c2) Calculating the distance between the data point and each corresponding nearest neighbor data point to obtain a first distance.

[0065] (c3) Calculate the distance between any two nearest neighbor data points among the k nearest neighbor data points to obtain a second distance.

[0066] (c4) Calculating a local outlier factor of the data point based on the first distance and the second distance.

[0067] (c5) Marking the data points whose local anomaly factors exceed a preset threshold as abnormal data.

[0068] In another exemplary embodiment of the present application, in step 104, TukeyBoxplot is used to identify abnormal values ​​in the monitoring data, such as Figure 3 As shown, Q 1 is the lower quartile; Q 2 is the median; Q 3 is the upper quartile; IQR is the interquartile range, i.e. Q 3 -Q 1 ;Q 1 -1.5*IQR is the upper limit; Q 3 +1.5*IQR is the lower limit. In the anomaly identification of groundwater monitoring data, TukeyBoxplot selects the monitoring data of a whole month as the data set, that is, the time window is set to a whole month. It can be considered to shift the time window with a step length of half a month to increase the frequency of anomaly identification, and the monitoring data above the upper limit of Boxplot and below the lower limit are marked as outliers. Therefore, in step 104, the denoised monitoring data is deeply analyzed according to the TukeyBoxplot algorithm to determine the abnormal data in the data, specifically including:

[0069] (d1) Dividing the denoised monitoring data into a plurality of monitoring data sets according to a preset time window.

[0070] (d2) for each monitoring data set, determining the median of the monitoring data set.

[0071] (d3) determining the upper quartile and the lower quartile of the monitoring data set based on the median.

[0072] (d4) determining an interquartile range according to the upper quartile and the lower quartile.

[0073] (d5) Determine the upper limit and lower limit of Boxplot according to the interquartile range.

[0074] (d6) Marking the data above the upper limit of the Boxplot and the data below the lower limit of the Boxplot in the denoised monitoring data as abnormal data.

[0075] In another exemplary embodiment of the present application, in step 104, STL (Seasonal and Trend decomposition using Loess) time series decomposition is used to identify outliers in the monitoring data: after the monitoring data is loaded, the original data sequence is decomposed into a seasonal cycle sequence, a trend sequence, and a residual sequence by weighted least squares regression. Figure 4 As shown, Figure 4In the data, data is the denoised monitoring data; seasonal is the seasonal cycle sequence; trend is the trend sequence; remainder is the residual sequence. In the anomaly identification of groundwater monitoring data, the STL model assigns a robust weight to each monitoring time point based on the median of the residual sequence and the residual value of the monitoring data, and the monitoring data corresponding to the monitoring time point with a low weight is marked as an outlier. Therefore, in step 104, the denoised monitoring data is deeply analyzed according to the STL time series decomposition algorithm to determine the abnormal data in the data, specifically including:

[0076] (e1) Using the STL time series decomposition algorithm, the denoised monitoring data is decomposed into a seasonal cycle series, a trend series and a residual series.

[0077] (e2) assigning a weight value to each monitoring time point in the denoised monitoring data according to the median of the trend sequence and the residual sequence.

[0078] (e3) Marking the data whose weight value at the monitoring time point in the denoised monitoring data is lower than the preset weight value as abnormal data.

[0079] In another exemplary embodiment of the present application, in step 104, iForest (Isolation Forest) is used to identify outliers in monitoring data: the historical monitoring database (already compiled) is called, and a certain number of samples are randomly extracted from the historical monitoring data to complete subsampling. For each subsampling set, at each node, the algorithm randomly selects a feature and randomly selects a split value between the maximum and minimum values ​​of the feature value. The data is divided into a left subtree or a right subtree according to the split value, and recursively until the tree reaches a limited height, the number of samples in the node reaches a certain number, or the feature values ​​selected by all samples are the same. Repeat the above process until a specific number of isolated trees are constructed, and the collection is an isolation forest.

[0080]

[0081] Where: S(x, n) is the anomaly score of data point x, which ranges from [0, 1]; n is the number of samples; h(x) is the path length, which refers to the number of edges required for the sample to reach the node where the sample is isolated from the root node of the tree through the feature selection method in the isolated tree construction stage; E(h(x)) is the average path length, that is, the average path length of all trees in the forest for the sample; H() is the harmonic number, which can be approximated as ln()+0.5772156649; c(n) is the average path length of the tree, calculated as:

[0082]

[0083] In the anomaly identification of groundwater monitoring data, the number of trees is set to 100, the path length is calculated, and the anomaly score of each data point is given to determine the anomaly: if the anomaly scores of all data points involved in the test are around 0.5, then there are no anomalies in the corresponding data sequence; if the anomaly score of the data point is much less than 0.5, then the data point is a normal value; if the anomaly score of the data point is close to 1, it is determined to be an anomaly and marked. In this regard, in step 104, the denoised monitoring data is deeply analyzed according to the isolation forest to determine the abnormal data in the data, specifically including:

[0084] (f1) Randomly select Ψ sample points from the denoised monitoring data as a sample subset and put them into the root node of the tree.

[0085] (f2) Randomly specify a dimension and randomly generate a cutting point p in the current node data; the cutting point p is generated between the maximum and minimum values ​​of the specified dimension in the current node data.

[0086] (f3) A hyperplane is generated with the cutting point p, and then the data space of the current node is divided into two subspaces. The data with a cutting point less than p in the sample subset is placed in the left child node of the current node, and the data greater than or equal to p is placed in the right child node of the current node.

[0087] (f4) Let the left child node and the right child node be the current node data respectively, and return to the step "randomly specify a dimension and randomly generate a cutting point p in the current node data" until there is only one data in the child node or the child node has reached the specified height.

[0088] (f5) Return to step "randomly select Ψ sample points from the denoised monitoring data as a sample subset" until T isolated trees are generated to obtain an isolation forest.

[0089] (f6) For each data point in the denoised monitoring data, allow the data point to traverse each isolated tree and calculate the path length of the data point in the isolated tree.

[0090] (f7) Normalizing the path lengths of all data points in the denoised monitoring data to obtain an average path length.

[0091] (f8) Calculate the anomaly score of each data point based on the path length and average path length of the data point in the isolation tree.

[0092] (f9) Mark data points whose anomaly scores are higher than a preset score threshold as abnormal data.

[0093] In another exemplary embodiment of the present application, in step 105, the manual identification result of the abnormal data is obtained. For the manual identification process, through the above steps 101 to 104, the abnormal identification and noise removal of the groundwater monitoring data are preliminarily completed. The abnormal data marked by the computer is displayed through the human-computer interface, and is transferred to professional and technical personnel for manual identification and manual trimming of the abnormal data. Manual identification of abnormal data is an important and complex process, which depends on the analyst's professional knowledge, experience and judgment.

[0094] Manual identification should be carried out in conjunction with the operation and maintenance of field monitoring wells and monitoring equipment, clarify the source and acquisition method of the data, and analyze the multi-year annual dynamic changes of groundwater reflected in the original monitoring data of the monitoring wells one by one. For periods of deviation from the basic laws, the specific causes of the anomalies should be analyzed in conjunction with on-site water level calibration results, extreme precipitation events, and surrounding pumping (irrigation) and drainage events that affect groundwater dynamics.

[0095] ① If the marked data is consistent with the on-site water level calibration result, the marked data will be retained, otherwise the marked data will be discarded.

[0096] ②Through the trend analysis method, according to the time series curve shape and change trend of the abnormal data at the corresponding time point, determine whether it belongs to periodic fluctuations or emergencies. If the curve shape of the marked data is similar to the historical curve shape of the same period, or its change trend shows a significant positive correlation with the precipitation change trend in the same period, the marked data will be retained, otherwise the marked data will be removed.

[0097] ③Use the causal analysis method and combine it with the field operation and maintenance records to determine whether the data anomaly is caused by groundwater pumping (irrigation) and drainage events. If the change trend of the marked data is closely related to the groundwater pumping (irrigation) and drainage events in the same period, the marked data will be retained, otherwise the marked data will be removed.

[0098] ④ Based on the on-site water level calibration results and combined with the equipment maintenance records, determine whether the data drift is caused by equipment failure. If it is indeed caused by equipment failure, the marked data during the failure period can be shifted as a whole based on the manual calibration data to ensure that the marked data after the shift is consistent with the previous and subsequent time periods and the trend is consistent.

[0099] ⑤ According to the on-site water level calibration results and combined with the equipment maintenance records, determine whether the data gaps or other uncorrectable data anomalies are caused by equipment failure. Eliminate the uncorrectable abnormal data corresponding to the time node when the abnormal data appears, and re-record the manual calibration data after elimination, and re-record the probe storage data during the data gap period.

[0100] In another exemplary embodiment of the present application, different levels of data review are performed based on the above-mentioned compilation process of step 101 to step 106. With one year as the compilation cycle, a four-level review system of daily-monthly-quarterly-annual is established.

[0101] ① Daily monitoring: Monitor the operating status and data reception of the monitoring equipment on the previous day every day, and start the daily compilation process for the abnormal data marked in steps 101-104. The manual identification and compilation method is shown in the above manual identification process.

[0102] ② Monthly compilation: It includes two aspects. One is to confirm whether the basic information of the monitoring site in the current month, such as whether the measuring point elevation, ground elevation, longitude and latitude and other information have changed; the other is to execute steps 101-104 on the data received in the previous month on a monthly basis, and start the monthly compilation process for the marked abnormal data. The manual identification and compilation method can be found in the above-mentioned manual identification process.

[0103] ③ Quarterly verification: Conduct a quarterly review of the compiled monthly data, execute steps 101-104 on the data received in the previous quarter on a quarterly basis, start the quarterly compilation process for the marked abnormal data, and manually identify the compilation method using the manual identification process mentioned above.

[0104] ④Annual verification: Annual review of compiled data, including two aspects: first, final verification and confirmation of basic information of all monitoring stations in the current year, such as whether there are any changes in information such as measuring point elevation, ground elevation, longitude and latitude; second, compilation and processing of all received data in the current year in units of years, and analysis of trend and periodicity of annual water level data point by point in combination with multi-year water level, precipitation and other data, to determine its rationality and eliminate marked data that violate hydrogeological laws.

[0105] In another exemplary embodiment of the present application, after the compilation process based on the above steps 101 to 106, data verification can be further performed. Monitoring stations in different regions and typical regions are selected at a ratio of 10%, and while using automated monitoring equipment to carry out automatic monitoring, manual water level calibration is carried out throughout the year, and the compiled annual calibration data is verified using the manual calibration data.

[0106] ① Using real-time verification methods, during the acquisition of original monitoring data, the accuracy and completeness of the data are verified in real time, and feedback is given immediately to improve the accuracy and efficiency of data compilation.

[0107] ② Use multi-source data comparison methods to compare monitoring data obtained from different sources within the same time period to verify the accuracy and consistency of the compiled data.

[0108] ③Use the sampling verification method to randomly select a part of the samples from a large amount of data for verification to evaluate the accuracy and reliability of the overall data.

[0109] ④ Using statistical test methods, the t-test was used to determine whether there was a significant difference in the mean of the compiled data and the manually calibrated data, the F-test was used to determine whether the compiled data and the manually calibrated data were derived from the same monitoring site in terms of data performance, and the chi-square test was used to determine whether the difference between the compiled data and the manually calibrated data was significant.

[0110] In this application, a method for compiling and processing massive groundwater monitoring data is disclosed. This method is based on data consistency, data cleaning, anomaly identification and other methods, combined with the unique dynamic change trends and laws of groundwater monitoring data, analyzes the common problems that cause data anomalies, and innovatively uses a technical method that interacts with machine learning and manual identification. After multiple rounds of cleaning and compiling massive monitoring data, the verification and testing of the compiled data are completed. This method improves the efficiency of large-scale data compilation, reduces time and labor costs, and solves the technical difficulties of identifying and locating massive data anomalies. At present, this method has been applied to the compilation of groundwater monitoring data in the National Groundwater Monitoring Project, and has achieved good results.

[0111] The present application also provides an application scenario, which applies the above-mentioned compilation method for massive groundwater monitoring data. Specifically: The compilation method for massive groundwater monitoring data provided by the present application can be applied in the groundwater monitoring data compilation scenario of the national groundwater monitoring project, including data collection link, data compilation link and data display link. The data collection link is used to collect groundwater monitoring data, the data compilation link is used to compile and process the groundwater monitoring data, and the data display link is used to display the compiled data. The compilation method for massive groundwater monitoring data of the present application belongs to the data compilation link.

[0112] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the compiled massive groundwater monitoring data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a compilation method for massive groundwater monitoring data is implemented.

[0113] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0114] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0115] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0117] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0118] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0119] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for compiling massive groundwater monitoring data, characterized in that: The method for compiling massive groundwater monitoring data includes: Obtain groundwater monitoring data to be compiled; Cleaning the groundwater monitoring data to be compiled according to the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database, marking abnormal data, and obtaining cleaned monitoring data; the cleaned monitoring data is marked with abnormal data; According to the historical groundwater monitoring data stored in the groundwater monitoring database, the cleaned monitoring data is subjected to normality test and noise processing to obtain denoised monitoring data; the denoised monitoring data is marked with abnormal data; The local anomaly factor algorithm, TukeyBoxplot algorithm, STL time series decomposition algorithm and isolation forest are used to perform in-depth anomaly analysis on the denoised monitoring data to determine the abnormal data in the data; Obtaining manual identification results of the abnormal data; In the denoised monitoring data, abnormal data in the manual identification results are eliminated to obtain the compiled groundwater monitoring data.

2. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: The groundwater monitoring data to be compiled is cleaned according to the basic information of the groundwater monitoring site and the historical groundwater monitoring data stored in the groundwater monitoring database, and abnormal data is marked to obtain the cleaned monitoring data, specifically including: According to the basic information of the groundwater monitoring site, identifying and removing the data whose groundwater level is lower than the bottom elevation of the monitoring well in the groundwater monitoring data to be compiled; According to the basic information of the groundwater monitoring site, identifying and removing the data whose actual groundwater depth measured value exceeds the probe range of the monitoring equipment in the groundwater monitoring data to be compiled; According to the basic information of the groundwater monitoring site, identifying the data in which the groundwater level of the non-artesian monitoring wells in the groundwater monitoring data to be compiled is higher than the ground elevation, and marking it as abnormal data; According to the historical groundwater monitoring data stored in the groundwater monitoring database, identifying the data in the groundwater monitoring data to be compiled whose groundwater level is higher than the historical highest water level, and marking it as abnormal data; According to the historical groundwater monitoring data stored in the groundwater monitoring database, the data in which the groundwater level is lower than the historical lowest water level in the groundwater monitoring data to be compiled are identified and marked as abnormal data; the cleaned monitoring data refers to the data obtained after data elimination and abnormal data marking of the groundwater monitoring data to be compiled.

3. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: According to the historical groundwater monitoring data stored in the groundwater monitoring database, the cleaned monitoring data is subjected to normality test and noise processing to obtain the denoised monitoring data, specifically including: Performing a normality test on historical groundwater monitoring data in the same period as the groundwater monitoring data to be compiled in the groundwater monitoring database, identifying data outside a preset normality test range in the groundwater monitoring data to be compiled, and marking them as abnormal data; The frequency components of the historical groundwater monitoring data stored in the groundwater monitoring database are obtained by using fast Fourier transform; Integrating the cleaned monitoring data with the historical groundwater monitoring data to obtain an integrated data sequence; Acquiring the time-frequency spectrum of the integrated data sequence by wavelet transform; identifying high-frequency data other than the frequency components in the integrated data sequence according to the time-frequency spectrum; Determining a time point at which the high-frequency data appears in the time domain; If the time point of the high-frequency data appears both in the time domain of the historical monitoring data and in the time domain of the cleaned monitoring data, the corresponding high-frequency data is not processed; If the time point of the high-frequency data only appears in the time domain of the cleaned monitoring data, the corresponding high-frequency data is marked as noise data; The marked noise data is removed by using a moving average method to obtain the denoised monitoring data.

4. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: The local anomaly factor algorithm is used to perform in-depth anomaly analysis on the denoised monitoring data to determine the abnormal data in the data, specifically including: For each data point in the denoised monitoring data, determining k nearest neighboring data points of the data point; Calculate the distance between the data point and each corresponding nearest neighbor data point to obtain a first distance; Calculate the distance between any two nearest neighbor data points among the k nearest neighbor data points to obtain the second distance; Calculate the local anomaly factor of the data point according to the first distance and the second distance; The data points whose local abnormality factors exceed a preset threshold are marked as abnormal data.

5. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: The TukeyBoxplot algorithm is used to perform in-depth anomaly analysis on the denoised monitoring data to determine the abnormal data in the data, specifically including: Dividing the denoised monitoring data into multiple monitoring data sets according to a preset time window; For each monitoring data set, determining the median of the monitoring data set; Determine the upper quartile and the lower quartile of the monitoring data set according to the median; Determine an interquartile range according to the upper quartile and the lower quartile; Determine the upper limit and lower limit of Boxplot according to the interquartile range; The data above the upper limit of the Boxplot and the data below the lower limit of the Boxplot in the denoised monitoring data are marked as abnormal data.

6. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: The STL time series decomposition algorithm is used to perform in-depth anomaly analysis on the denoised monitoring data to determine the abnormal data in the data, specifically including: Using the STL time series decomposition algorithm, the denoised monitoring data is decomposed into a seasonal cycle series, a trend series and a residual series; According to the median of the trend sequence and the residual sequence, a weight value is assigned to each monitoring time point in the denoised monitoring data; The data whose weight value at the monitoring time point in the denoised monitoring data is lower than the preset weight value is marked as abnormal data.

7. The method for compiling massive groundwater monitoring data according to claim 1, characterized in that: The isolation forest is used to perform deep anomaly analysis on the denoised monitoring data to determine the abnormal data in the data, specifically including: Randomly select Ψ sample points from the denoised monitoring data as a sample subset and put them into the root node of the tree; Randomly specify a dimension and randomly generate a cutting point p in the current node data; the cutting point p is generated between the maximum and minimum values ​​of the specified dimension in the current node data; A hyperplane is generated with the cutting point p, and then the data space of the current node is divided into two subspaces, the data with a cutting point less than p is placed in the left child node of the current node, and the data greater than or equal to p is placed in the right child node of the current node; Let the left child node and the right child node be the current node data respectively, and return to step "randomly specify a dimension and randomly generate a cutting point p in the current node data" until there is only one data in the child node or the child node has reached the specified height; Return to step "randomly select Ψ sample points from the denoised monitoring data as a sample subset" until T isolated trees are generated to obtain an isolation forest; For each data point in the denoised monitoring data, let the data point traverse each isolated tree and calculate the path length of the data point in the isolated tree; Normalizing the path lengths of all data points in the denoised monitoring data to obtain an average path length; Calculate the anomaly score of each data point based on the path length and average path length of the data point in the isolation tree; Data points with anomaly scores higher than a preset score threshold are marked as abnormal data.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for compiling massive groundwater monitoring data as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for compiling massive groundwater monitoring data described in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for compiling massive groundwater monitoring data described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Geological test data abnormal value detection method, system, equipment and medium

    CN120804640A