Hydrological element data cleaning processing method and system

By combining the slope difference and mean filtering algorithms, abnormal fluctuations in hydrological data are identified and cleaned, which improves the accuracy and reliability of the data and solves the problem of the impact of abnormal fluctuations in hydrological data on analysis.

CN120804530AActive Publication Date: 2025-10-17XIAN ERJI ENVIRONMENTAL PROTECTION TECH CO LTD

Patent Information

Application Number
CN202511308269.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify and process abnormal fluctuations or mutations in hydrological data, resulting in decreased data analysis accuracy and erroneous decision-making.

Method used

The slope difference algorithm is used for primary filtering, combined with the mean filtering algorithm for secondary cleaning. Abnormal data are identified through slope difference and water level difference, and fine calibration is performed using the mean characteristics of hydrological data.

Benefits of technology

It improves the accuracy of hydrological data, reduces the impact of erroneous data on analysis results, and provides a reliable data basis for water resources management and water conservancy project decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804530A_ABST
    Figure CN120804530A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hydrological monitoring, in particular to a hydrological element data cleaning processing method and system, and the method comprises the steps: carrying out the initial filtering of hydrological data through employing a slope difference algorithm, and obtaining the initial cleaning data; acquiring fluctuation parameters of the preliminary cleaning data; judging whether the preliminary cleaning data is fluctuation data or not according to the fluctuation parameters; when the preliminary cleaning data is not fluctuation data, cleaning the preliminary cleaning data by adopting a first mean filtering algorithm to obtain final cleaning data; when the preliminary cleaning data is fluctuation data, detecting whether the fluctuation data is normal fluctuation data or not; if the fluctuation data is normal fluctuation data, cleaning the preliminary cleaning data by adopting a second mean filtering algorithm to obtain final cleaning data; and if the fluctuation data is abnormal fluctuation data, determining the preliminary cleaning data as final cleaning data. According to the method, the accuracy of the data is comprehensively improved, and the influence of wrong data on the accuracy of an analysis result is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hydrological monitoring, in particular to a hydrological element data cleaning processing method and system. BACKGROUND

[0002] In the field of hydrological monitoring, the quality of data is crucial for flood and drought disaster prevention, water resource management, and environmental protection decision-making. With the development of technology in the water industry and the increasing demand for water resource management, the accuracy and reliability of hydrological data, which is an important basis for water resource management, become particularly critical. These data are not only used for daily water resource scheduling, flood and drought warning, but also widely used in long-term water resource planning and environmental impact assessment.

[0003] Due to equipment failure, environmental interference, network transmission, etc., the data recorded by the hydrological station may have abnormal fluctuations or mutations.

[0004] However, these hydrological data with abnormal fluctuations or mutations not only affect the accuracy of data analysis, but also mislead the decision-makers' judgment, leading to inappropriate problem handling measures.

[0005] Due to the unique properties of hydrological data, traditional data cleaning methods may not be able to effectively identify and process abnormal values in such data, such as unrealistic data fluctuations or mutations, which may lead to incorrect data affecting the accuracy of analysis results. SUMMARY

[0006] In order to solve the technical problem that abnormal fluctuations or mutations of hydrological data affect the accuracy of data analysis, the purpose of the present application is to provide a hydrological element data cleaning processing method and system, and the technical solution adopted is as follows: The first aspect of the present disclosure provides a hydrological element data cleaning processing method, which comprises: using a slope difference algorithm to initially filter the hydrological data and obtain preliminary cleaning data; obtaining the fluctuation parameters of the preliminary cleaning data; determining whether the preliminary cleaning data is fluctuation data according to the fluctuation parameters; when the preliminary cleaning data is not the fluctuation data, using a first mean filter algorithm to clean the preliminary cleaning data to obtain final cleaning data; when the preliminary cleaning data is the fluctuation data, detecting whether the fluctuation data is normal fluctuation data; if the fluctuation data is the normal fluctuation data, using a second mean filter algorithm to clean the preliminary cleaning data to obtain the final cleaning data; If the fluctuation data is abnormal fluctuation data, the preliminary cleaning data is determined as final cleaning data.

[0007] In one embodiment, after determining the final cleaning data, the method further comprises: storing the final cleaning data.

[0008] In one embodiment, the preliminary filtering of the hydrological data using the slope difference algorithm to obtain the preliminary cleaning data comprises: obtaining the change amount of two adjacent hydrological data in time sequence in the hydrological data; obtaining the time interval of the two adjacent hydrological data; obtaining the slope of the two adjacent hydrological data according to the change amount and the corresponding time interval of the two adjacent hydrological data; obtaining a target slope of the two adjacent hydrological data, the target slope being greater than a preset slope threshold, and the hydrological data corresponding to the target slope being a preliminary abnormal point; obtaining the water level difference of two adjacent preliminary abnormal points; obtaining a target abnormal point, the water level difference of the target abnormal point being greater than a preset change amount threshold, and the target abnormal point being real abnormal data; filtering the real abnormal data from the hydrological data to obtain the preliminary cleaning data.

[0009] In one embodiment, the obtaining of the fluctuation parameter of the preliminary cleaning data comprises: obtaining the hydrological data difference of all two adjacent hydrological data in time sequence in the preliminary cleaning data; obtaining the fluctuation parameter of the preliminary cleaning data according to the number of all the preliminary cleaning data and the sum of all the hydrological data differences.

[0010] In one embodiment, the determining of whether the preliminary cleaning data is fluctuation data according to the fluctuation parameter comprises: detecting whether the fluctuation parameter is greater than a preset fluctuation threshold; if the fluctuation parameter is greater than the preset fluctuation threshold, determining that the preliminary cleaning data is the fluctuation data; if the fluctuation parameter is not greater than the preset fluctuation threshold, determining that the preliminary cleaning data is not the fluctuation data.

[0011] In one embodiment, the detecting of whether the fluctuation data is normal fluctuation data comprises: obtaining the peak value of all peaks in the fluctuation data; obtaining the sharpness corresponding to the peak according to the peak value; determining the fluctuation data corresponding to the wave peak with the sharp degree greater than the preset sharp threshold as the abnormal fluctuation data; determining the fluctuation data corresponding to the wave peak with the sharp degree not greater than the preset sharp threshold as the normal fluctuation data.

[0012] In one embodiment, the sharp degree corresponding to the wave peak is obtained according to the peak value, comprising: performing the following steps on each wave peak: obtaining a preset number of other wave peaks adjacent to the current wave peak; obtaining the peak center of the current wave peak and the other wave peaks; obtaining the sharp degree of the current wave peak according to the peak value of the peak center and the peak value of the current wave peak.

[0013] In one embodiment, the first mean filtering algorithm is used to clean the preliminary cleaning data, comprising: determining the abnormal data in the preliminary cleaning data; filtering out the abnormal data from the preliminary cleaning data to obtain the final cleaning data; The determination of the abnormal data in the preliminary cleaning data comprises: performing the following steps on each cleaning data in the preliminary cleaning data: obtaining the occurrence time of the current cleaning data; obtaining a preset number of historical hydrological data filtered by the slope difference algorithm at the same time as the occurrence time; obtaining the mean value of the historical hydrological data; obtaining the mean value feature difference corresponding to the current cleaning data according to the mean value and the current cleaning data; If the mean value feature difference is greater than a preset mean value threshold, the current cleaning data is determined as the abnormal data.

[0014] In one embodiment, the preset mean value threshold corresponding to the first mean filtering algorithm is less than the preset mean value threshold corresponding to the second mean filtering algorithm.

[0015] The second aspect of the present disclosure provides a hydrological element data cleaning processing system, comprising: a collection module for collecting hydrological data, the hydrological data comprising water level data; a primary filtering module for performing primary filtering on the hydrological data using a slope difference algorithm to obtain preliminary cleaning data; a secondary filtering module, configured to: acquire a fluctuation parameter of the preliminary cleaning data; determine whether the preliminary cleaning data is fluctuation data according to the fluctuation parameter; when the preliminary cleaning data is not the fluctuation data, clean the preliminary cleaning data by using a first mean value filtering algorithm to obtain final cleaning data; when the preliminary cleaning data is the fluctuation data, detect whether the fluctuation data is normal fluctuation data; when the fluctuation data is the normal fluctuation data, clean the preliminary cleaning data by using a second mean value filtering algorithm to obtain the final cleaning data; when the fluctuation data is abnormal fluctuation data, determine that the preliminary cleaning data is the final cleaning data; a storage module, configured to store the final cleaning data.

[0016] In one embodiment, the primary filtering module is specifically configured to: acquire a change amount of two hydrological data adjacent in time sequence in the hydrological data; acquire a time interval of the two adjacent hydrological data; acquire a slope of the two adjacent hydrological data according to the change amount and the corresponding time interval of the two adjacent hydrological data; acquire a target slope greater than a preset slope threshold from the slope of the two adjacent hydrological data, the hydrological data corresponding to the target slope being a preliminary abnormal point; acquire a water level difference of two adjacent preliminary abnormal points; acquire a target abnormal point greater than a preset change amount threshold from the water level difference, the target abnormal point being the real abnormal data.

[0017] In one embodiment, the secondary filtering module is specifically configured to: acquire a hydrological data difference of two hydrological data adjacent in time sequence in the preliminary cleaning data; acquire a fluctuation parameter of the preliminary cleaning data according to a number of all the preliminary cleaning data and a sum of all the hydrological data differences.

[0018] In one embodiment, the secondary filtering module is specifically configured to: detect whether the fluctuation parameter is greater than a preset fluctuation threshold; when the fluctuation parameter is greater than the preset fluctuation threshold, determine that the preliminary cleaning data is the fluctuation data; when the fluctuation parameter is not greater than the preset fluctuation threshold, determine that the preliminary cleaning data is not the fluctuation data.

[0019] In one embodiment, the secondary filtering module is specifically configured to: acquire a peak value of all peaks in the fluctuation data; According to the peak, the sharpness corresponding to the wave peak is obtained; The wave fluctuation data corresponding to the wave peak with the sharpness greater than the preset sharpness threshold is determined as the abnormal wave fluctuation data; The wave fluctuation data corresponding to the wave peak with the sharpness not greater than the preset sharpness threshold is determined as the normal wave fluctuation data.

[0020] In one embodiment, the secondary filtering module is specifically used for: The following steps are performed on each wave peak: A preset number of other wave peaks adjacent to the current wave peak are obtained; The center of the current wave peak and the wave peak in the other wave peaks are obtained; According to the peak value of the wave peak center and the peak value of the current wave peak, the sharpness of the current wave peak is obtained.

[0021] In one embodiment, the secondary filtering module is specifically used for: Abnormal data in the preliminary cleaning data is determined; The abnormal data is filtered out from the preliminary cleaning data to obtain the final cleaning data; The determination of the abnormal data in the preliminary cleaning data includes: The following steps are performed on each cleaning data in the preliminary cleaning data: The occurrence time of the current cleaning data is obtained; A preset number of historical hydrological data filtered by the slope difference value algorithm and the same occurrence time are obtained; The mean value of the historical hydrological data is obtained; According to the mean value and the current cleaning data, the mean value feature difference corresponding to the current cleaning data is obtained; If the mean value feature difference is greater than a preset mean value threshold, the current cleaning data is determined as the abnormal data.

[0022] In one embodiment, the preset mean value threshold corresponding to the first mean value filtering algorithm is less than the preset mean value threshold corresponding to the second mean value filtering algorithm.

[0023] The present application has the following advantages: According to the characteristics of hydrological data in the present disclosure, the slope difference algorithm is used for primary filtering, and the slope difference of the water level data before and after the change can be used to acutely understand the data fluctuation and mutation deviating from the normal change track, and the abnormal data is accurately locked and preliminarily removed from the original data, so as to curb the interference of the error data from the source. Then, the mean filtering algorithm is used for secondary fine cleaning. The mean filtering algorithm takes the mean value of the water level data as the reference to comprehensively calibrate the primary cleaned water level data, can smooth the slight data noise and residual deviation, comprehensively improves the accuracy of the data, greatly reduces the influence of the error data on the accuracy of the analysis result, makes the analysis conclusion based on the cleaned data more reliable, and provides a solid data basis for hydrological research and water conservancy engineering decision. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0025] Figure 1 The flowchart of the hydrological element data cleaning processing method provided by an embodiment of the present application Figure 1 Figure 2 The flowchart of the hydrological element data cleaning processing method provided by an embodiment of the present application Figure 2 Figure 3 The slope provided by an embodiment of the present application Figure 1 Figure 4 The slope provided by an embodiment of the present application Figure 2 Figure 5 The flowchart of the hydrological element data cleaning processing method provided by an embodiment of the present application Figure 3 Figure 6 The fluctuation parameter and the water level fluctuation provided by an embodiment of the present application Figure 7 The schematic diagram of a wave crest in the curve corresponding to the fluctuation data provided by an embodiment of the present application Figure 8 The structural schematic diagram of the hydrological element data cleaning processing system provided by an embodiment of the present application DETAILED DESCRIPTION

[0026] ​​​​​In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined object of the application, the specific implementation, structure, features and effects of the hydrological element data cleaning processing method and system according to the present application are described in detail below in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0028] Abnormal fluctuations or mutations of hydrological data not only affect the accuracy of data analysis, but also mislead the judgment of decision makers, leading to inappropriate problem handling measures.

[0029] Most existing data cleaning processing tools and methods do not fully consider the unique properties of hydrological data, such as river water level and flow, reservoir water level and storage capacity, etc. These parameters have specific time series characteristics and physical laws, and can form certain historical characteristics combined with historical data. The traditional data cleaning processing method may not be able to effectively identify and process abnormal values in such data, such as unrealistic data fluctuations or mutations. This may lead to errors in data affecting the accuracy of analysis results, while causing bias in stored historical data.

[0030] To solve the above technical problems, the present disclosure provides a hydrological element data cleaning processing method and system.

[0031] The specific scheme of the hydrological element data cleaning processing method and system provided by the present application is described in detail below in combination with the drawings.

[0032] Please refer to Figure 1 which shows the flow chart of the hydrological element data cleaning processing method provided by one embodiment of the present application, which includes the following steps S101-S104: Step S101, collecting hydrological data.

[0033] In this step, the relevant hydrological data recorded by the monitoring station of the hydrological monitoring station is obtained. For example: water level data.

[0034] The hydrological data of reservoirs, rivers, etc. needs to be obtained by monitoring stations. Therefore, through the smart water system of the municipal water bureau, the original data of the city's hydrological stations is obtained, including rainfall, surface water, groundwater, water environment, etc.

[0035] In this example, the water level data monitored by the river station of a certain city is taken as an example to describe the subsequent data cleaning processing method.

[0036] Because the river station site measures and collects water level data without interruption, the above river water level data is a streaming data, so the subsequent data cleaning processing method is above the river water level time series data.

[0037] Due to the demand for water resources management, high-quality hydrological data is needed for water resources scheduling, flood control and drought warning, and long-term water resources planning and scheduling. Therefore, before using these data, the collected data needs to be cleaned first, and the error data such as large fluctuations and data mutations that do not conform to the actual situation are removed, and the real data is retained, so as to improve the accuracy of subsequent data analysis.

[0038] Regarding data cleaning, there are many mature technologies, but most of the existing data cleaning tools and methods do not consider the unique properties of hydrological data, such as river water level and flow parameters. These data have specific time series characteristics and interrelated physical laws, and can form certain historical characteristics combined with historical data.

[0039] Therefore, next, the hydrological data cleaning processing method proposed by the present application is introduced in detail.

[0040] Step S102, using a slope difference algorithm to perform initial filtering on the hydrological data to obtain preliminary cleaning data.

[0041] Specifically, as shown in Figure 2 The step S102 includes the following sub-steps S1021-S1027: S1021, obtaining the change amount of two adjacent hydrological data in time sequence.

[0042] S1022, obtaining the time interval of the two adjacent hydrological data.

[0043] S1023, obtaining the slope of the two adjacent hydrological data according to the change amount and the corresponding time interval of the two adjacent hydrological data.

[0044] S1024, obtaining the target slope of the two adjacent hydrological data whose slope is greater than the preset slope threshold, and the hydrological data corresponding to the target slope is the preliminary abnormal point.

[0045] S1025, obtaining the water level difference of the two preliminary abnormal points.

[0046] S1026, obtaining the target abnormal point whose water level difference is greater than the preset change amount threshold, and the target abnormal point is the real abnormal data.

[0047] S1027, filtering the real abnormal data from the hydrological data to obtain the preliminary cleaning data.

[0048] Specifically, taking the water level data of the river detection station as an example, the two water level data adjacent in time sequence have strong correlation, and the change rule therebetween can well reflect the characteristics of the water level data in this time period. Therefore, according to the two water level data adjacent in time sequence and time as the basis, the slope of the two adjacent water level data is calculated: In the formula, is the slope of the two adjacent water level data, is the change amount of the two adjacent water level data in time sequence, is the time interval of the two adjacent water level data in time sequence.

[0049] The calculated slope can represent the change trend of the two adjacent water level data.

[0050] Under normal circumstances, the river water level in a short time will not have a large change, so should also be at a small level. Therefore, in the present disclosure, a preset slope threshold can be set, and for the two adjacent water level data , it is considered that the water level data is abnormal, and it is determined as a preliminary abnormal point.

[0051] The above method determines whether the data is abnormal by calculating the slope of the water level-time between the adjacent points in time sequence. In this way, the hydrological data with a large slope can generally be considered as abnormal data, however, in reality, there are the following situations: if the two hydrological data adjacent in time sequence have data loss, the water level change between the two data may be large, but the calculated slope may not be high; if the time interval between the two collected data is small, although the difference of the water level data may be small, the calculated slope will be large. As shown in Figure 3 and Figure 4 , the horizontal coordinate T represents time, and the vertical coordinate WL represents water level. Figure 3 and Figure 4 .

[0052] The water level difference and the slope need to be treated together to filter the data to be processed. The specific method is as follows: First, for a group of water level data to be processed, the slope between each point and its adjacent water level data is calculated by the above method, and the water level data with greater than the preset slope threshold is recorded as a preliminary abnormal point.

[0053] Next, the water level difference value of the preliminary abnormal point filtered in S1024 is further calculated , the water level difference value The greater the preliminary abnormal point, the greater the likelihood of being a real abnormal data; and the smaller the water level difference value , the more likely the preliminary abnormal point is a normal data that is misjudged as abnormal data.

[0054] The water level difference value A preset change threshold is set , if the water level difference value of the preliminary abnormal point calculated in the previous step is less than the preset change threshold, it is considered that the preliminary abnormal point is real abnormal data, and it is filtered out. The preset change threshold , needs to be selected by the experimenter according to experience and the actual situation of the local river, and the smaller the preset change threshold, the more sensitive the algorithm is to the change of the river water level.

[0055] In this way, the slope difference filtering of the hydrological data is completed. The algorithm mainly uses data slope and auxiliary difference value, which can ensure the identification of abnormal data while avoiding incorrect processing of correct data.

[0056] In step S102, the current real-time collected hydrological data is directly analyzed and calculated to judge and clean the abnormal data. Next, in this step, the real-time collected hydrological data is taken as the object, and the preliminary cleaning data obtained after the slope difference cleaning is combined to continue analyzing the hydrological data, find and remove abnormal data, and further complete the cleaning of the hydrological data.

[0057] Step S103, the mean filtering algorithm is used to filter the preliminary cleaning data again to obtain the final cleaning data; In step S102, the current real-time collected hydrological data is directly analyzed and calculated to judge and clean the abnormal data. Next, in this step, the real-time collected hydrological data is taken as the object, and the preliminary cleaning data obtained after the slope difference cleaning is combined to continue analyzing the hydrological data, find and remove abnormal data, and further complete the cleaning of the hydrological data.

[0058] Specifically, step S103 can be implemented by the following way: A1, determine the abnormal data in the preliminary cleaning data.

[0059] A2, filter out the abnormal data from the preliminary cleaning data to obtain the final cleaning data.

[0060] Among them, determining the abnormal data in the preliminary cleaning data includes executing the following steps A11-A15 on each cleaning data in the preliminary cleaning data: A11, obtain the occurrence time of the current cleaning data.

[0061] A12, obtain a preset number of historical water level data filtered by the slope difference value algorithm at the same occurrence time.

[0062] A13, obtain the mean value of the historical water level data.

[0063] A14, obtain the mean value feature gap corresponding to the current cleaning data according to the mean value and the current cleaning data.

[0064] A15, if the mean value feature gap is greater than a preset mean value threshold, determine that the current cleaning data is abnormal data.

[0065] Still taking the water level data of the river site as an example, the data processed by the above algorithm is a certain amount of correct data that can be referenced, so the mean value of the data calculated using a certain amount of historical correct data of the same period can reflect the river water level feature of the site to a certain extent, and further, it can be considered that if the gap between the real-time streaming water level data collected next time and the mean value of the current calculated water level data is very small, then the next water level data is correct; when the gap between the real-time streaming water level data collected next time and the mean value of the current calculated water level data is very large, it is considered that the water level data is incorrect.

[0066] According to the above analysis, the specific calculation formula of the mean value filtering is: ; Wherein, represents the mean value feature gap, represents the mean value of the water level data stream filtered by the slope difference value at the same period; represents the water level data collected in real time at the current time.

[0067] The formula uses the absolute value of the difference between the mean value of the water level data filtered by the slope difference value at the same period and the water level data collected in real time at the current time to represent the mean value feature gap used to measure the difference between the current calculated water level data and the mean value of the water level data at the same period , the larger the value, the more likely it is that the current calculated water level data is abnormal data.

[0068] Here, the mean value feature gap can also be set to a preset mean value threshold, and the water level data exceeding the preset mean value threshold is considered as abnormal data, and then the abnormal data obtained by the mean value filtering algorithm is filtered out again from the preliminary cleaning data to obtain the final cleaning data.

[0069] Step S104, store the final cleaning data.

[0070] After the data cleaning process of the above steps is completed, the next step is to store the cleaned effective data in the database in a high-concurrent and orderly manner. Each piece of data is marked with a data state, and stored in the corresponding database table according to the monitoring site to which it belongs, the data type (such as river water level, flow, reservoir water level and capacity, etc.), and other related attributes.

[0071] For data information identified as abnormal in the cleaning process, it is extracted and marked for storage in the abnormal data table. The time, type and corresponding hydrological elements of the data are recorded in the abnormal data table, and the reason why the current data is identified as abnormal needs to be recorded. This is important for subsequent analysis of abnormal reasons, problem troubleshooting and algorithm improvement, and can also serve as a historical reference to assist in more accurate identification of abnormal situations in the future.

[0072] In view of the characteristics of hydrological data in the present disclosure, the slope difference algorithm is used for preliminary filtering. According to the slope difference of the change before and after the water level data, the data fluctuation and mutation deviating from the normal change track can be intuitively perceived, and the abnormal data can be accurately locked and preliminarily excluded from the original data, thereby suppressing the interference of error data from the source. Then, the mean filtering algorithm is used for secondary fine cleaning. The mean filtering algorithm takes the mean value of the water level data as the reference to comprehensively calibrate the water level data after preliminary cleaning, which can smooth out the subtle data noise and residual deviation, and improve the accuracy of the data in all directions. The influence of error data on the accuracy of the analysis result is greatly reduced, and the analysis conclusion based on the cleaned data is more reliable, providing a solid data foundation for hydrological research and water conservancy engineering decision-making.

[0073] In the above S102 and S103 steps, two cleaning methods for hydrological data, slope difference filtering and mean filtering algorithm, are proposed, and the input data used by the mean filtering algorithm is obtained by slope difference filtering. It can be found that the cleaning effect of the two algorithms on hydrological data is determined by their respective thresholds, and the core idea of the above-mentioned mean filtering algorithm is that the water level data showing stability (the closer to the mean value, the better) is normal data, and the data showing large fluctuations (the data deviating from the mean value) is abnormal data. However, taking the river water level data as an example, it may show a large range of fluctuations in the near future (here, the fluctuation refers to the correct and normal data fluctuation, such as the reasonable rise and fall of water level caused by heavy rain, rather than abnormal data fluctuation). At this time, the use of the mean value of the same period water level data to measure the stability of the data does not conform to the actual situation. Therefore, when combining the above two algorithms, the development trend of the input hydrological data needs to be analyzed first, and then the filtering effect of the mean filtering algorithm needs to be dynamically enhanced or weakened, so that the abnormal data can be cleaned under the premise that the normal data is not incorrectly processed.

[0074] Therefore, as Figure 5 shown, the above step S103 uses a mean filter algorithm to filter the preliminary cleaned data again to obtain the final cleaned data, including the following sub-steps S1031-S1036: S1031, obtaining the fluctuation parameter of the preliminary cleaned data.

[0075] In one embodiment, the fluctuation parameter of the preliminary cleaned data is obtained, including the following sub-steps S10311-S10312: S10311, obtaining the hydrological data difference value of all time-series adjacent two hydrological data in the preliminary cleaned data; S10312, obtaining the fluctuation parameter of the preliminary cleaned data according to the fluctuation base, the number of all preliminary cleaned data and the sum of all hydrological data difference values.

[0076] First, the volatility of the input hydrological data needs to be judged. In the judgment, as Figure 6 shown, if the hydrological data fluctuates greatly, the fluctuation parameter is large, and if the hydrological data is relatively stable, the fluctuation parameter is low.

[0077] Here, the river water level data is taken as an example, and the fluctuation parameter calculation method of the water level data in time series is as follows: Among them, represents the fluctuation parameter of the input water level data, which can measure the overall volatility of the data; represents the water level data difference value of the time-series adjacent two water level data; is the fluctuation base, which represents the water level data fluctuation noise when the water level data is in a normal state, i.e. the river water regime is in a stable state; is the number of input preliminary cleaned data.

[0078] S1032, judging whether the preliminary cleaned data is fluctuation data according to the fluctuation parameter, executing step S1033 when the preliminary cleaned data is not fluctuation data, and executing step S1034 when the preliminary cleaned data is fluctuation data.

[0079] Specifically, if the fluctuation parameter is greater than the preset fluctuation threshold, it is determined that the preliminary cleaned data is fluctuation data; if the fluctuation parameter is not greater than the preset fluctuation threshold, it is determined that the preliminary cleaned data is not fluctuation data.

[0080] The water level data in the stable period also has fluctuation phenomenon of rising and falling. Therefore, the data fluctuation base of the water level data in the stable period can be obtained , that is, the fluctuation range of the water level data is less than At this time, the water level data is considered to be "stable". Then, the difference between the current input adjacent two water level data and the fluctuation noise of the data in the above stable period is determined. Since the difference is a very small value itself, the greater the above difference, the less the similarity between the water level data at this time and the past stable period, that is, the greater the fluctuation of the water level data at this time.

[0081] The preset fluctuation threshold value here can be 1. For the adjacent two water level data , it is considered to be fluctuation data rather than stable data.

[0082] S1033, when the preliminary cleaned data is not fluctuation data, the first mean filtering algorithm is used to clean the preliminary cleaned data to obtain the final cleaned data.

[0083] In an embodiment, step S1033 uses the first mean filtering algorithm to clean the preliminary cleaned data, including the following sub-steps B1-B2: B1, determine the abnormal data in the preliminary cleaned data.

[0084] B2, filter out the abnormal data from the preliminary cleaned data to obtain the final cleaned data.

[0085] Among them, determining the abnormal data in the preliminary cleaned data includes executing the following steps B11-B15 on each cleaned data in the preliminary cleaned data: B11, obtain the occurrence time of the current cleaned data.

[0086] B12, obtain a preset number of historical hydrological data filtered by the slope difference value algorithm with the same occurrence time as the occurrence time.

[0087] B13, obtain the mean value of the historical hydrological data.

[0088] B14, according to the mean value and the current cleaned data, obtain the mean value feature difference corresponding to the current cleaned data.

[0089] B15, if the mean value feature difference is greater than a preset mean value threshold, determine that the current cleaned data is abnormal data.

[0090] S1034, when the preliminary cleaned data is fluctuation data, it is detected whether the fluctuation data is normal fluctuation data. If the fluctuation data is normal fluctuation data, step S1035 is executed. If the fluctuation data is not normal fluctuation data, step S1036 is executed.

[0091] S1035, if the fluctuation data is normal fluctuation data, using a second mean filtering algorithm to clean the preliminary cleaning data to obtain the final cleaning data.

[0092] The implementation of the second mean filtering algorithm is similar to the mean filtering algorithm in step S103 of the above embodiment, and will not be described here. It is worth noting that the preset mean threshold of the first mean filtering algorithm is less than the preset mean threshold of the second mean filtering algorithm.

[0093] S1036, if the fluctuation data is abnormal fluctuation data, determining the preliminary cleaning data as the final cleaning data.

[0094] Through step S1031, the fluctuation parameter that can measure the fluctuation state of the current water level data is obtained. Taking the fluctuation data as an object for further analysis. The fluctuation parameter only judges the macro volatility of the data, so it is necessary to judge whether the group of fluctuation data is normal fluctuation data or abnormal fluctuation data.

[0095] Taking the water level data collected at a fixed detection point on the river as an example, the river width at a certain point on the river is fixed, that is, the maximum runoff of the river at the point is determined. If it rains at this point, the fluctuation curve of the water level from rising to falling should be similar to the fluctuation of the water level data after the rain in the historical data. Further, there are two groups of water level data after the rain. If the water volume of the river after the rain does not exceed the maximum runoff, the curves of the two groups of water level data should show that only the amplitude of fluctuation is different, and the trend and shape are similar. According to common sense, the river water level after the rain is a clear and uniform trend, while abnormal data caused by collection errors, sensor failures and other factors will show as discrete points in the whole data, and it is almost impossible to find similar data in the historical data.

[0096] In one embodiment, step S1034 of detecting whether the fluctuation data is normal fluctuation data includes the following sub-steps S10341-S10344: S10341, obtaining the peak value of all peaks in the fluctuation data.

[0097] S10342, obtaining the sharpness corresponding to the corresponding peak of the peak value.

[0098] Specifically, obtaining the sharpness corresponding to the corresponding peak of the peak value includes executing the following sub-steps C1-C3 for each peak: C1, obtaining the peak value of a preset number of other peaks adjacent to the current peak.

[0099] C2, obtaining the peak center of the current peak and other peaks.

[0100] C3, the sharpness of the peak value of the current wave crest is obtained according to the peak value of the wave crest center and the peak value of the current wave crest.

[0101] After obtaining the peak values of all wave crests in the fluctuation data, taking one wave crest in the current collected fluctuation data as an example, if the wave crest is sharper, the point of the peak value of the wave crest is more likely to be abnormal data.

[0102] Therefore, the sharpness calculation method of the peak value of each wave crest in the fluctuation data is: Wherein, represents the sharpness of the peak value of each wave crest in the fluctuation data, represents the current peak value in the curve corresponding to the fluctuation data, is all peak values in the curve corresponding to the fluctuation data, represents a wave crest adjacent to the current peak value, represents the peak value of the wave crest . For a certain peak value , represents all discrete regions contained in the current peak value, the wave crests in the discrete regions are other wave crests adjacent to the current peak value in a preset number, and is the position of the wave crest center. As shown in Figure 7 , it shows a wave crest in the curve corresponding to the fluctuation data.

[0103] If the values on both sides of the peak value of a wave crest are far away from the peak value, it means that the wave crest is narrower, that is, the peak value is sharper. represents that for the peak value , the sum of squares of differences between all on both sides of the peak value and the wave crest center is calculated, and then multiplied by , that is, all peak values in the curve corresponding to the abnormal data.

[0104] S10343, determining that the fluctuation data corresponding to the wave crest with a sharpness greater than a preset sharpness threshold is abnormal fluctuation data.

[0105] S10344, determining that the fluctuation data corresponding to the wave crest with a sharpness not greater than a preset sharpness threshold is normal fluctuation data.

[0106] The greater the value is, the greater the difference between the values in the range on both sides of the peak value of a wave crest in the curve corresponding to the fluctuation data and the wave crest is, which means that the peak is steeper, and then the water level peak value data is more likely to be abnormal data.

[0107] Thus, the abnormality of the extreme points in the collected fluctuation data is obtained by setting a preset sharp threshold For the peak of the fluctuation data is considered to be abnormal fluctuation data rather than normal fluctuation data.

[0108] The calculation method of the sharp degree of the trough is the same as that of the peak. Specifically, the detection of whether the fluctuation data is normal fluctuation data in step S1034 includes the following sub-steps D1-D4: D1, obtaining the trough value of all troughs in the fluctuation data.

[0109] D2, obtaining the sharp degree corresponding to the trough according to the trough value.

[0110] Specifically, obtaining the sharp degree corresponding to the trough according to the trough value includes performing the following sub-steps D11-D13 on each trough: D11, obtaining the trough values of a preset number of other troughs adjacent to the current trough.

[0111] D12, obtaining the trough center of the current trough and the other troughs.

[0112] D13, obtaining the sharp degree of the trough value of the current trough according to the trough value of the trough center and the trough value of the current trough.

[0113] D3, determining that the fluctuation data corresponding to the trough with a sharp degree greater than the preset sharp threshold is abnormal fluctuation data.

[0114] D4, determining that the fluctuation data corresponding to the trough with a sharp degree not greater than the preset sharp threshold is normal fluctuation data.

[0115] According to the foregoing analysis, we know that when the data fluctuation is large, i.e., the fluctuation parameter is large, the water level data may be in two cases: reasonable rise and fall of water level due to natural phenomena, or abnormal data fluctuation caused by other reasons. At the same time, we judge and distinguish the two kinds of fluctuation data by calculating the sharp degree of the extreme value in the fluctuation data curve . Therefore, in the subsequent data cleaning process, the feature mean calculated by the mean filtering cannot accurately clean the above-mentioned fluctuation data. Specifically, for normal and reasonable data fluctuation, we only need to process the extreme value; for abnormal fluctuation, the influence of the abnormal point on the feature mean is large, so the calculated feature mean cannot well represent the correct data. Therefore, it is necessary to reduce the effect (weight) of the mean filtering, and when the fluctuation is very large, the feature mean calculated according to the water level at the same period has no meaning, so the mean filtering should be stopped and only the slope difference filtering is used on the original water level data.

[0116] According to the above analysis, the hydrological data cleaning processing method combined with the two algorithms is as follows: First, for the collected hydrological data, the initial data is cleaned using the slope difference filtering.

[0117] Next, the data processed by the slope difference filtering algorithm is calculated for its fluctuation parameter , to determine whether it is fluctuation data; if not, the above data is further cleaned using the mean filtering. If so, it is further determined whether it belongs to normal fluctuation data or abnormal fluctuation data. For normal fluctuation data, we need to improve the threshold of the feature mean, and only filter those extremely special values far away from the majority data. For abnormal fluctuation data, the threshold of the feature mean is directly taken as a large number, that is, the mean filtering is not used in the processing.

[0118] In this way, the combination of the above slope difference filtering and mean filtering is completed, that is, the collected data is first filtered using the slope difference filtering, then the fluctuation parameter of the cleaned data is calculated, and the data type is determined according to the fluctuation parameter and the sharpness of the peak value in the fluctuation data, and the threshold of the feature mean in the mean filtering is dynamically adjusted to control the weight of the mean filtering, and the effect of the hydrological data cleaning processing is intelligently controlled. This can ensure the processing of abnormal data while avoiding incorrect processing of correct data as much as possible.

[0119] In the present disclosure, the relevant hydrological data recorded by the hydrological station monitoring record is obtained, and a hydrological data slope difference filtering method suitable for hydrological elements is proposed to combine the recent data trend to predict future data and judge the rationality of the current data. Then a hydrological mean filtering algorithm is proposed, and the effect of the cleaning algorithm is intelligently adjusted according to the data trend and waveform to complete the storage of the cleaned data and record the abnormal data.

[0120] Figure 8 The structure diagram of a hydrological element data cleaning processing system provided by an embodiment of the present application is shown in Figure 8 , which comprises: The acquisition module 11 is used for acquiring hydrological data, and the hydrological data includes water level data. The primary filtering module 12 is used for primary filtering of the hydrological data by using the slope difference algorithm to obtain preliminary cleaning data. The secondary filtering module 13 is used for re-filtering the preliminary cleaning data by using the mean filtering algorithm to obtain the final cleaning data. The storage module 14 is used for storing the final cleaning data.

[0121] In one embodiment, the secondary filtering module 13 is specifically configured to: obtain a fluctuation parameter of the preliminary cleaning data; determine whether the preliminary cleaning data is fluctuation data according to the fluctuation parameter; when the preliminary cleaning data is not the fluctuation data, clean the preliminary cleaning data by using a first mean value filtering algorithm to obtain final cleaning data; when the preliminary cleaning data is the fluctuation data, detect whether the fluctuation data is normal fluctuation data; if the fluctuation data is the normal fluctuation data, clean the preliminary cleaning data by using a second mean value filtering algorithm to obtain the final cleaning data; if the fluctuation data is abnormal fluctuation data, determine that the preliminary cleaning data is the final cleaning data.

[0122] In one embodiment, the primary filtering module 12 is specifically configured to: obtain a change amount of two hydrological data adjacent in time sequence in the hydrological data; obtain a time interval of the two adjacent hydrological data; obtain a slope of the two adjacent hydrological data according to the change amount and the corresponding time interval of the two adjacent hydrological data; obtain a target slope greater than a preset slope threshold, the target slope corresponding to the hydrological data being a preliminary abnormal point; obtain a water level difference value of two adjacent preliminary abnormal points; obtain a target abnormal point greater than a preset change amount threshold, the target abnormal point being the real abnormal data.

[0123] In one embodiment, the secondary filtering module 13 is specifically configured to: obtain a hydrological data difference value of all two hydrological data adjacent in time sequence in the preliminary cleaning data; obtain a fluctuation parameter of the preliminary cleaning data according to a number of all the preliminary cleaning data and a sum of all the hydrological data difference values.

[0124] In one embodiment, the secondary filtering module 13 is specifically configured to: detect whether the fluctuation parameter is greater than a preset fluctuation threshold; if the fluctuation parameter is greater than the preset fluctuation threshold, determine that the preliminary cleaning data is the fluctuation data; if the fluctuation parameter is not greater than the preset fluctuation threshold, determine that the preliminary cleaning data is not the fluctuation data.

[0125] In one embodiment, the secondary filtering module 13 is specifically configured to: obtain a peak value of each peak in the fluctuation data; obtain a sharp degree of the peak according to the peak value; determine that the fluctuation data corresponding to the peak with the sharp degree greater than a preset sharp threshold value is the abnormal fluctuation data; determine that the fluctuation data corresponding to the peak with the sharp degree not greater than the preset sharp threshold value is the normal fluctuation data.

[0126] In one embodiment, the secondary filtering module 13 is specifically configured to: perform the following steps on each peak: obtain a preset number of other peaks adjacent to the current peak; obtain a peak center of the current peak and the other peaks; obtain a sharp degree of the current peak according to a peak value of the peak center and a peak value of the current peak.

[0127] In one embodiment, the secondary filtering module 13 is specifically configured to: determine abnormal data in the preliminary cleaning data; filter out the abnormal data from the preliminary cleaning data to obtain the final cleaning data; The determination of the abnormal data in the preliminary cleaning data comprises: perform the following steps on each cleaning data in the preliminary cleaning data: obtain an occurrence time of the current cleaning data; obtain a preset number of historical hydrological data filtered by the slope difference value algorithm and the same occurrence time as the occurrence time; obtain a mean value of the historical hydrological data; obtain a mean value feature gap corresponding to the current cleaning data according to the mean value and the current cleaning data; if the mean value feature gap is greater than a preset mean value threshold, determine that the current cleaning data is the abnormal data.

[0128] In one embodiment, the preset mean value threshold corresponding to the first mean value filtering algorithm is less than the preset mean value threshold corresponding to the second mean value filtering algorithm.

[0129] It is to be noted that the sequential order of the above-described embodiments of the present application only for the purpose of description, but not the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.

[0130] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments.

Claims

1. A method for cleaning and processing hydrological element data, characterized in that: The method comprises: The slope difference algorithm is used to perform initial filtering on the hydrological data to obtain preliminary cleaned data; Obtaining fluctuation parameters of the preliminary cleaning data; determining whether the preliminary cleaning data is fluctuation data according to the fluctuation parameter; When the preliminary cleaned data is not the fluctuation data, the preliminary cleaned data is cleaned using a first mean filtering algorithm to obtain final cleaned data; When the preliminary cleansing data is the fluctuation data, detecting whether the fluctuation data is normal fluctuation data; If the fluctuation data is the normal fluctuation data, the preliminary cleaned data is cleaned using a second mean filtering algorithm to obtain the final cleaned data; If the fluctuation data is abnormal fluctuation data, the preliminary cleaning data is determined to be final cleaning data.

2. The hydrological element data cleaning and processing method according to claim 1, characterized in that: After the final cleansing data is determined, it also includes: The final cleansing data is stored.

3. The hydrological element data cleaning and processing method according to claim 1, characterized in that: The slope difference algorithm is used to perform initial filtering on the hydrological data to obtain preliminary cleaned data, including: Obtaining a change in two hydrological data that are adjacent in time series in the hydrological data; Obtaining the time interval of the two adjacent hydrological data; Obtaining the slope of the two adjacent hydrological data according to the change amount and the corresponding time interval of the two adjacent hydrological data; Obtaining a target slope whose slope of the two adjacent hydrological data is greater than a preset slope threshold, the hydrological data corresponding to the target slope being a preliminary abnormal point; Obtaining a water level difference between two adjacent preliminary abnormal points; Obtain a target abnormal point where the water level difference is greater than a preset change threshold, where the target abnormal point is real abnormal data; The real abnormal data is filtered from the hydrological data to obtain the preliminary cleaned data.

4. The hydrological element data cleaning and processing method according to claim 3, characterized in that: The obtaining of the fluctuation parameter of the preliminary cleaning data includes: Obtaining the hydrological data difference between all two hydrological data that are adjacent in time sequence in the preliminary cleaned data; The fluctuation parameter of the preliminary cleaning data is obtained according to the quantity of all the preliminary cleaning data and the sum of all the hydrological data differences.

5. The hydrological element data cleaning and processing method according to claim 3, characterized in that: The determining, based on the fluctuation parameter, whether the preliminary cleaning data is fluctuation data includes: Detecting whether the fluctuation parameter is greater than a preset fluctuation threshold; If the fluctuation parameter is greater than the preset fluctuation threshold, determining that the preliminary cleaning data is the fluctuation data; If the fluctuation parameter is not greater than the preset fluctuation threshold, it is determined that the preliminary cleaning data is not the fluctuation data.

6. The hydrological element data cleaning and processing method according to claim 5, characterized in that: The detecting whether the fluctuation data is normal fluctuation data includes: Obtaining the peak values ​​of all peaks in the fluctuation data; Acquire the sharpness corresponding to the corresponding peak according to the peak value; Determining the fluctuation data corresponding to the peak whose sharpness is greater than a preset sharpness threshold as the abnormal fluctuation data; The fluctuation data corresponding to the peak whose sharpness is not greater than the preset sharpness threshold is determined to be the normal fluctuation data.

7. The hydrological element data cleaning and processing method according to claim 6, characterized in that: The acquiring the sharpness corresponding to the corresponding peak according to the peak value includes: For each of the peaks, perform the following steps: Get a preset number of other peaks adjacent to the current peak; Obtaining the peak centers of the current peak and the other peaks; The sharpness of the current peak is obtained according to the peak value of the peak center and the peak value of the current peak.

8. The hydrological element data cleaning and processing method according to claim 7, characterized in that: The step of using a first mean filtering algorithm to clean the preliminary cleaned data includes: Determining abnormal data in the preliminary cleansed data; filtering out the abnormal data from the preliminary cleaned data to obtain the final cleaned data; The determining of abnormal data in the preliminary cleansed data includes: Perform the following steps on each cleaned data in the preliminary cleaned data: Get the occurrence time of the current cleaning data; Acquire a preset number of historical hydrological data having the same occurrence time as the occurrence time and filtered by the slope difference algorithm; Obtaining the mean of the historical hydrological data; Obtaining a mean feature gap corresponding to the current cleaning data based on the mean and the current cleaning data; If the mean feature difference is greater than a preset mean threshold, the current cleaned data is determined to be the abnormal data.

9. The hydrological element data cleaning and processing method according to claim 8, characterized in that: The preset mean threshold corresponding to the first mean filtering algorithm is smaller than the preset mean threshold corresponding to the second mean filtering algorithm.

10. A hydrological element data cleaning and processing system, characterized in that: The system comprises: A collection module is used to collect hydrological data, wherein the hydrological data includes: water level data; A primary filtering module, configured to perform a primary filtering on the hydrological data using a slope difference algorithm to obtain preliminary cleaned data; a secondary filtering module, configured to obtain a fluctuation parameter of the preliminary cleaned data; determine whether the preliminary cleaned data is fluctuation data based on the fluctuation parameter; if the preliminary cleaned data is not fluctuation data, clean the preliminary cleaned data using a first mean filtering algorithm to obtain final cleaned data; if the preliminary cleaned data is fluctuation data, detect whether the fluctuation data is normal fluctuation data; if the fluctuation data is normal fluctuation data, clean the preliminary cleaned data using a second mean filtering algorithm to obtain the final cleaned data; if the fluctuation data is abnormal fluctuation data, determine that the preliminary cleaned data is final cleaned data; A storage module is used to store the final cleaning data.

Citation Information

Patent Citations

  • Urban flood ponding monitoring data cleaning method

    CN113111056A

  • Method and device for monitoring sludge level increment of black and odorous water body sediment

    CN113155231A

  • Water conservancy project anomaly detection method

    CN116429220A

  • Method and system for treating odor generated by sewage treatment

    CN117874583A

  • Integrated configuration verification method and system for station digital water level navigation data

    CN119293461A

Cited By

  • Data exception alarm method and device, electronic equipment and storage medium

    CN121637335A