A method for quickly evaluating the running state of a small reservoir dam safety monitoring facility
By employing singular value identification and a hierarchical progressive identification architecture, environmental fluctuations are isolated, enabling high-precision assessment of dam monitoring facilities. This solves the problem of insufficient accuracy in anomaly identification caused by environmental fluctuation interference in existing technologies, and improves the accuracy and robustness of equipment fault identification.
Patent Information
- Application Number
- CN202611130140.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies struggle to effectively distinguish between fluctuations caused by non-stationary environmental factors and equipment malfunctions when processing dam monitoring data. This results in insufficient accuracy and robustness in anomaly identification, making it impossible to accurately assess the true availability of monitoring facilities.
A hierarchical, progressive identification architecture combining singular value identification, moving median detrending, MAD robust detection, and dual thresholds is adopted to isolate environmental fluctuations. By extrapolating the local difference median and performing spatiotemporal continuity analysis, a missing duration quantification strategy is constructed to achieve compatibility assessment of high-frequency automated equipment and low-frequency manual equipment.
It improves the accuracy of rapid assessment of the operational status of dam monitoring facilities, reduces the risk of underestimating equipment downtime due to occasional data loss, and ensures the ability to automatically classify and characterize multi-dimensional physical fault patterns.
Smart Images

Figure CN122634461A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of dam safety monitoring technology, and in particular to a method for rapid assessment of the operational status of safety monitoring facilities for small reservoir dams. Background Technology
[0002] The safe operation of reservoir dams is highly dependent on continuous and accurate data feedback from monitoring facilities. Efficient and quantitative assessment of the operational status of these facilities can promptly identify underlying physical and link issues such as sensor malfunctions, communication interruptions, or data drift. This is not only a technical prerequisite for ensuring the availability of monitoring data on dam environmental parameters, deformation, and seepage flow, but also a core foundation for the effective operation of the engineering structure safety status analysis and early warning system.
[0003] Currently, quality control and facility status assessment of dam monitoring data typically employ a combination of fixed numerical thresholds and conventional statistical distribution tests for anomaly screening. However, in real-world applications, the complex environment of dams often results in monitoring time-series signals superimposed with strong seasonal water level fluctuations and temperature cycles—non-stationary low-frequency trends. Directly applying detection limits based on conventional means and standard deviations to raw signals containing macroscopic physical trends can easily lead to misjudgments of normal high and low water level responses, or the masking of genuine minor fault characteristics by significant environmental fluctuations. Furthermore, conventional methods often equate missing data or minor anomalies simply with transient equipment failure. In engineering environments where high-frequency automated equipment and low-frequency manual observations coexist, a severe imbalance often arises between the sensitivity of evaluation indicators and the assessment of the true condition of the equipment.
[0004] In summary, existing monitoring facility assessment mechanisms face significant bottlenecks in anomaly identification accuracy and overall robustness when processing time-series data strongly disturbed by non-stationary environmental factors, and struggle to scientifically accommodate diverse equipment acquisition characteristics. Therefore, it is necessary to research an assessment method that can effectively adapt to non-stationary operating conditions and accurately quantify the true availability of facilities. Summary of the Invention
[0005] Purpose of the invention: To provide a rapid assessment method for the operational status of safety monitoring facilities for small reservoir dams, in order to solve the aforementioned problems in the prior art.
[0006] Technical solution: A method for rapid assessment of the operational status of safety monitoring facilities for small reservoir dams, comprising:
[0007] Obtain the initial monitoring point status and on-site detection status of the target dam safety monitoring facilities, and determine the status of the monitoring point itself based on the initial monitoring point status or on-site detection status;
[0008] The monitoring time series collected by the safety monitoring facilities of the target dam within a set statistical period is obtained, and the outlier identification results are obtained by performing outlier identification on the monitoring time series.
[0009] Based on the monitoring time series and outlier identification results, the percentage of abnormal data, the current data stoppage duration, and the duration of missing valid data are calculated.
[0010] The working status of the measurement points is determined based on preset threshold judgment rules, taking into account the percentage of abnormal data, the duration of current data suspension, and the duration of missing valid data.
[0011] By combining the physical status of the measuring points with their operational status, the operational status evaluation results of the target dam safety monitoring facilities are determined and output.
[0012] Beneficial effects:
[0013] A hierarchical, progressive identification architecture is adopted, which includes physical coarse screening, moving median detrending, and MAD robust detection. This helps to isolate low-frequency baselines affected by environmental fluctuations, decouple anomaly identification under complex operating conditions from macroscopic fluctuations, and helps to avoid baseline collapse of conventional models under high anomaly ratios.
[0014] By employing endpoint extrapolation based on the local difference median, phase lag distortion at sequence boundaries can be suppressed, thereby improving the fidelity of the determination of the latest monitoring node status.
[0015] By introducing a spatiotemporal continuity analysis mechanism that combines soft and hard thresholds with continuous segment length, the system is equipped with the ability to automatically classify and characterize anomalies from a single point in time to multi-dimensional physical fault modes such as "isolation, jump, and drift".
[0016] A missing duration quantification strategy based on a minimum effective data ratio threshold is constructed, supplemented by a low-frequency observation degradation branch. This helps reduce the risk of systematically underestimating equipment downtime due to occasional single effective data points, achieving compatibility and characterization of the actual operational reliability of high-frequency automated equipment and low-frequency manual equipment. Attached Figure Description
[0017] Figure 1 This is a flowchart of a method for rapid assessment of the operational status of a safety monitoring facility for a small reservoir dam, according to the present invention.
[0018] Figure 2 This is a flowchart illustrating the process of identifying outliers in the time series monitored by this invention, and obtaining the outlier identification results.
[0019] Figure 3 This is a flowchart illustrating how the present invention filters a monitoring time series based on physical constraints to obtain a preliminary purified sequence.
[0020] Figure 4 This is a flowchart illustrating the process of extracting local trend estimates from the initial purification sequence and calculating the residual sequence.
[0021] Figure 5 This is a flowchart of the extension steps for eliminating end-boundary distortion of the sequence according to the present invention. Detailed Implementation
[0022] like Figure 1 As shown, a rapid assessment method for the operational status of a safety monitoring facility for a small reservoir dam is provided, which mainly includes the following steps:
[0023] Step 101: Obtain the initial measurement point status and on-site detection status of the target dam safety monitoring facilities, and determine the status of the measurement point itself based on the initial measurement point status or on-site detection status.
[0024] The target dam safety monitoring facilities specifically cover continuous numerical monitoring facilities, encompassing various measuring points for automated acquisition of reservoir water level, seepage pressure, surface deformation, and seepage flow. The initial measuring point status refers to the default physical baseline status registered in the dam's daily maintenance system, specifically divided into normal and failed states. The on-site testing status refers to the updated status entered by engineers after physical verification of the measuring point equipment in the field. The measuring point status serves as the fundamental qualitative benchmark for evaluation, and its determination logic follows the principle of prioritizing the latest information, contributing to the real-time accuracy of the benchmark information.
[0025] In some specific implementations, if the target dam safety monitoring facility has not undergone on-site physical testing recently, or if on-site testing confirms that its physical benchmark has not changed, the initial measurement point status obtained is directly identified as the measurement point's physical status.
[0026] Step 102: Obtain the monitoring time series collected by the target dam safety monitoring facilities within the set statistical period, perform singular value identification on the monitoring time series, and obtain the singular value identification results;
[0027] Specifically, the statistical period is typically set to a complete hydrological year to cover the full cycle of high and low water level changes. The monitoring time series consists of continuous monitoring data nodes with timestamps.
[0028] Because dam monitoring instruments may generate distorted data that does not represent the true physical changes of the engineering entity due to power supply anomalies, communication interference, or sensor aging, it is necessary to use identification algorithms to remove this distorted data. The process of performing outlier identification involves locating and marking various abnormal data nodes from the original monitoring time series to prevent false fluctuations from interfering with subsequent reliability quantitative assessments.
[0029] For example, in the pre-execution stage of acquiring monitoring time series, the system establishes a unique mapping association table between the physical number of the measurement point and the acquisition channel number in the database, which helps to prevent channel misalignment during the extraction of time series data from multi-source heterogeneous measurement points.
[0030] Step 103: Based on the monitoring time series, the current assessment time, and the outlier identification results, calculate the percentage of abnormal data, the current data stoppage duration, and the duration of missing valid data.
[0031] Specifically, this includes: statistically monitoring the total amount of data in the time series within a set statistical period, and statistically analyzing the corresponding abnormal data based on the outlier identification results, and calculating the ratio of the abnormal data to the total data as the percentage of abnormal data.
[0032] Extract the latest monitoring time from the monitoring time series, calculate the difference between the current assessment time and the latest monitoring time, and use it as the current data stoppage duration;
[0033] Based on the monitoring time series and outlier identification results, the retention status of valid data for each time unit within the set statistical period is determined, and the time units where valid data is determined to be missing are accumulated to obtain the duration of valid data missing.
[0034] Among them, the percentage of data anomalies is a quantitative indicator characterizing the overall degree of contamination of the monitoring data within the statistical period. The current data outage duration refers to the time span between the current assessment execution time and the timestamp corresponding to the latest valid measurement in the monitoring time series. The valid data missing duration is the sum of the total time span during the entire statistical period during which unrepresentative valid data was returned due to instrument malfunction.
[0035] By simultaneously extracting time-series indicators from these three dimensions, a complete characterization of the dynamic operation characteristics of the measurement points can be achieved from three aspects: data continuity, data validity, and current activity.
[0036] In some embodiments, the method may also be: based on the monitoring time series and the outlier identification results, determine the monitoring time corresponding to the latest valid data within a set statistical period, and calculate the difference between the current evaluation time and the monitoring time corresponding to the latest valid data as the current data stoppage duration.
[0037] Optionally, the percentage of data anomalies is obtained by calculating the ratio of the total number of outliers to the total amount of data in the monitored time series.
[0038] Step 104: Based on the percentage of abnormal data, the current duration of data suspension, and the duration of missing valid data, determine the working status of the measurement point according to the preset threshold judgment rules.
[0039] Specifically, this includes comparing the percentage of abnormal data, the current duration of data suspension, and the duration of missing valid data with the pre-configured corresponding state classification thresholds, and classifying the working status of the target dam safety monitoring facility's monitoring points into one of the following states: normal state, unstable state, suspension state, or failure state, based on the comparison results.
[0040] In this stage, the system maps continuous quantitative indicators to four discrete categories based on the division of numerical intervals: normal operating state, unstable operating state, stopped operating state, and failed operating state. For example, when the percentage of abnormal data at a measurement point is less than 10% and other indicators meet the minimum requirements, it is determined to be in a normal operating state; when the percentage of abnormal data is greater than or equal to 10% but less than 50%, it is determined to be in an unstable operating state. By setting these hard judgment boundaries with engineering tolerance, the system can automatically and quickly screen out monitoring nodes with potential hidden dangers.
[0041] Furthermore, when assessing the status of work stoppage, for automated observation points, a stoppage duration of ≥7 days is sufficient to trigger a judgment; while for manual observation and monitoring projects, the stoppage duration benchmark for triggering the judgment is adaptively replaced with the statutory observation cycle corresponding to the observation point.
[0042] Step 105: Combining the physical status of the measuring point with its operational status, determine and output the operational status evaluation results of the target dam safety monitoring facilities.
[0043] The operational status evaluation results are divided into three levels: reliable, basically reliable, and unreliable. The specific judgment logic employs a superimposed verification mechanism: only when the measurement point's physical state and operational status are both normal, is the output evaluation result considered reliable; when the measurement point's operational status deteriorates to an unstable state, the output result is downgraded to basically reliable; once the measurement point's physical state is determined to be faulty, or although its physical state is normal, its operational status drops to the failure range due to prolonged downtime or data anomalies (≥50%), an unreliable evaluation result is forcibly output. This mechanism strictly adheres to stringent evaluation criteria to prevent the existence of faulty instrument nodes in the dam safety monitoring system.
[0044] Specifically, in some embodiments, the process is as follows: based on the pre-configured state mapping rules, when the state of the measuring point body is normal and the working state of the measuring point is determined to be normal, the running state evaluation result is output as reliable.
[0045] When the status of the measuring point itself is normal and the working status of the measuring point is determined to be unstable, the operating status evaluation result will be output as basically reliable.
[0046] When the measurement point itself is in a state of failure, or when the measurement point is determined to be in a state of stoppage or failure and meets the pre-configured adverse continuous conditions, the operation status evaluation result will be output as unreliable.
[0047] As an improvement to the above scheme, considering that the threshold determination rule relies on discrete numerical boundary divisions, there is an objective step jump effect that can cause the evaluation result to drop directly from the best to the worst. When the specific values of the various operational evaluation indicators involved in the calculation fall within the preset boundary sensitive range, the system will generate a manual review warning message for that measurement point while generating the corresponding operational status evaluation result.
[0048] In some embodiments, when the measurement point body status is in failure, the operation status evaluation result is output as unreliable; when the measurement point body status is normal, and the measurement point working status is determined to be in a stopped state and the current data stop time exceeds the pre-configured stop time tolerance threshold, or when the measurement point working status is determined to be in failure, the operation status evaluation result is output as unreliable.
[0049] According to one aspect of this application, singular value identification is performed on a monitoring time series to obtain singular value identification results, including:
[0050] like Figure 2 As shown, in step 201, the monitoring time series is filtered based on physical constraints, and data points that do not meet the physical constraints are marked as physical outliers, thus obtaining a preliminary purified sequence.
[0051] Dam safety monitoring data are subject to long-term interference from objective physical factors such as reservoir water level fluctuations and alternating environmental temperature changes, resulting in a clear non-stationary trend over time.
[0052] In this physical context, conventional single statistical tests, due to the implicit prior assumption that the data as a whole follows a specific distribution, are prone to classifying normal physical response data during periods of high water levels as outliers. To prevent the interference of the non-stationary nature of the data itself on statistical discrimination criteria, a sequential and progressive purification pipeline must be constructed.
[0053] By utilizing the spatial dimensions of the structure itself and the range limits of the instrument, absolute numerical boundaries are set to filter out extreme outliers caused by short circuits in the front-end sensors or severe packet loss in the communication link.
[0054] By eliminating these distorted values that have no physical meaning in the early stages, it can be ensured that when calculating trend estimates based on local data windows in subsequent levels, there will be no serious deviation due to the presence of extreme values.
[0055] Step 202: Extract the local trend estimate of the preliminary cleanup sequence and calculate the residual sequence;
[0056] After completing the initial screening based on physical constraints, the system initiates the second level of processing in the serial pipeline. This step utilizes a sliding window algorithm to continuously calculate robust eigenvalues on the initially purified sequence, thereby approximating the slowly changing baseline of the actual engineering physical response over time. By performing a difference operation to remove this slowly changing baseline, the system forcibly transforms the non-stationary time series into a stationary residual series with a mean approaching zero, thus satisfying the mathematical prerequisites for subsequent statistical testing methods.
[0057] Step 203: Calculate the dispersion index of the residual sequence, and detect the residual sequence based on the dispersion index and using a dual threshold mechanism to obtain the out-of-limit classification labeling results.
[0058] This step is the third level of processing in the serial pipeline. The system extracts a robust scatter metric from this stationary sequence that is not affected by individual outliers, establishing a statistical scale with tolerance space. By introducing a stepped dual decision threshold, the system not only identifies isolated data that deviates absolutely from the normal, but also accurately captures data nodes with small deviations but showing anomaly tendencies, outputting a set of labels with severity distinctions.
[0059] Step 204: Perform continuity analysis based on the out-of-limit classification labeling results to determine the singular value classification results. Singular value classification results and physical out-of-limit singular values are summarized, and the summarized results are used as singular value identification results.
[0060] This step constitutes the fourth level of pipeline processing, completing the dimensional leap from point-in-time detection to spatiotemporal structure analysis over a time period. The system scans continuous data blocks labeled as exceeding limits, and based on the duration of the abnormal state, maps individual numerical anomalies to anomaly categories with clear physical diagnostic significance. All categories are then aggregated to form the final output.
[0061] Based on the above embodiments, it is clear that through four sequential processing levels with a strict causal inheritance relationship, the system not only solves the interference of environmental quantity changes on recognition accuracy, but also provides a refined classification and diagnostic basis for equipment failure modes. Specific algorithmic details regarding local trend estimation, residual detection, and continuous classification will be further detailed and disclosed in subsequent embodiments.
[0062] According to one aspect of this application, a monitoring time series is filtered based on physical constraints to obtain a preliminary purified sequence, including:
[0063] like Figure 3 As shown, in step 301, based on the engineering design parameters of the target dam and the range information of the monitoring instruments used, the upper and lower limits of the fixed threshold are set.
[0064] Specifically, different monitoring items possess different physical dimensions and ranges of change, requiring the system to establish independent extreme value constraint intervals for each monitoring dimension. The engineering design parameters of the target dam include various characteristic elevations, while the measurement range information is determined by the sensor's factory calibration parameters. For example, for the seepage pressure monitoring dimension, the system extracts the head value converted from the orifice elevation of the corresponding piezometer as the upper limit and extracts the head value corresponding to the piezometer's installation elevation as the lower limit. In extreme cases where early verification parameters are missing, the system adaptively adopts the dam crest elevation as the upper limit and the dam foundation elevation as the lower limit. For the temperature monitoring dimension, the upper limit is limited to 50 degrees Celsius, and the lower limit is -20 degrees Celsius, based on the instrument's measurement range.
[0065] Step ①: Obtain the physical parameters of the maximum dam height of the target dam;
[0066] For deformation monitoring projects such as surface displacement, the normal structural deformation amplitude exhibits a strong positive correlation with the overall scale of the dam. Directly setting a globally invariant absolute numerical boundary can easily lead to missed or false detections. Therefore, the system reads the dam foundation information database to obtain the elevation difference between the highest and lowest points of the main dam cross section, using this as the physical parameter for the maximum dam height. This physical parameter constitutes the benchmark for adaptively generating allowable displacement boundaries.
[0067] Step 2: Extract the pre-configured empirical proportional coefficients, and multiply the physical parameter of the maximum dam height by the corresponding empirical proportional coefficients to calculate the extreme value of the allowable displacement for the corresponding monitoring dimension.
[0068] For surface deformation monitoring projects, the system maintains a configuration table storing empirical proportional coefficients corresponding to various displacement dimensions. The system extracts pre-determined proportional values and amplifies them to the current physical scale of the dam through multiplication operations. The specific calculation relationship is expressed as follows:
[0069] D limit =c*H;
[0070] Among them, D limit Here, c represents the extreme value of the allowable displacement, c is the empirical proportionality coefficient, and H is the physical parameter of the maximum dam height.
[0071] Step ③: Use the positive and negative interval boundaries of the allowable displacement extreme values as the upper and lower limits for filtering obvious instrument fault data, respectively.
[0072] Based on the above calculation results, a safe deformation range is symmetrically constructed in the system. For example, for the surface horizontal displacement monitoring dimension, the corresponding empirical scaling factor is extracted as 0.5%, and the positive allowable displacement extreme value obtained by substituting it into the calculation is the upper limit, and the negative allowable displacement extreme value is the lower limit. The corresponding extreme value calculation is expressed as follows:
[0073] L high=+0.5%*H;
[0074] L low =-0.5%*H;
[0075] Among them, L high L is the upper limit of a fixed threshold. low Here, H represents the lower limit of a fixed threshold, and H is the physical parameter for the maximum dam height. Similarly, for the surface vertical displacement monitoring dimension, the corresponding empirical proportionality coefficient is extracted as 0.3%. The calculated upper limit is the product of 0.3% and the maximum dam height, and the lower limit is the product of -0.3% and the maximum dam height. These physical upper and lower limits, derived from engineering experience, are specifically used to intercept obviously unreasonable jumps in data caused by displacement gauge jamming or reference point displacement.
[0076] Step 302: Remove values in the monitoring time series that exceed the upper limit or fall below the lower limit and mark them as physical outliers. Use the remaining data series as the initial cleanup series.
[0077] The system executes a traversal comparison logic, comparing each acquired sequence point with its corresponding upper and lower limits. If a value exceeds the limits at any given time, an outlier attribute label is immediately added to that data node, and it is removed from the current processing sequence. After one complete traversal, the remaining data set constitutes the initial cleaned sequence. This sequence has eliminated the most severe physical limit-crossing interference, providing a data quality foundation for the next level of algorithm to extract stationary residuals.
[0078] In some specific implementations, for three-dimensional deformation monitoring points with multiple spatial components, the system independently performs the above-mentioned physical constraint filtering operation on the monitoring time series of each coordinate axis component, which helps to prevent cross-contamination of anomalies in different dimensions.
[0079] like Figure 4 As shown, according to one aspect of this application, a local trend estimate of the preliminary cleanup sequence is extracted, and a residual sequence is calculated, including:
[0080] Step 401: For the actual data points in the preliminary purification sequence, select a local window that covers the actual data points based on the set window width;
[0081] In this step, the true data points refer to the valid measurement nodes retained in the time series after filtering by physical constraints. The selection of the window width directly determines the smoothness of the local trend estimation, and its specific value is calibrated based on the typical physical change time scales of various monitoring projects.
[0082] For example, in a reservoir water level monitoring project, the window width can be configured to correspond to the number of data points over 7 days. For a seepage pressure monitoring project, the window width can be configured to correspond to the number of data points over 15 days. For a surface deformation monitoring project, the window width can be configured to correspond to the number of data points over 30 days. For a seepage flow monitoring project, the window width can be configured to correspond to the number of data points over 7 days. These configurations ensure that the local window contains sufficient samples across a given time dimension to reflect the long-term trend evolution of objective engineering variables.
[0083] Step 402: Calculate the median of the data within the local window and use it as the local trend estimate for the corresponding real data points;
[0084] All values falling within the local window are sorted, and the value at the median position is extracted as the local trend estimate. The median, rather than the arithmetic mean, is used for local trend extraction because the median estimation algorithm has a 50% failure rate. When the proportion of outliers remaining within the local window that are not filtered out by physical constraints is less than 50%, the extracted trend baseline is unaffected by extreme outliers. This extraction mechanism ensures the robustness of the baseline estimation process.
[0085] Step 403: The difference between the preliminary purified sequence and the corresponding local trend estimate is calculated to obtain the residual sequence with the trend component removed.
[0086] The residual sequence is obtained by subtracting the corresponding local trend estimate from the actual data point value at the same moment in a time series. After the difference calculation, the output sequence is the residual sequence. By removing the low-frequency components that fluctuate with environmental factors, the residual sequence is transformed into a stationary sequence that approximately fluctuates around zero. This stationarization transformation allows subsequent statistical detection algorithms to operate within their established mathematical premises.
[0087] Based on the above embodiments, another implementation method including endpoint extension processing is provided to address the issue of incomplete sliding windows at the ends of the sequence. When a local window slides to the beginning or end of a time series, the number of available real data points on one side of the window is insufficient, causing the calculated median to be biased towards samples from the historical or future directions.
[0088] Systematic phase shifts can cause tail trend estimates to lag, artificially amplifying the residuals of the most recently acquired data and triggering false anomaly markers. To eliminate these boundary distortions, an extension step to eliminate end-of-sequence boundary distortions is included before calculating the median of the data within the local window.
[0089] like Figure 5 As shown, according to one aspect of this application, before calculating the median of the data within a local window, an extension step to eliminate end-boundary distortion of the sequence is further included, as follows:
[0090] Step 501: Extract the first and last data segments of the preliminary purification sequence, respectively.
[0091] Specifically, the system uses a set window half-width as a benchmark, extracting a certain number of consecutive data nodes from the beginning of the preliminary purified sequence to form the head data segment, and extracting the same number of consecutive data nodes from the end of the sequence to form the tail data segment. The number of data nodes extracted can be configured to be the smaller of twice the window half-width and the total sequence length. By limiting the reference range, the system obtains the local true trend of change at the ends, preventing fluctuation signals at the far end of the time series from interfering with extrapolation calculations.
[0092] Step 502: Calculate the median of the point-by-point difference sequence within the first data segment as the left-end local trend slope, and virtually extrapolate a preset number of data points to the outside of the first data segment based on the left-end local trend slope.
[0093] Within the first data segment, the difference between two adjacent time points is calculated to form a pointwise difference sequence. Subsequently, the median of this pointwise difference sequence is calculated, and the output is the left-hand local trend slope. Based on this, the system uses a determined slope parameter to generate virtual node data at the leftmost edge of the time series. The specific calculation process is as follows:
[0094] y leftk =y start -k*d L ;
[0095] Among them, y leftk For the k-th data point virtually extrapolated to the left, y start The first real data point in the first data segment, k is the extrapolation step size index, and d L This is the calculated local trend slope on the left. The system controls the step size index to increment, generating a number of virtual extrapolated data points equal to half the window width. Using the median difference to estimate the trend slope also leverages its collapse point characteristic, providing a robust local trend unaffected by residual outliers.
[0096] Step 503: Calculate the median of the point-by-point difference sequence within the tail data segment as the right-end local trend slope, and virtually extrapolate a preset number of data points to the outside of the tail data segment based on the right-end local trend slope.
[0097] Using a similar mathematical processing procedure as the left end, the system processes the tail data segment to extract the local trend slope on the right end. The virtual extrapolation of the right end endpoints ensures that data nodes reflecting the latest operating status of the project can undergo symmetrical window calculations.
[0098] The specific calculation is as follows: y rightk =y end+k*d R ;
[0099] Among them, y rightk For the k-th data point virtually extrapolated to the right, y end Here, k represents the last real data point in the tail data segment, k is the extrapolation step size index, and d is the last real data point in the tail data segment. R This is the calculated local trend slope on the right side.
[0100] Step 504: The data points obtained by virtual extrapolation are spliced with the preliminary purification sequence to form an extended sequence;
[0101] The system combines the data array by concatenating the virtual nodes generated on the left, the sequence of real data nodes in the center, and the virtual nodes generated on the right in chronological order according to their timestamps. The resulting extended sequence expands on both sides by a number of data units equal to half the width of the window.
[0102] Step 505: Select a local window that covers the real data point. Specifically, select a symmetrical local window that covers the corresponding real data point on the extended sequence; and after completing the calculation of the local trend estimate, remove the calculation results corresponding to the virtual extrapolated data point.
[0103] By utilizing the pre-completed end-filling operation, when performing a sliding operation on any real data point on the extended sequence, it is possible to extract complete data nodes before and after the point, thus constructing a symmetrical local window.
[0104] Within these symmetrical windows containing virtual data, the median is calculated to output the trend baseline. After the baseline calculation for all real data points has been completed, the trend calculation return values corresponding to the first and last virtual nodes are actively discarded, and only the local trend estimate corresponding to the real data index is retained and output.
[0105] In some optional implementations, when the total number of valid data nodes in the initial purified sequence is less than or equal to the set overall width of the local window, it indicates that the sequence sample length is insufficient to support the sliding window's traversal. In this case, the system activates a degradation processing mechanism and no longer performs the aforementioned endpoint extension and sliding median extraction calculations.
[0106] The median of the entire set of features of the entire effective data sequence is directly calculated, and this constant is regarded as the local trend estimate corresponding to all real data points.
[0107] According to one aspect of this application, a dispersion index of the residual sequence is calculated, and a dual-threshold mechanism is used to detect the residual sequence to obtain out-of-limit classification labeling results, including:
[0108] Step 601: Calculate the median absolute deviation of the residual sequence as a robust dispersion index;
[0109] After obtaining the residual sequence after removing non-stationary trends, it is necessary to extract a benchmark scale parameter that reflects the characteristics of the data fluctuation distribution. Since the conventional standard deviation calculation method based on the mean square deviation is easily skewed by a few outliers with large values, resulting in the compression of the relative dispersion of normal data, this scheme adopts the median absolute deviation, a robust statistical feature with a 50% collapse point.
[0110] Robust dispersion index MAD = median(|r1-r) tilde |,|r2-r tilde |,...,|r m -r tilde |);
[0111] Where MAD is the robust dispersion index, median is the function for calculating the median of the dataset, and r1 to r m r represents the residual value corresponding to each real data point in the residual sequence. tilde is the median of the entire residual sequence, m is the total number of valid data points in the residual sequence, and | is the mathematical operator for taking the absolute value.
[0112] Based on this, and considering the possibility of extreme physical failures that may render the monitoring facilities unresponsive, this solution introduces a foolproof interception design based on the aforementioned robust dispersion index.
[0113] Step ①: Obtain the preset instrument resolution parameters of the monitoring facilities used;
[0114] The system accesses the device's physical attribute database, extracts the minimum resolvable numerical gradient of the hardware sensor corresponding to the current processing sequence that was confirmed during the factory or calibration test phase, and uses it as the preset instrument resolution parameter.
[0115] Step 2: Determine whether the calculated robust dispersion index is less than a fixed multiple of the preset instrument resolution parameter;
[0116] The system compares the robust dispersion index with the preset instrument resolution parameter after being scaled by a fixed factor. In a specific implementation, this fixed factor can be set to 0.5.
[0117] Step ③: If yes, determine that there is no effective physical variation in the residual sequence within the corresponding time period, directly mark the monitoring data of that time period as a suspected sensor failure state, and directly add its corresponding time span to the effective data missing duration. At the same time, terminate the subsequent steps of calculating the corrected standard score and classifying the data for that time period. If no, continue to execute the subsequent steps of calculating the corrected standard score and classifying the data exceeding the limit.
[0118] When the robust dispersion index falls below half of the preset instrument resolution parameter, it objectively indicates that the time series has completely lost the subtle physical noise fluctuations caused by the real natural environment within the corresponding analysis window. This absolute numerical stagnation state highly matches the physical fault characteristics of sensor malfunction or signal transmission freeze. By triggering this circuit breaker mechanism in a timely manner, fake stable false normal data can be accurately intercepted, while saving the subsequent redundant statistical classification computing power consumption.
[0119] Step 602: Based on the ratio of the deviation of each data point from the median of the residual to the robust dispersion index, and combined with the normal distribution correction coefficient, calculate the corrected standard score of each data point.
[0120] For residual sequences that have not triggered the aforementioned circuit breaker mechanism, the system performs a standardization transformation operation on each data node one by one. The specific operation logic is expressed as follows:
[0121] M i =0.6745*(r i -r tilde ) / MAD;
[0122] Among them, M i Let r be the corrected standard score for the i-th data point, and 0.6745 be the normal distribution correction coefficient, which is equal to the third and fourth quartile of the standard normal distribution, i.e., the quartile corresponding to the 75th percentile of the standard normal distribution. i Let r be the residual value corresponding to the i-th true data point in the residual sequence. tilde The median of the entire residual sequence is given, and MAD is the robust dispersion index. The purpose of multiplying by a coefficient of 0.6745 is to ensure that, when the residual data in an ideal state follows a standard normal distribution, the output corrected standard score is numerically equivalent to the traditional standard score, thus being compatible with existing conventional statistical decision limits.
[0123] Furthermore, a computational example is provided to illustrate the robust computational characteristics described above. A local residual array *r* is obtained containing five numerical points: 0.1, 0.2, 0.15, 1.5, and 0.1, where 1.5 is a simulated extreme outlier. The median *r* of the entire array is calculated. tilde The median is 0.15. The absolute deviations of each point from the median are 0.05, 0.05, 0, 1.35, and 0.05, respectively. The median of these absolute deviations is 0.05. Substituting the extreme value of 1.5 into the score calculation formula, the corrected standard score is 0.6745*(1.5-0.15) / 0.05=18.2115.
[0124] If we use the traditional mean of 0.41 and standard deviation of 0.61 as an alternative verification method, the traditional standard score obtained at this extreme point is only (1.5-0.41) / 0.61=1.79. Based on this example, it can be seen that in data environments containing outliers, the median absolute deviation, as a dispersion index, can still maintain an objective benchmark and output a more accurate deviation response value.
[0125] Step 603: Preset a hard threshold for characterizing absolute anomalies and a soft threshold for characterizing systematic shifts, wherein the hard threshold is greater than the soft threshold.
[0126] Unlike methods that rely solely on a single numerical threshold for identification, this scheme employs a two-tiered numerical threshold system. Specifically, a hard threshold, configurable at 3.5 (corresponding to the 99.95% confidence level quantile under a normal distribution), is used to intercept independent data nodes exhibiting anomalies at a single time point. The soft threshold, configurable at 2.0 (corresponding to the 99.95% confidence level quantile under a normal distribution), is used to capture data nodes bordering on normal fluctuations and exhibiting a tendency to deviate.
[0127] Step 604: Compare the absolute value of the corrected standard score of each data point with the hard threshold and the soft threshold respectively. Based on the comparison results, assign the corresponding hard over-limit label, soft over-limit label or normal label to the data points in the residual sequence to obtain the over-limit classification label results.
[0128] The system assigns discrete state labels to sequence nodes based on numerical comparison logic. When the absolute value of a data node's corrected standard score is strictly greater than a preset hard threshold, the system assigns it a hard out-of-limit label; when the absolute value of the corrected standard score is greater than a soft threshold but less than or equal to the hard threshold, the system assigns it a soft out-of-limit label; and when the absolute value of the corrected standard score is less than or equal to the soft threshold, the system assigns it a normal label. At this point, the continuous numerical time series is transformed into a symbolic label sequence containing hierarchical state information, laying the analytical foundation for the next level of inferring structured anomaly categories.
[0129] According to one aspect of this application, in one possible embodiment, a continuity analysis is performed based on the out-of-limit classification labeling results to determine the singular value classification results, including:
[0130] Step 701: In the out-of-limit classification and marking results, extract the data segments that are continuously composed of soft out-of-limit markers or hard out-of-limit markers, and count the number of data points contained in the extracted continuous out-of-limit segments;
[0131] Specifically, the system scans the out-of-limit classification results output from the previous processing level in timestamp order. When a data node with a non-normal marking status is encountered, the system uses it as a starting point and continues to trace data nodes with adjacent timestamps until it encounters a data node marked as normal. The set of data nodes consecutively marked as soft or hard out-of-limit constitutes a continuous anomaly segment. The system counts the total number of nodes contained within this set, defining it as the number of data points, denoted as parameter l. Through this step, the system transforms isolated single-point anomaly discrimination results into one-dimensional sequence structure information containing time span characteristics.
[0132] Step 702: Preset drift detection length threshold and jump detection length threshold, wherein the drift detection length threshold is greater than the jump detection length threshold;
[0133] To accurately distinguish the physical response duration corresponding to different failure modes, the system is configured with two levels of judgment criteria. The jump detection length threshold is usually configured as a fixed constant, denoted as parameter L. j The drift detection length threshold is dynamically calculated based on the daily sampling frequency of the monitoring facility, and is denoted as parameter L. d The specific computational mapping relationship is expressed as follows:
[0134] L d =max(2*f,5);
[0135] Among them, L d is the drift detection length threshold, max is a mathematical function to select the maximum value, f is the daily sampling frequency of the corresponding monitoring facility, and the constant 5 represents the minimum number of data nodes required to determine the drift phenomenon.
[0136] For example, assuming a seepage pressure monitoring point is configured to collect data once per hour, its daily sampling frequency f equals 24. Substituting into the above formula, we can calculate L. d The value is 48, meaning the drift phenomenon must last for at least two days. Simultaneously, a pre-configured threshold L for the jump detection length is set. j The value is 3, meaning that at least 3 consecutive out-of-limit data points are needed to constitute a jump phenomenon segment.
[0137] Step 703: The number of data points counted is compared with the drift detection length threshold and the jump detection length threshold respectively. Based on the comparison results and the label type contained in the continuous anomaly segment, the data points in the continuous anomaly segment are classified as drift-type singular values, jump-type singular values or isolated singular values, thereby generating singular value classification results.
[0138] Existing dam monitoring systems typically only perform binary judgments, lacking the ability to classify and diagnose the physical forms of sensor faults. By coupling the length of a continuous segment with dual threshold types, the system can distinguish three distinct abnormal spatiotemporal structural features: single-point deviation, short-term multi-point deviation, and multi-point sustained moderate deviation.
[0139] According to one aspect of this application, specifically:
[0140] Step 801: When the number of data points is greater than or equal to the drift detection length threshold, all data points in the corresponding continuous abnormal segment are determined to be drift-type singular values.
[0141] Execute logical conditional judgment l≥L d When this condition is met, it objectively indicates that the monitoring data has experienced a prolonged and continuous deviation, consistent with the physical characteristics of sensor reference drift or measurement channel interference from a prolonged environment. Therefore, the system defines the entire continuous abnormal segment as a drift phenomenon and forcibly classifies all data points within this segment into the drift-type singular value set.
[0142] Step 802: When the number of data points is greater than or equal to the jump detection length threshold and less than the drift detection length threshold, all data points in the corresponding continuous abnormal segment are determined to be jump-type singular values.
[0143] System execution logical condition judgment L j ≤l <L d When this condition is met, it indicates that the abnormal state spans multiple acquisition cycles but has not yet reached the scale of long-term drift, and its manifestation is between isolated anomalies and system drift. This short-term multi-point deviation is usually caused by transient circuit failures inside the sensor or sudden external electromagnetic shocks. The system defines the entire continuous segment that meets this condition as a jump phenomenon and classifies all data points within this segment into a jump-type singular value set.
[0144] Step 803: When the number of data points is less than the jump detection length threshold, only the data points with hard overlimit markers in the corresponding continuous abnormal segment are judged as isolated singular values, and the data points with soft overlimit markers are restored as normal data points.
[0145] System execution logical condition judgment l <L jThis condition corresponds to a relatively short, continuous sequence, typically containing only one or two data nodes. In this scenario, the system executes a refined filtering and adjudication mechanism. If a data node within this short, continuous segment has a hard out-of-limit marker, it is considered to have generated an absolute single-point deviation and is classified into the isolated singular value set. Conversely, if the data node only has a soft out-of-limit marker, objective physical evidence determines that the moderate deviation of this small number of nodes falls within the normal statistical fluctuation range caused by the superposition of natural environmental factors. Therefore, the system proactively withdraws its anomaly judgment, restoring it to the valid data sequence and not including it in any singular value set. This restoration mechanism effectively filters out false anomalies introduced by the relatively lenient soft threshold condition, improving the robustness of the overall evaluation system.
[0146] In one possible embodiment, for automated monitoring facilities, the retention status of valid data for each time unit within a set statistical period is determined, and the time units where valid data is determined to be missing are accumulated to obtain the duration of valid data loss, specifically including:
[0147] Step 1001: Divide the set statistical period into multiple consecutive calendar days;
[0148] Specifically, the system discretizes continuous statistical periods using the calendar day as the basic time granularity. It obtains the evaluation start and end timestamps and generates a set of calendar days according to the time window sequence from midnight to midnight each day. This set of calendar days constitutes the basic accumulation unit for calculating equipment downtime due to faults.
[0149] Step 1002: Obtain the total number of measured data entries for the monitored time series within a single calendar day, and count the number of abnormal data entries within that calendar day based on the outlier identification results;
[0150] For each independent calendar day, the system accesses the database to extract the total number of monitoring data records with independent timestamps actually received within that time window, recording this as the total number of measured data entries. Simultaneously, it reads the outlier classification results output from the previous processing stage, calculates the total number of data nodes assigned hard or soft out-of-limit markers within that time window, and records this as the number of abnormal data entries.
[0151] Step 1003: Subtract the number of abnormal data entries from the total number of measured data entries to obtain the number of valid data entries for the corresponding calendar day;
[0152] The system performs data culling and accounting operations on a daily basis. The specific calculation formula is expressed as follows:
[0153] n validj =n totalj -n outlierj ;
[0154] Where, nvalidj Let n be the number of valid data entries for the j-th calendar day. totalj Let n be the total number of measured data points for the j-th calendar day. outlierj Let be the number of outlier data entries within the j-th calendar day, where j is the calendar day sequence index variable. The calculated result represents the scale of clean data that truly reflects the physical operating status of the dam within that day.
[0155] Step 1004: Obtain the pre-configured theoretical acquisition frequency of the monitoring facilities, and determine the theoretical number of data entries to be acquired per day accordingly;
[0156] The system reads the hardware configuration file of the current measuring point equipment and extracts the sampling time interval set by the equipment. Based on the 24-hour time span of the day and this sampling time interval, the quotient is calculated and rounded down to obtain the theoretical number of data points to be collected per day. For example, for an automated monitoring facility with a set sampling frequency of once every two hours, the calculated theoretical number of data points to be collected per day is 12.
[0157] Step 1005: Calculate the ratio of the number of valid data entries to the theoretical number of data entries collected per day to obtain the valid data ratio;
[0158] The system divides the daily clean data size calculated in step 1003 by the theoretical receivable data size calculated in step 1004. The specific calculation formula is as follows:
[0159] R validj =n validj / n expectj ;
[0160] Among them, R validj Let n be the proportion of valid data for the j-th calendar day. validj n represents the number of valid data entries. expectj This represents the theoretical number of data points collected per day. This percentage eliminates benchmark differences between devices with different acquisition frequencies, providing a unified data quality evaluation dimension in the zero-to-one range.
[0161] Step 1006: Compare the effective data ratio with the pre-configured minimum effective data ratio threshold. When the effective data ratio is lower than the minimum effective data ratio threshold, the corresponding calendar day is determined as a day with missing effective data.
[0162] In previous technologies, a single day's data was considered valid as long as it contained at least one unmarked data point. For example, a facility might theoretically collect 24 data points per day, but a malfunction would generate 23 abnormal data points. Judging the day's equipment operation as normal based on just one randomly occurring data point significantly underestimated downtime. To overcome this deficiency, the system presets a minimum valid data retention threshold, specifically configurable to 0.2. When the daily valid data retention rate is less than 20%, the system, based on statistical representativeness principles, determines that the remaining small amount of data represents random noise or transient quantities during fault recovery, insufficient to accurately support engineering safety assessments. Therefore, the system directly classifies the entire calendar day as a day with missing valid data. Using 0.2 as the threshold parameter achieves an engineering balance between recognition sensitivity and fault tolerance conservatism.
[0163] Step 1007: Accumulate all the valid data missing days identified within the set statistical period to obtain the valid data missing duration.
[0164] The system iterates through all calendar days within the defined statistical period, retrieving time units marked as days with missing valid data. For each marked unit retrieved, a counter increments by calendar days. The cumulative sum output after the iteration is complete represents the duration of the missing valid data. This value is directly input into the subsequent state mapping matrix to define the final operational reliability level of the equipment.
[0165] According to one aspect of this application, it also includes:
[0166] Step ①: Preset the acquisition frequency limit to distinguish between high-frequency acquisition and low-frequency observation;
[0167] Because both high-frequency automatic transmission equipment and low-frequency manual portable reading equipment exist at the engineering site, a single missing evaluation standard cannot be compatible with all scenarios. The system sets a constant value as the sampling frequency limit in the environmental configuration file. In a specific implementation, this sampling frequency limit can be set to 5 to distinguish between continuous daytime monitoring items and discrete point measurement items.
[0168] Step 2: Compare the determined daily theoretical number of data entries with the collection frequency limit;
[0169] The system executes the numerical value judgment logic, compares the theoretical number of daily data collections calculated in step 1004 with the collection frequency limit parameter set in step 1501, and guides the data flow into the corresponding judgment branch pipeline according to the numerical value relationship.
[0170] Step ③: When the theoretical number of data entries collected on a given day is greater than or equal to the collection frequency limit, the step of comparing the proportion of valid data with the pre-configured minimum valid data proportion threshold is executed to determine the days with missing valid data.
[0171] The fulfillment of this condition indicates that the measurement points being analyzed belong to automated high-frequency monitoring facilities. Under this branch, the system activates a data validity accounting mechanism based on relative proportions, replacing the absolute zero data missing standard in traditional evaluation methods.
[0172] Step 4: When the theoretical number of data entries collected on a given day is less than the collection frequency limit, stop calculating the proportion of valid data and directly verify whether the number of valid data entries is zero. If it is zero, the corresponding calendar day is determined to be a day with missing valid data. If it is not zero, the corresponding calendar day is determined to be a day without missing data.
[0173] The condition being met indicates that the measurement point being analyzed belongs to a manual observation or low-frequency monitoring project. In this scenario, due to the small denominator, any percentage-based proportioning loses statistical significance. The system blocks the proportion calculation pipeline through a degenerate branch, reverting to the judgment criterion of using absolute counts. Only when the number of valid data points acquired throughout the day is completely equal to zero does the system determine that a valid data missing event has occurred on that manual observation calendar day.
[0174] According to one aspect of this application, a secondary verification optimization mechanism is described as a preferred scheme for a hierarchical progressive identification architecture.
[0175] The screening logic for preventing false positives based on the statistical model of dam effect includes the following steps:
[0176] Step 1101: Based on the outlier classification results, exclude the data points marked as outliers in the monitoring time series and extract the corresponding valid monitoring data.
[0177] Specifically, the system receives the initial classification labels output by the hierarchical progressive identification pipeline. To prevent the statistical model from learning incorrect anomalous features during the parameter estimation stage, the system performs a masking filtering operation on the original acquired sequences. All data nodes identified as isolated singularities, jump singularities, and drifting singularities are removed. The time series sample set retained after the above filtering operation constitutes the effective monitoring data, which is used to support the subsequent regression fitting of the mathematical statistical model.
[0178] Step 1102: Obtain synchronously collected dam environmental data, and construct and fit an effect quantity statistical model based on the effective monitoring data and dam environmental data; the effect quantity statistical model can adopt the HST (water pressure-temperature-time) model.
[0179] In this step, the dam environmental data is a set of external environmental variables that are causally related to the physical response at the measuring points, specifically including parameters such as upstream reservoir water level and calendar time. Considering the classical physical laws of dam structural engineering response, the system uses a statistical model of dam effect quantities that includes water pressure, temperature, and time-dependent components for fitting and calculation. For datasets with the aforementioned multidimensional characteristic variables, the system obtains the undetermined regression coefficients through mathematical solution algorithms, completing the training and construction of the model.
[0180] Furthermore, to ensure the required degrees of freedom and statistical power of the regression analysis, specific preconditions must be met before performing this statistical modeling step. The effective monitoring data must cover at least one complete hydrological year, typically with a time span of no less than 365 days. Simultaneously, upstream reservoir water level data must be available and valid, and the water level fluctuations must cover the main range of the dam's normal operating conditions. In addition, the sample size N used for fitting must satisfy N > 2*p, where p is the total number of regression coefficients in the model.
[0181] Step 1103: Calculate the theoretical values of the monitoring time series using the fitted effect size statistical model, and subtract the actual collected values from the theoretical calculated values to obtain the model residual sequence.
[0182] The system inputs complete time-axis environmental parameters, including suspected anomaly data, into the fitted effect size statistical model. Based on a predetermined function mapping, the model outputs the theoretically calculated value for each time point. Subsequently, the system subtracts the corresponding theoretically calculated value from the actual collected values point by point. Through this difference calculation, the system removes the interpretable physical fluctuations caused by changes in the external environment, outputting only the model residual sequence containing unexplained random measurement errors or inherent sensor fault biases. This model residual sequence differs physically from the residual sequence extracted using the moving median in the previous stage; its baseline is jointly determined by multiple external environmental variables.
[0183] Step 1104: Perform a secondary screening of the model residual sequence based on the pre-configured residual judgment criteria, and add data points with abnormal deviations to the corresponding type of singular value set, or remove data segments that are misjudged as abnormal from the corresponding singular value set based on correlation analysis, thereby updating and outputting the final singular value identification result.
[0184] Correlation analysis refers to the correlation between corresponding data segments in the monitoring time series and the environmental data of the dam.
[0185] This step is a multi-dimensional fallback mechanism for anomaly detection. It corrects potential missed or false positives that might occur when relying on a single time series analysis. First, the system calculates the residual standard deviation of the model's residual sequence and sets the residual tolerance limits.
[0186] The specific boundary calculation is expressed as: E limit =K*S;
[0187] Among them, E limit Here, K represents the allowable limit for residuals, K is the reliability multiplier, and S is the residual standard deviation. Typically, the reliability multiplier K used for dam safety monitoring data can be set to 3.
[0188] When the absolute value of the model residual of a certain data node exceeds the residual tolerance limit, the system determines that the data point has an abnormal deviation independent of the normal response model, and performs a supplementary screening action to add it to the corresponding singular value set.
[0189] On the other hand, for continuous outlier segments identified as drifting singularities in the previous processing, the system performs correlation analysis. The system extracts the measured value sequence within this continuous outlier segment and simultaneously extracts the upstream reservoir water level sequence for the corresponding time period, calculating the Pearson correlation coefficient between the two sets of sequences. When the absolute value of the correlation coefficient is greater than 0.7, objective physical laws indicate that the continuous numerical shift in this data segment exhibits a high degree of synchronous linkage with the rise and fall of the reservoir water level.
[0190] At this point, it was determined that the continuous shift was not caused by a malfunction in the monitoring instrument, but rather reflected the actual physical response of the engineering structure to changes in water level. Based on this physical judgment, the system identified it as a data segment that had been mistakenly identified as abnormal, proactively removed it from the drifting singularity set, and restored its valid data attributes. After the above two-way correction, addition, and removal judgments, the system outputs the final updated singularity identification results for subsequent use by the evaluation index module.
[0191] According to one aspect of this application, a monitoring time series collected by the safety monitoring facilities of the target dam within a set statistical period is obtained, and singular value identification is performed on the monitoring time series to obtain the singular value identification results, as follows:
[0192] In specific engineering applications, such as when the monitoring time series has a short time span and is not affected by seasonal cyclical changes in dam reservoir water levels or ambient temperature, the series itself does not contain significant non-stationary trends. Under the physical premise that such data exhibits a stationary distribution and does not contain systematic errors, multiple conventional statistical detection methods can be run in parallel to identify and verify outliers, replacing the aforementioned progressive identification pipeline that involves complex trend estimation. Specifically, the system can run box plot analysis and traditional standard score statistical analysis in parallel.
[0193] When using box plot analysis, the system extracts the upper and lower quartiles of the monitored time series and calculates the difference between them as the interquartile range. Based on this, the system sets upper and lower limits for judgment.
[0194] Limit upper =Q3 + 1.5 * IQR;
[0195] Limit lower =Q1-1.5*IQR;
[0196] Among them, Limit upper To determine the upper limit, Q3 is the upper quartile of the monitored time series, IQR is the interquartile range, and Limit is the upper limit. lower To determine the lower limit, Q1 is the lower quartile of the monitoring time series. The system iterates through the monitoring time series and identifies monitoring values outside the aforementioned upper and lower limits as outliers.
[0197] When using the traditional standard score statistical analysis method, the system first employs a distribution test algorithm to verify whether the single-point dam safety monitoring data follows a Gaussian distribution. Under the assumption of a Gaussian distribution, the system calculates the mean and standard deviation of the sequence variable, and for each data point, calculates the difference between its value and the mean, then divides this difference by the standard deviation of the variable sequence to obtain the corresponding standard score. The specific calculation formula is as follows:
[0198] Z score =(y i -y mean ) / y std ;
[0199] Among them, Z score Let y be the standard score for the i-th data point. i To monitor the actual value of the i-th data point in the time series, y mean To monitor the average value of a time series, y std To monitor the standard deviation of the time series.
[0200] According to the Raida criterion in statistics, the system selects data points with an absolute value of standard score greater than 3 as outliers, that is, extracts data nodes with standard scores greater than +3 and less than -3. This threshold condition corresponds to a 99.73% confidence level under a normal distribution.
[0201] After performing calculations using multiple conventional statistical detection methods in parallel, the system integrates the detection outputs from these methods. In some specific implementations, the system can be configured to take the union of the identification results from multiple methods as the final anomaly judgment set to capture suspicious isolated singular values, drifting singular values, and jump singular values to the greatest extent possible; or it can be configured to take the intersection to reduce the false positive rate, and finally output the integrated label set as the singular value identification result for subsequent operation evaluation index calculation modules to use.
[0202] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A method for rapid assessment of the operational status of safety monitoring facilities for small reservoir dams, characterized in that, include: Obtain the initial monitoring point status and on-site detection status of the target dam safety monitoring facilities, and determine the status of the monitoring point itself based on the initial monitoring point status or on-site detection status; The monitoring time series collected by the safety monitoring facilities of the target dam within a set statistical period is obtained, and the outlier identification results are obtained by performing outlier identification on the monitoring time series. Based on the monitoring time series and outlier identification results, the percentage of abnormal data, the current data stoppage duration, and the duration of missing valid data are calculated. The working status of the measurement points is determined based on preset threshold judgment rules, taking into account the percentage of abnormal data, the duration of current data suspension, and the duration of missing valid data. By combining the physical status of the measuring points with their operational status, the operational status evaluation results of the target dam safety monitoring facilities are determined and output.
2. The method according to claim 1, characterized in that, Singular value identification was performed on the monitored time series, and the results of the singular value identification were obtained, including: The monitoring time series was filtered based on physical constraints to obtain a preliminary purified sequence; Extract the local trend estimate of the preliminary cleanup sequence and calculate the residual sequence; The dispersion index of the residual sequence is calculated, and the residual sequence is detected using a double threshold mechanism to obtain the out-of-limit classification labeling results. Based on the results of the out-of-limit classification, a continuous analysis is performed to determine the singular value classification results. The singular value classification results are then combined with the singular values marked in the physical constraint filtering stage, and the combined results are used as the singular value identification results.
3. The method according to claim 2, characterized in that, Based on physical constraints, the monitoring time series is filtered to obtain a preliminary purified sequence, including: Based on the engineering design parameters of the target dam and the range information of the monitoring instruments used, set the upper and lower limits of the fixed threshold. Values exceeding the upper or lower limits in the monitored time series are removed and marked as outliers, and the remaining data series are used as the initial cleaned series.
4. The method according to claim 2, characterized in that, Extract the local trend estimate of the preliminary cleanup sequence and calculate the residual sequence, including: For the actual data points in the preliminary purification sequence, a local window covering the actual data point is selected based on the set window width; Calculate the median of the data within the local window and use it as a local trend estimate for the corresponding real data points; The difference between the initial purified sequence and the corresponding local trend estimate is calculated to obtain the residual sequence with the trend component removed.
5. The method according to claim 4, characterized in that, Before calculating the median of the data within the local window, an extension step is included to eliminate end-boundary distortion of the sequence: Extract the first and last data segments of the preliminary purification sequence respectively; Calculate the median of the point-by-point difference sequence within the first data segment, using it as the slope of the local trend on the left, and virtually extrapolate a preset number of data points to the outside of the first data segment based on the slope of the local trend on the left. Calculate the median of the point-by-point difference sequence within the tail data segment as the right-end local trend slope, and virtually extrapolate a preset number of data points to the outside of the tail data segment based on the right-end local trend slope. The data points obtained by virtual extrapolation are spliced with the preliminary purification sequence to form an extended sequence.
6. The method according to claim 2, characterized in that, The dispersion index of the residual sequence is calculated, and a dual-threshold mechanism is used to detect the residual sequence, obtaining the out-of-limit classification labeling results, including: Calculate the median absolute deviation of the residual sequence as a robust dispersion index; Based on the ratio of the deviation of each data point from the median of the residual to the robust dispersion index, and combined with the normal distribution correction coefficient, the corrected standard score of each data point is calculated. A hard threshold is preset to characterize absolute anomalies and a soft threshold is preset to characterize systematic shifts, wherein the hard threshold is greater than the soft threshold. The absolute value of the corrected standard score of each data point is compared with the hard threshold and the soft threshold respectively. Based on the comparison results, the data points in the residual sequence are assigned the corresponding hard over-limit label, soft over-limit label or normal label to obtain the over-limit classification label results.
7. The method according to claim 2, characterized in that, Based on the out-of-limit classification results, a continuity analysis is performed to determine the singular value classification results, including: In the results of the out-of-limit classification and marking, extract the data segments that are continuously composed of soft out-of-limit or hard out-of-limit markers, and count the number of data points contained in the extracted continuous out-of-limit segments; Preset drift detection length threshold and jump detection length threshold, wherein the drift detection length threshold is greater than the jump detection length threshold; The number of data points is compared with the drift detection length threshold and the jump detection length threshold respectively. Based on the comparison results and the label type contained in the continuous anomaly segment, the data points in the continuous anomaly segment are classified as drift-type singular values, jump-type singular values, or isolated singular values, thereby generating singular value classification results.
8. The method according to claim 7, characterized in that, Data points within consecutive outlier segments are classified as drifting singularities, jump singularities, or isolated singularities, specifically including: When the number of data points is greater than or equal to the drift detection length threshold, all data points in the corresponding continuous abnormal segment are determined to be drift-type singular values. When the number of data points is greater than or equal to the jump detection length threshold and less than the drift detection length threshold, all data points in the corresponding continuous abnormal segment are determined to be jump-type singular values. When the number of data points is less than the jump detection length threshold, only data points with hard out-of-limit markers within the corresponding continuous abnormal segment are identified as isolated singular values, and data points with soft out-of-limit markers are restored as normal data points.
9. The method according to claim 1, characterized in that, Based on the monitoring time series and outlier identification results, the following calculations are performed to obtain the percentage of data anomalies, the current data outage duration, and the duration of missing valid data: The total amount of data in the time series is monitored within a set statistical period. Based on the outlier identification results, the corresponding amount of abnormal data is counted, and the ratio of the amount of abnormal data to the total amount of data is calculated as the proportion of abnormal data. Extract the latest monitoring time from the monitoring time series, calculate the difference between the current assessment time and the latest monitoring time, and use it as the current data stoppage duration; Based on the monitoring time series and outlier identification results, the retention status of valid data for each time unit within the set statistical period is determined, and the time units where valid data is determined to be missing are accumulated to obtain the duration of valid data missing.
10. The method according to claim 9, characterized in that, For automated monitoring facilities, the retention status of valid data for each time unit within a set statistical period is determined, and the time units where valid data is determined to be missing are accumulated to obtain the duration of valid data loss, specifically including: Divide the set statistical period into multiple consecutive calendar days; Obtain the total number of measured data entries for the monitored time series within a single calendar day, and count the number of abnormal data entries within that calendar day based on the outlier identification results; Subtract the number of abnormal data entries from the total number of measured data entries to obtain the number of valid data entries for the corresponding calendar day; Obtain the pre-configured theoretical data collection frequency of the monitoring facilities, and determine the theoretical number of data entries to be collected per day accordingly; The ratio of valid data entries to the theoretical daily data collection count is used to obtain the valid data ratio. The effective data ratio is compared with the pre-configured minimum effective data ratio threshold. When the effective data ratio is lower than the minimum effective data ratio threshold, the corresponding calendar day is determined as a day with missing effective data. The effective data missing duration is obtained by summing up all the days with missing data identified within the set statistical period.
11. The method according to claim 1, characterized in that, Based on the percentage of abnormal data, the current duration of data suspension, and the duration of missing valid data, the working status of the measurement points is determined according to preset threshold judgment rules, including: The abnormal data ratio, current data outage duration, and valid data missing duration are compared with the pre-configured corresponding state classification thresholds. Based on the comparison results, the working status of the target dam safety monitoring facility's measuring points is classified into one of the following: normal state, unstable state, outage state, or failure state.