Safety monitoring data reliability evaluation method based on multi-stage pre-screening and feature fusion

CN121524818BActive Publication Date: 2026-09-25POWER CHINA KUNMING ENG CORP LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511570826.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-09-25
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

效率低下:大坝监测仪器数量庞大,历史数据浩如烟海,人工评判工作量巨大

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524818B_ABST
    Figure CN121524818B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of safety monitoring of water conservancy and hydropower engineering, and particularly relates to a safety monitoring data reliability evaluation method based on multi-stage pre-screening and feature fusion, which, through a three-stage pre-screening mechanism, sequentially checks recent data loss, overall data loss rate and preliminary abnormal value rate of instruments, quickly determines and screens out instruments that have failed or have poor data quality; for the instruments that pass the pre-screening, multi-dimensional deep features thereof are further extracted, and expert rule engines and machine learning models are fused for comprehensive intelligent evaluation, and finally a reliability conclusion is output. The present application realizes efficient, automatic and intelligent evaluation of the reliability of historical monitoring data, the evaluation result is objective and accurate, and can be continuously self-optimized through a man-machine cooperation mechanism, greatly improving the efficiency and intelligent level of dam safety monitoring management, and providing reliable technical support for scientific management and maintenance decision of the dam safety monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety monitoring technology for water conservancy and hydropower projects, and in particular to a method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion. Background Technology

[0002] The dam safety monitoring system is equipped with numerous monitoring instruments of various types to continuously collect data reflecting the operational status of the dam and its ancillary structures. The reliability of these instruments directly affects the validity of the monitoring data, and consequently, the accuracy of the dam safety assessment. Currently, the evaluation of the reliability of monitoring instruments mainly relies on human experience, involving visual inspection and judgment of historical data. This approach has the following drawbacks: Inefficient: The number of dam monitoring instruments is enormous, the historical data is vast, and the workload for manual evaluation is huge.

[0003] Highly subjective: the evaluation criteria are difficult to unify, and it relies heavily on the personal experience of experts, so different experts may reach different conclusions.

[0004] Not comprehensive enough: Manual analysis makes it difficult to deeply mine multi-dimensional information such as time sequence patterns and mutation characteristics hidden in the data.

[0005] High lag: It cannot achieve automated real-time evaluation and it is difficult to detect instrument failure in a timely manner.

[0006] Therefore, there is an urgent need for an automated, intelligent, objective, and efficient method to standardize the evaluation of the reliability of monitoring instruments. The purpose of this invention is to address the shortcomings of the existing technology by providing a method and system that can automatically, intelligently, efficiently, and accurately evaluate the reliability of safety monitoring data. This method rapidly identifies failed instruments through a multi-level pre-screening mechanism and utilizes multi-feature fusion and intelligent algorithms to perform in-depth analysis of the data from the pre-screened instruments, ultimately outputting a reliability conclusion. It also enables human-machine collaboration and self-iterative optimization. Summary of the Invention

[0007] To achieve the above objectives, the present invention provides the following technical solution: According to a first aspect of the present invention, the present invention claims protection for a method for evaluating the reliability of security monitoring data based on multi-level pre-screening and feature fusion, comprising the following steps: S1: Perform data acquisition and preprocessing, acquire historical time series data of the target monitoring instrument and its corresponding timestamps, and acquire the current system time; S2: Three-level unreliability pre-screening judgment: The historical time series data is sequentially checked for recent data missing, overall data missing rate, and preliminary outlier ratio according to priority. If any check meets the unreliability condition, the target monitoring instrument is directly determined to be unreliable and the process is terminated. S3: Multi-dimensional deep feature extraction. For historical time series data that has been pre-screened by S2, multi-dimensional deep features including statistical features, temporal regularity features, and abrupt change features are extracted to construct feature vectors. S4: Comprehensive intelligent reliability evaluation, based on the feature vector, uses a combination of rule engine and machine learning model to evaluate and generate reliability conclusions. The rule engine makes judgments based on a preset expert rule base, and the machine learning model is a pre-trained multi-classification model that outputs a probability distribution. S5: Retrain the machine learning model based on human feedback and iteratively optimize it to continuously improve model performance.

[0008] Furthermore, the recent data missing check in step S2 specifically includes: Calculate the time interval between the last valid data timestamp and the current system time; Determine whether the time interval is greater than a preset threshold, wherein the preset threshold is 5 to 10 days; If so, the target monitoring instrument is deemed unreliable.

[0009] Furthermore, the preliminary outlier ratio check in step S2 specifically includes: The interquartile range method is used to detect outliers in the historical time series data, and the ratio of the number of outlier points to the total number of data points is calculated. Determine whether the ratio is greater than a preset threshold, wherein the preset threshold is 0.4; If so, the target monitoring instrument is deemed unreliable.

[0010] Furthermore, the temporal regularity features extracted in step S3 include: Trend strength, periodicity strength, and noise level; The trend strength is obtained by fitting a nonlinear model and calculating the ratio of the residual variance to the original data variance; The periodic intensity is obtained by calculating the first significant peak of the autocorrelation function in the non-zero hysteresis region; The noise level is obtained by calculating the standard deviation of the detrended residual sequence.

[0011] Furthermore, the mutation point features extracted in step S3 include the number of mutation points and the average mutation magnitude; The number of mutation points was obtained using the PELT (Problem Point Detection) algorithm. The average mutation magnitude is obtained by calculating the average value of the data changes at the mutation point.

[0012] Furthermore, the evaluation by the rule engine in step S4 specifically includes: Based on the conditional statements in the expert rule base, the relationship between each feature value in the feature vector and the preset threshold is compared and combined to output the reliability conclusion of the rule engine. The expert rule base is constructed based on the evaluation specifications for dam safety monitoring systems.

[0013] Furthermore, the machine learning model evaluation in step S4 specifically includes: The feature vector is input into a pre-trained LightGBM classification model to obtain the probability distribution of the target monitoring instrument belonging to the categories of "reliable", "basically reliable" and "unreliable". The decision on whether to adopt the conclusion of the machine learning model is based on a comparison between the highest probability value of the probability distribution and the confidence threshold.

[0014] Furthermore, step S4 also includes: If the highest probability value output by the machine learning model is greater than the confidence threshold, then that category is directly used as the final reliability conclusion. If the highest probability value is lower than the confidence threshold, or if the conclusion of the rule engine conflicts with the conclusion of the machine learning, a manual judgment process will be initiated, and experts will make a decision and store the result in the sample database.

[0015] Furthermore, step S5 also includes: New rulings are periodically obtained from the manual review process and used as annotation samples; The labeled samples are used to update the training dataset, and the machine learning model is retrained and its parameters optimized to achieve adaptive learning.

[0016] This invention relates to the field of safety monitoring technology for water conservancy and hydropower projects, and particularly to a reliability evaluation method for safety monitoring data based on multi-level pre-screening and feature fusion. Through a three-level pre-screening mechanism, the method sequentially checks for recent data loss, overall data loss rate, and preliminary outlier ratio, quickly identifying and filtering instruments that are faulty or have extremely poor data quality. For instruments that pass the pre-screening, multi-dimensional deep features are further extracted, and a comprehensive intelligent evaluation is performed by integrating an expert rule engine and a machine learning model, ultimately outputting a reliability conclusion. This invention achieves efficient, automatic, and intelligent evaluation of the reliability of historical monitoring data. The evaluation results are objective and accurate, and the method can continuously self-optimize through a human-machine collaborative mechanism, greatly improving the efficiency and intelligence level of dam safety monitoring management, and providing reliable technical support for the scientific management and maintenance decision-making of dam safety monitoring systems. Attached Figure Description

[0017] Figure 1 The flowchart is for a safety monitoring data reliability evaluation method based on multi-level pre-screening and feature fusion, which is claimed in this invention. Figure 2 This is a schematic diagram of the multi-dimensional deep feature extraction and comprehensive intelligent evaluation process of a safety monitoring data reliability evaluation method based on multi-level pre-screening and feature fusion, which is claimed in this invention. Figure 3 This is a schematic diagram of test data for a safety monitoring data reliability evaluation method based on multi-level pre-screening and feature fusion, which is claimed in this invention. Figures 4-6 This is a schematic diagram illustrating the operational results of a safety monitoring data reliability evaluation method based on multi-level pre-screening and feature fusion, for which this invention is claimed. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] The terms "first," "second," and "third" used in this invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this invention are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the accompanying drawings). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0020] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0021] According to the first embodiment of the present invention, referring to Figure 1 This invention claims protection for a method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion, comprising the following steps: S1: Perform data acquisition and preprocessing, acquire historical time series data of the target monitoring instrument and its corresponding timestamps, and acquire the current system time; S2: Three-level unreliability pre-screening judgment: The historical time series data is sequentially checked for recent data missing, overall data missing rate, and preliminary outlier ratio according to priority. If any check meets the unreliability condition, the target monitoring instrument is directly determined to be unreliable and the process is terminated. S3: Multi-dimensional deep feature extraction. For historical time series data that has been pre-screened by S2, multi-dimensional deep features including statistical features, temporal regularity features, and abrupt change features are extracted to construct feature vectors. S4: Comprehensive intelligent reliability evaluation, based on the feature vector, uses a combination of rule engine and machine learning model to evaluate and generate reliability conclusions. The rule engine makes judgments based on a preset expert rule base, and the machine learning model is a pre-trained multi-classification model that outputs a probability distribution. S5: Retrain the machine learning model based on human feedback and iteratively optimize it to continuously improve model performance.

[0022] In this embodiment, historical time-series data of the target monitoring instrument is acquired. and its corresponding timestamp and obtain the current system time. ; Three judgment conditions are set in descending order of priority. If any condition is met, the instrument is directly determined to be unreliable and the process is terminated. The number of measurement points to be calculated by the instrument within a historical time period. and the actual number of measurement points Calculate the overall missing rate :

[0023] judge Is it greater than the preset threshold? ( (Values ​​range from 0.3 to 0.4). If so, the instrument is deemed unreliable because historical data is severely lacking and cannot accurately reflect the building's operational status.

[0024] Furthermore, the recent data missing check in step S2 specifically includes: Calculate the time interval between the last valid data timestamp and the current system time; Determine whether the time interval is greater than a preset threshold, wherein the preset threshold is 5 to 10 days; If so, the target monitoring instrument is deemed unreliable.

[0025] In this embodiment, the last valid data timestamp is calculated. With current time time interval :

[0026] judge Is it greater than the preset threshold? ( (The value is 5-10 days). If so, the instrument is considered unreliable because there is no recent data, and it is considered to have failed.

[0027] Furthermore, the preliminary outlier ratio check in step S2 specifically includes: The interquartile range method is used to detect outliers in the historical time series data, and the ratio of the number of outlier points to the total number of data points is calculated. Determine whether the ratio is greater than a preset threshold, wherein the preset threshold is 0.4; If so, the target monitoring instrument is deemed unreliable.

[0028] In this embodiment, the interquartile range (IQR) method is used to analyze the actual data. Perform preliminary outlier detection:

[0029]

[0030]

[0031]

[0032]

[0033] In the above formula, The function is used to calculate quantiles. express 25% of the values ​​are less than , express 75% of the values ​​are less than , Interquartile range, and These represent the lower and upper limits of the normal value, respectively.

[0034] Statistics are not available Number of data points within the range And calculate the ratio :

[0035] judge Is it greater than the preset threshold? ( (The value is 0.4). If so, the instrument is deemed unreliable because the percentage of gross errors in the data is too high, and it cannot accurately reflect the building's operating status.

[0036] Furthermore, the temporal regularity features extracted in step S3 include: Trend strength, periodicity strength, and noise level; The trend strength is obtained by fitting a nonlinear model and calculating the ratio of the residual variance to the original data variance; The periodic intensity is obtained by calculating the first significant peak of the autocorrelation function in the non-zero hysteresis region; The noise level is obtained by calculating the standard deviation of the detrended residual sequence.

[0037] Furthermore, the mutation point features extracted in step S3 include the number of mutation points and the average mutation magnitude; The number of mutation points was obtained using the PELT (Problem Point Detection) algorithm. The average mutation magnitude is obtained by calculating the average value of the data changes at the mutation point.

[0038] In this embodiment, feature vectors for deep analysis are extracted from the data pre-screened in stage S2. ,include: Statistical characteristics: variance skewness kurtosis .

[0039] Temporal regularity characteristics: Trend strength : ,in Represents the original data The variance of the residuals after nonlinear fitting depends on the type of instrument used, such as Gaussian model, logarithmic model, trigonometric function, multiple regression, etc.

[0040] Periodic intensity : Calculate the first significant peak of the autocorrelation function in the non-zero lag region.

[0041] noise level : Calculate the standard deviation of the residual sequence after detrending.

[0042] Mutation point characteristics: The number of mutation points in a sequence is detected using the PELT (Prototype Exception Language) algorithm. and average mutation magnitude ; Reference Figure 2 This is a schematic diagram of the multi-dimensional deep feature extraction and comprehensive intelligent evaluation process in this embodiment.

[0043] Furthermore, the evaluation by the rule engine in step S4 specifically includes: Based on the conditional statements in the expert rule base, the relationship between each feature value in the feature vector and the preset threshold is compared and combined to output the reliability conclusion of the rule engine. The expert rule base is constructed based on the evaluation specifications for dam safety monitoring systems.

[0044] Furthermore, the machine learning model evaluation in step S4 specifically includes: The feature vector is input into a pre-trained LightGBM classification model to obtain the probability distribution of the target monitoring instrument belonging to the categories of "reliable", "basically reliable" and "unreliable". The decision on whether to adopt the conclusion of the machine learning model is based on a comparison between the highest probability value of the probability distribution and the confidence threshold.

[0045] Furthermore, step S4 also includes: If the highest probability value output by the machine learning model is greater than the confidence threshold, then that category is directly used as the final reliability conclusion. If the highest probability value is lower than the confidence threshold, or if the conclusion of the rule engine conflicts with the conclusion of the machine learning, a manual judgment process will be initiated, and experts will make a decision and store the result in the sample database.

[0046] In this embodiment, based on feature vectors The evaluation is conducted using a combination of rule engines and machine learning models. Rule engine evaluation: Judgments are based on a pre-set expert rule base. Rules consist of a series of conditional statements. According to relevant standards for dam safety monitoring system evaluation, the relationship between various feature values ​​and their limiting thresholds is compared to assess data reliability. For example: "If..." and "Then the data is rated as reliable."

[0047] Machine learning model evaluation: This involves evaluating the feature vectors. The input is fed into the pre-trained classification model LightGBM to obtain the probability distribution of its belonging to the three categories of "reliable", "mostly reliable", and "unreliable". .

[0048] If the machine learning model outputs the highest probability Greater than the confidence threshold If so, then that category will be used as the final conclusion.

[0049] If the highest probability is lower than the confidence threshold, or if the conclusion of the rule engine conflicts with the conclusion of the machine learning, the instrument data and all feature analysis results will be pushed to the manual judgment interface for experts to make the final decision, and the decision will be stored in the sample database.

[0050] Furthermore, step S5 also includes: New rulings are periodically obtained from the manual review process and used as annotation samples; The labeled samples are used to update the training dataset, and the machine learning model is retrained and its parameters optimized to achieve adaptive learning.

[0051] In this embodiment, the machine learning model is periodically retrained using newly added manually judged sample data to continuously optimize the model's performance.

[0052] The algorithm described in this invention was implemented in MATLAB. A total of 5 sets of data were tested, as shown in the following figures. Figure 3 As shown, after running the code, the result is as follows: Figure 4-6 As shown; The above results show that the method described in this invention can accurately assess the reliability of various types of data and provides detailed assessment criteria.

[0053] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0054] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

[0055] The specific embodiments of the invention have been described in detail above, but these are merely examples, and the invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this invention. Therefore, all equivalent transformations, modifications, and improvements made without departing from the spirit and principles of this invention should be included within the scope of this invention.

Claims

1. A method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion, characterized in that, Includes the following steps: S1: Perform data acquisition and preprocessing, acquire historical time series data of the target monitoring instrument and its corresponding timestamps, and acquire the current system time; S2: Three-level unreliability pre-screening judgment: The historical time series data is sequentially checked for recent data missing, overall data missing rate, and preliminary outlier ratio according to priority. If any check meets the unreliability condition, the target monitoring instrument is directly determined to be unreliable and the process is terminated. S3: Multi-dimensional deep feature extraction. For historical time series data that has been pre-screened by S2, multi-dimensional deep features including statistical features, temporal regularity features, and abrupt change features are extracted to construct feature vectors. S4: Comprehensive intelligent reliability evaluation, based on the feature vector, uses a combination of rule engine and machine learning model to evaluate and generate reliability conclusions. The rule engine makes judgments based on a preset expert rule base, and the machine learning model is a pre-trained multi-classification model that outputs a probability distribution. S5: Retrain the machine learning model based on human feedback and iteratively optimize it to continuously improve model performance.

2. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The recent data missing check in step S2 specifically includes: Calculate the time interval between the last valid data timestamp and the current system time; Determine whether the time interval is greater than a preset threshold, wherein the preset threshold is 5 to 10 days; If so, the target monitoring instrument is deemed unreliable.

3. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The preliminary outlier ratio check in step S2 specifically includes: The interquartile range method is used to detect outliers in the historical time series data, and the ratio of the number of outlier points to the total number of data points is calculated. Determine whether the ratio is greater than a preset threshold, wherein the preset threshold is 0.4; If so, the target monitoring instrument is deemed unreliable.

4. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The temporal regularity features extracted in step S3 include: Trend strength, periodicity strength, and noise level; The trend strength is obtained by fitting a nonlinear model and calculating the ratio of the residual variance to the original data variance; The periodic intensity is obtained by calculating the first significant peak of the autocorrelation function in the non-zero hysteresis region; The noise level is obtained by calculating the standard deviation of the detrended residual sequence.

5. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The mutation point features extracted in step S3 include the number of mutation points and the average mutation magnitude. The number of mutation points was obtained using the PELT (Problem Point Detection) algorithm. The average mutation magnitude is obtained by calculating the average value of the data changes at the mutation point.

6. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The evaluation by the rule engine in step S4 specifically includes: Based on the conditional statements in the expert rule base, the relationship between each feature value in the feature vector and the preset threshold is compared and combined to output the reliability conclusion of the rule engine. The expert rule base is constructed based on the evaluation specifications for dam safety monitoring systems.

7. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, The machine learning model evaluation in step S4 specifically includes: The feature vector is input into the pre-trained LightGBM classification model to obtain the probability distribution of the target monitoring instrument belonging to the categories of "reliable", "basically reliable" and "unreliable". The decision on whether to adopt the conclusion of the machine learning model is based on a comparison between the highest probability value of the probability distribution and the confidence threshold.

8. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 7, characterized in that, Step S4 further includes: If the highest probability value output by the machine learning model is greater than the confidence threshold, then that category is directly used as the final reliability conclusion. If the highest probability value is lower than the confidence threshold, or if the conclusion of the rule engine conflicts with the conclusion of the machine learning, a manual judgment process will be initiated, and experts will make a decision and store the result in the sample database.

9. The method for evaluating the reliability of safety monitoring data based on multi-level pre-screening and feature fusion according to claim 1, characterized in that, Step S5 also includes: New rulings are periodically obtained from the manual review process and used as annotation samples; The labeled samples are used to update the training dataset, and the machine learning model is retrained and its parameters optimized to achieve adaptive learning.

Citation Information

Patent Citations

  • A Credit Reporting Evaluation System Based on Mixed Machine Learning

    AU2019100968A4

  • Bus load prediction method and system based on multi-scale time sequence convolutional neural network

    CN116805173A