A method and system for detecting measurement abnormal data based on a rule engine

Through the detection method of measurement abnormal data based on the rules engine, the abnormal points in the grid measurement data are identified and corrected in real time, and the problems of insufficient flexibility and low processing efficiency in the prior art are solved, and efficient and accurate data processing and abnormal detection are achieved.

CN119807985BActive Publication Date: 2025-06-24STATE GRID ANHUI ELECTRIC POWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510302400.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-24
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing technology lacks flexibility and adaptability in dynamically changing data environments, resulting in fixed rules being unable to adapt to emerging exceptions in a timely manner, and requires a long processing time when processing large-scale data, resulting in delay problems.

Method used

Through the detection method of measuring abnormal data based on the rules engine, the transmission path and timestamp sequence are extracted, the abnormality method with the fitting degree of risk path exceeding the preset threshold is identified, the abnormal flow feature set is obtained, and the abnormal points in the data are identified and corrected in real time through the dynamic abnormality threshold and data calibration set.

Benefits of technology

It significantly improves the response speed and accuracy of data processing, enhances the ability to capture abnormal trends, reduces data redundancy and error repetition, and ensures data quality and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807985B_ABST
    Figure CN119807985B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of anomaly detection, and specifically to a method and system for detecting measurement anomaly data based on a rule engine, including the following steps: Based on the source point number, transmission link number, and real-time status code of the measurement data packet, extract the transmission path and timestamp sequence, compare the data flow path with the preset abnormal flow direction, identify the abnormal mode with the risk path fitting degree exceeding the preset threshold, and obtain the abnormal flow direction feature set. In the present invention, by monitoring the changes in the submission frequency and activity of the real-time measurement data of the power grid, as well as the data synchronization, it is possible to effectively identify and correct the abnormal points in the data. This dynamic abnormal threshold setting and real-time data analysis further enhance the ability to capture abnormal trends, can timely adjust the processing strategy, avoid the accumulation and spread of incorrect data, the generation of the data calibration set and the application of data fingerprints effectively reduce data redundancy and the repetition of errors, and ensure data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and particularly to a method and system for detecting measurement anomaly data based on a rule engine. Background Art

[0002] The technical field of anomaly detection includes methods and systems for identifying outliers or abnormal patterns in a dataset. The core content is to analyze the normal behavior patterns of data through various algorithms and technical means, and detect abnormal data points that deviate from the patterns. Anomaly detection is widely applied in multiple industries, such as financial fraud detection, network security, health monitoring, power grid measurement, and industrial systems. In this field, the quality of data directly affects the accuracy and reliability of anomaly detection, and efficient data processing and analysis technologies are required to process and optimize a large amount of data.

[0003] Among them, the method for detecting measurement anomaly data based on a rule engine refers to using predefined rules to identify abnormal data in a real-time power grid measurement dataset. This method sets a series of rules to judge whether the data is abnormal according to specific business logics and data characteristics. The design of the rule engine includes the detection of problems such as missing data acquisition points, abnormal data values, and data duplication, and performs data verification based on clear rule conditions and thresholds, without involving complex data processing or model adjustment. Through the rule engine, anomalies can be directly identified and processed at the data input stage, improving the efficiency and accuracy of data processing.

[0004] The prior art mainly relies on predefined rules and static thresholds to identify anomalies in actual operations, lacking flexibility and adaptability in a dynamically changing data environment. For example, in the fields of financial fraud detection or network security, fixed rules cannot adapt to newly emerging fraud means or attack patterns in a timely manner, resulting in detection failures or increased false alarm rates. When dealing with large-scale data, the existing methods require a long processing time, causing latency problems in application scenarios that require quick responses, affecting the overall system performance and user experience. Therefore, this processing mode relying on fixed rules faces the dual challenges of poor adaptability and low processing efficiency in a modern data-driven environment. Summary of the Invention

[0005] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a method and system for detecting measurement anomaly data based on a rule engine.

[0006] To achieve the above purpose, the present invention adopts the following technical solution. A method for detecting measurement anomaly data based on a rule engine includes the following steps:

[0007] S1: Based on the source point number, transmission link number, and real-time status code of the measurement data packet, extract the transmission path and timestamp sequence, compare the data flow path with the preset abnormal flow, identify the abnormal modes with the risk path fitting degree exceeding the preset threshold, and obtain the abnormal flow feature set;

[0008] S2: Based on the abnormal flow feature set, monitor the active status of the measurement collection points, identify the number of data submissions and the change in activity, analyze the data synchronization between adjacent points, extract the associated point information, identify the abnormal points, and form an abnormal point fluctuation graph;

[0009] S3: Use the abnormal point fluctuation graph to monitor the data flow, set a dynamic abnormal threshold, compare the deviation of the real-time data, analyze the fluctuation amplitude, and obtain an abnormal trend identification table;

[0010] S4: Use the abnormal trend identification table to extract the real-time data of the neighboring points, compensate for the continuous abnormal points, mark the measurement collection points in the cooling state, and generate a data calibration set;

[0011] S5: Through the data calibration set, combined with the ID, timestamp, and key fields of the collection points, extract the data fingerprint, compare the cached fingerprint, eliminate duplicate data, and adjust the sampling frequency to obtain the intelligent processing result of the measurement data.

[0012] As a further solution of the present invention, the abnormal flow feature set includes paths with excessive flow deviation, points with abnormal transmission time intervals, and abnormal modes of risk path matching. The abnormal point fluctuation graph includes the active state fluctuation trend, the data submission frequency change trend, and the abnormal distribution of data synchronization between adjacent points. The abnormal trend identification table includes points with excessive data fluctuation amplitude, points with abnormal change rate, and points with data distribution deviation exceeding the threshold. The data calibration set includes repaired data records, compensated abnormal point data, and marked points in the cooling state. The intelligent processing result of the measurement data includes the statistics of the duplicate data ratio, the sampling frequency adjustment strategy, and the device self-check trigger record.

[0013] As a further solution of the present invention, the specific steps for obtaining the abnormal flow feature set are as follows:

[0014] S111: Based on the source point number, transmission link number, and real-time status code of the measurement data packet, extract the transmission path and timestamp sequence, and establish a data flow sequence;

[0015] S112: Use the data flow sequence to compare with the preset abnormal flow, calculate the flow deviation degree, identify the transmission time interval between adjacent nodes, perform deviation calculation with the standard time interval, extract the packet numbers and paths exceeding the time threshold range, establish the corresponding relationship between the path deviation degree and the time deviation, and obtain the set of path deviation and time deviation;

[0016] S113: Call the set of path deviation and time deviation, evaluate the correlation between the path deviation degree and the time deviation, and use the formula:

[0017] ;

[0018] Calculate the risk path fitness, compare with the threshold value, screen the abnormal modes, extract the features of the abnormal modes, and obtain the abnormal flow direction feature set;

[0019] Wherein, is the risk path fitness, is the path deviation degree, and are weight parameters, is the time deviation.

[0020] As a further solution of the present invention, the steps for obtaining the abnormal point position fluctuation diagram are specifically as follows:

[0021] S211: Based on the abnormal flow direction feature set, monitor the active state of the measurement and collection points, count the number of data submissions of each measurement and collection point within a fixed time window, identify the submission frequency, calculate the change rate of the point submission frequency, call the set submission frequency change threshold, and determine whether the activity of the point has mutated to obtain the activity change record;

[0022] S212: Through the activity change record, analyze the synchronization of data of adjacent points, compare the number of submissions of adjacent points within the same time window, calculate the time deviation value of adjacent points, and obtain the adjacent point time deviation identification result;

[0023] S213: Utilize the adjacent point time deviation identification result, extract the associated point information, screen the set of points whose time deviation value exceeds the set threshold, identify the distribution range of the measurement and collection points, and use the formula:

[0024] ;

[0025] Calculate the aggregation degree of abnormal points, screen the set of abnormal points, and obtain the abnormal point position fluctuation diagram;

[0026] Wherein, represents the aggregation degree of abnormal points, represents the th spatial coordinate of the abnormal point, represents the mean coordinate of the abnormal point, represents the number of abnormal points.

[0027] As a further solution of the present invention, the steps for obtaining the abnormal trend identification table are specifically as follows:

[0028] S311: Based on the abnormal point fluctuation graph, parse multiple time series intervals of the monitored data stream, extract the fluctuation amplitude, change rate, and data distribution characteristics within each interval, calculate the mean, standard deviation, and change rate of the fluctuation amplitude, and obtain the fluctuation characteristic coefficient by comparing the mean difference and standard deviation change trend of adjacent intervals;

[0029] S312: Use the fluctuation characteristic coefficient, combined with the set dynamic abnormal threshold, to analyze the deviation of real-time data, and use the formula:

[0030] ;

[0031] Calculate the data deviation degree. When exceeds the set threshold, mark the data point as an abnormal fluctuation point, and count the proportion of multiple abnormal fluctuation points in different time intervals to obtain the deviation degree quantization value;

[0032] Among them, represents the data deviation degree at the real-time moment , is the real-time data value, is the mean of the previous interval, is the standard deviation of the previous interval, is a constant, is the fluctuation weight factor;

[0033] S313: Combine the time distribution of abnormal fluctuation points, calculate the change rate of abnormal fluctuation trends in multiple time periods, and identify the deviation degree of abnormal distribution. Use the deviation degree quantization value and change amplitude for evaluation to obtain the abnormal trend identification table.

[0034] As a further solution of the present invention, the steps for obtaining the data calibration set are specifically as follows:

[0035] S411: Use the abnormal trend identification table to extract the real-time data of neighboring points, perform continuity detection on the real-time data, screen out abnormal data points, and identify the data fluctuation trend based on the data change amplitude and time series to obtain an abnormal data point sequence;

[0036] S412: Through the abnormal data point sequence, repair the abnormal points, identify the error range of the repaired data, and correct the data points whose deviation exceeds the error range. Use the formula:

[0037] ;

[0038] Calculate the data correction deviation degree to obtain the data repair record;

[0039] Among them, represents the data correction deviation degree, represents the The repair data value of a point, representing the repair data value of the th point, representing the data mean value, representing the total number of data points;

[0040] S413: Using the data repair record, compensate for continuous abnormal points, calculate the data recovery rate, and mark the measurement acquisition points in the cooling state to generate a data calibration set.

[0041] As a further solution of the present invention, the steps for obtaining the intelligent processing result of the measurement data are specifically as follows:

[0042] S511: Through the data calibration set, combine the ID, timestamp, and keyword fields of the acquisition points to extract data fingerprints, calculate hash values, compare with cached fingerprints, eliminate duplicate data, and obtain a set of filtered data fingerprints;

[0043] S512: Using the set of filtered data fingerprints, calculate the ratio of the total data volume before and after filtering to obtain the duplicate data ratio, compare with the set threshold, determine whether to adjust the sampling frequency. If it exceeds the set threshold, adjust the frequency, trigger device self-check, update the sampling parameters, and obtain the adjusted sampling frequency parameters;

[0044] S513: Based on the adjusted sampling frequency parameters, combine the device self-check record, identify the error range, analyze the error change, compare with the error threshold, filter the affected measurement data, and obtain the intelligent processing result of the measurement data.

[0045] The measurement abnormal data detection system based on the rule engine is used to execute the above-mentioned measurement abnormal data detection method based on the rule engine. The system includes:

[0046] The data flow analysis module obtains the source point number, transmission link number, and status code of the measurement data packet, calculates the flow deviation degree, filters the data packets exceeding the flow deviation threshold, identifies the transmission time interval of adjacent nodes, calculates the time interval deviation value, and obtains the abnormal flow feature set;

[0047] The abnormal point monitoring module uses the abnormal flow feature set to count the submission times and activity change rate of the measurement acquisition points, compares the data submission times of adjacent points, extracts associated point information, and obtains the abnormal point fluctuation diagram;

[0048] The data comparison module uses the abnormal point fluctuation diagram to set a dynamic abnormal threshold, calculates the data offset and compares it with the threshold, identifies the fluctuation amplitude and change situation, and obtains the abnormal trend identification table;

[0049] The offset degree recognition module uses the abnormal trend identification table to extract the real-time data of neighborhood points, identify the offset degree of data distribution, analyze the trend change, compensate for continuous abnormal points, and obtain a data calibration set;

[0050] The data evaluation module uses the data calibration set to compare the data fingerprint and the cache fingerprint, eliminate duplicate data, adjust the sampling frequency, trigger device self-check, and obtain the intelligent processing result of measurement data.

[0051] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0052] In the present invention, the abnormal detection is optimized by accurately tracking the real-time measurement data flow of the power grid and monitoring the active state. The risk path is determined by using the data flow deviation degree and the transmission time interval deviation to identify potential abnormal data streams in real time. This real-time monitoring mechanism significantly improves the response speed and accuracy of data processing. By monitoring the changes in the data submission times and activity, as well as the data synchronization, it can effectively identify and correct abnormal points in the data. This dynamic abnormal threshold setting and real-time data analysis further enhance the ability to capture abnormal trends, can timely adjust the processing strategy, and avoid the accumulation and spread of incorrect data. The generation of the data calibration set and the application of the data fingerprint effectively reduce data redundancy and the repetition of errors, ensure data quality, and bring higher data integrity and reliability to abnormal detection based on the application of detailed data flow monitoring and real-time processing logic. Brief Description of the Drawings

[0053] Figure 1 is a schematic diagram of the working process of the present invention;

[0054] Figure 2 is a flowchart of obtaining the abnormal flow feature set in the present invention;

[0055] Figure 3 is a flowchart of obtaining the abnormal point fluctuation diagram in the present invention;

[0056] Figure 4 is a flowchart of obtaining the abnormal trend identification table in the present invention;

[0057] Figure 5 is a flowchart of obtaining the data calibration set in the present invention;

[0058] Figure 6 is a flowchart of obtaining the intelligent processing result of measurement data in the present invention. Detailed Embodiment

[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0060] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.

[0061] Embodiment 1: Please refer to Figure 1 , the present invention provides a technical solution, a method for detecting measurement abnormal data based on a rule engine, including the following steps:

[0062] S1: Based on the source point number, transmission link number and real-time status code of the measurement data packet, extract the transmission path and timestamp sequence, record the node information in the measurement data acquisition, transmission and processing links, compare the data flow path with the preset abnormal flow, calculate the flow deviation degree, identify the deviation between the transmission time interval of adjacent nodes and the standard time interval, identify the abnormal modes with the risk path fitting degree exceeding the preset threshold, and obtain the abnormal flow feature set;

[0063] S2: Based on the abnormal flow feature set, monitor the active status of the measurement acquisition points, identify the number of data submissions and the change in activity, analyze the data synchronization between adjacent points, extract the associated point information, identify the abnormal points, and form an abnormal point fluctuation graph;

[0064] S3: Use the abnormal point fluctuation graph to monitor the data flow, set a dynamic abnormal threshold, compare the deviation of the real-time data, analyze the fluctuation amplitude, change rate and data distribution deviation degree, and obtain an abnormal trend identification table;

[0065] S4: Use the abnormal trend identification table to extract the real-time data of the neighboring points, compensate the continuous abnormal points, mark the measurement acquisition points in the cooling state, and generate a data calibration set;

[0066] S5: Through the data calibration set, combine the ID, timestamp and key fields of the acquisition points, extract the data fingerprint, compare the cached fingerprint, eliminate the duplicate data, count the duplicate data ratio, adjust the sampling frequency and trigger the device self-check, and obtain the intelligent processing result of the measurement data;

[0067] The abnormal flow feature set includes paths with excessive flow deviation, points with abnormal transmission time intervals, and abnormal ways of matching risk paths. The abnormal point fluctuation graph includes the fluctuation trend of the active state, the change trend of the data submission frequency, and the abnormal distribution of data synchronization between adjacent points. The abnormal trend identification table includes points with excessive data fluctuation amplitude, points with abnormal change rate, and points with data distribution deviation exceeding the threshold. The data calibration set includes repaired data records, data compensation for abnormal points, and points marked with the cooling state. The intelligent processing result of measurement data includes the statistics of the duplicate data ratio, the sampling frequency adjustment strategy, and the device self-check trigger record.

[0068] Please refer to Figure 2 , and the steps for obtaining the abnormal flow feature set are specifically as follows:

[0069] S111: Based on the source point number, transmission link number, and real-time status code of the measurement data packet, extract the transmission path and timestamp sequence, and establish a data flow sequence.

[0070] It is necessary to extract the source point numbers from multiple data packets, arrange them in chronological order to construct an initial data flow sequence. During the extraction process, store the source point numbers of each data packet, and sequentially read the source numbers of subsequent data packets, sort them according to the timestamp order. At the same time, extract the transmission link numbers and identify the association between the source point number and the target receiving point number. Determine the order of the nodes passed by each transmission path through cross-comparison, and establish a complete transmission path sequence. During the construction process, for the situation where multiple data packets are sent to the same target node simultaneously, it is necessary to distinguish the data packets according to the arrival time through timestamp sorting, and number each path to ensure the correct path order. The real-time status code is used to identify the transmission status of the data packet, including normal transmission, loss, delay, retransmission, etc. When constructing the data flow sequence, it is necessary to associate and store the status code of the data packet with the transmission path. For data packets with abnormal status, specifically mark the status and record the paths experienced, including the transmission path, transmission link, and timestamp of each data packet, and be able to accurately reflect the transfer of data between different transmission nodes to establish a data flow sequence.

[0071] S112: Use the data flow sequence to compare with the preset abnormal flow, calculate the flow deviation degree, identify the transmission time interval between adjacent nodes, calculate the deviation from the standard time interval, extract the data packet numbers and paths that exceed the time threshold range, establish the corresponding relationship between the path deviation degree and the time deviation, and obtain the set of path deviation and time deviation.

[0072] It is necessary to set preset abnormal flow directions, including abnormal path combinations and the abnormal transmission behaviors of data packets that occur. By traversing all data packet paths in the data flow sequence and comparing them with the preset abnormal flow directions, for the matching paths, calculate the flow deviation degree, calculate the number of node differences between the path of the data packet in the current data flow sequence and the normal path, and at the same time consider the hop deviation degree of the path, and express the path deviation degree in percentage form. For the case where the path difference exceeds the set deviation threshold, mark the data packet path as an abnormal path and record the deviation degree value. Analyze the transmission time interval between adjacent nodes, calculate the transmission time difference of each data packet between two adjacent nodes, and the calculation method is the receiving time minus the sending time, obtain the time interval value between each transmission node, and compare it with the standard time interval. The standard time interval is obtained according to the transmission data statistics, and the average value and standard deviation can be taken as the reference values. For the data packets that exceed the set standard time interval, extract the number and path information, and classify and store them according to the degree of exceeding the time threshold range, establish the corresponding relationship between the path deviation degree and the time deviation, and form the path deviation and time deviation set.

[0073] S113: Call the path deviation and time deviation set to evaluate the correlation between the path deviation degree and the time deviation, using the formula:

[0074] ;

[0075] Calculate the risk path fitting degree, compare with the threshold value, screen the abnormal methods, extract the characteristics of the abnormal methods, and obtain the abnormal flow direction feature set;

[0076] Among them, is the risk path fitting degree, is the path deviation degree, and are weight parameters, is the time deviation;

[0077] Parameter meaning:

[0078] Path deviation degree : Set that the normal path contains 10 nodes, and the node deviation of the abnormal path is 3, then calculate the path deviation degree:

[0079] ;

[0080] Time deviation : Set that the average value of the standard transmission time interval of the normal path is 50ms, the standard deviation is 10ms, and the average time deviation of the abnormal path is 20ms, then the normalization calculation of the time deviation is:

[0081] ;

[0082] Set weight parameters and , calculate the risk path fitness degree:

[0083] ;

[0084] Abnormality determination: Set the risk path fitness threshold , and compare the calculation results:

[0085] ;

[0086] Therefore, this path does not belong to the abnormal path;

[0087] The results show that under the set conditions, the deviation degree and time deviation of the path are still within the acceptable range and do not exceed the abnormal threshold. If the deviation degree or time deviation of the path increases, it will cause the value to exceed 0.5, and it is marked as an abnormal path.

[0088] Please refer to Figure 3 , and the specific steps for obtaining the abnormal point fluctuation diagram are as follows:

[0089] S211: Based on the abnormal flow feature set, monitor the active status of the measurement collection points, count the number of data submissions of each measurement collection point within a fixed time window, identify the submission frequency, calculate the change rate of the point submission frequency, call the set submission frequency change threshold, and determine whether the activity of the point has mutated to obtain the activity change record;

[0090] To monitor the active status of the measurement collection points, it is necessary to obtain the data submission records of all points within a specific time window. Set one hour as the time window, count the number of data submissions of each point within this time period. For different types of monitoring points, such as temperature monitoring points, voltage monitoring points, flow monitoring points, etc., their data submission frequencies are different. Therefore, it is necessary to calculate the data submission frequencies of the points separately. Set to represent the submission frequency of point within this time window. For multiple points in a certain area, calculate the historical mean and variance of the submission frequencies of all points. Compare the current submission frequency of each point with its historical mean and calculate the change rate:

[0091] ;

[0092] If the change rate exceeds the set submission frequency change threshold , it is determined that the activity of this point has mutated. For a monitoring point of an industrial device, the average historical submission frequency is 10 times per hour, and the variance is 2 times per hour. If the submission frequency in a current period rises to 15 times per hour, the change rate is calculated as follows:

[0093] ;

[0094] If the set threshold is 0.3 and the change rate of this point is greater than the threshold, then this point is determined to have abnormal activity, and an activity change record is obtained.

[0095] S212: Through the activity change record, analyze the synchronization of data at adjacent points, compare the submission times of adjacent points within the same time window, calculate the time deviation value of adjacent points, and obtain the time deviation recognition result of adjacent points;

[0096] Analyze the synchronization of data at adjacent points, compare the data submission situations of multiple adjacent points in the same area within the same time window, calculate the time deviation values of each adjacent point. For two adjacent points and , the times of submitting data are respectively recorded as , and the time deviation is defined as:

[0097] ;

[0098] Statistically calculate the average value and variance of the time deviations of all adjacent points in the entire area, call the set synchronization deviation threshold for judgment. If the time deviation of a certain point and its adjacent point exceeds the set threshold , then it is determined that the data synchronization of this point is abnormal. Set the synchronization deviation threshold seconds. For points and , their data submission times are 100 seconds and 120 seconds respectively, then the time deviation is:

[0099] ;

[0100] Since 20 seconds is greater than the threshold of 10 seconds, this pair of points is determined to be an abnormal synchronization point, and the time deviation recognition result of adjacent points is obtained.

[0101] S213: Using the time deviation recognition result of adjacent points, extract the information of associated points, screen the set of points whose time deviation values exceed the set threshold, identify the distribution range of measurement and collection points, and use the formula:

[0102] ;

[0103] Calculate the aggregation degree of abnormal points, screen the set of abnormal points, and obtain the abnormal point fluctuation diagram;

[0104] Among them, represents the aggregation degree of abnormal points, represents the th spatial coordinate of the abnormal point, represents the mean coordinate of the abnormal points, represents the number of abnormal points;

[0105] Parameter acquisition method:

[0106] Spatial coordinates of abnormal points : By collecting data on each point in the monitoring area, identifying the abnormal points, and recording their spatial coordinates. For example, using a GPS device to measure the longitude and latitude coordinates of each abnormal point;

[0107] Average coordinate : Average the coordinates of all abnormal points. The calculation method is as follows:

[0108] ;

[0109] Among them, is the total number of abnormal points, is the th coordinate of the abnormal point;

[0110] Actual example:

[0111] Set in a certain monitoring area, 5 abnormal points are detected, and their coordinates (represented by two-dimensional plane coordinates) are respectively: ;

[0112] Calculate the average coordinate :

[0113] ;

[0114] Calculate the sum of the absolute values of the distances from each point to the average coordinate:

[0115] ;

[0116] ;

[0117] Calculate the sum of the squares of the distances from each point to the average coordinate:

[0118] ;

[0119] ;

[0120] Calculate the value of the denominator:

[0121] ;

[0122] Substitute into the formula to calculate the aggregation degree:

[0123] ;

[0124] The results show that the aggregation degree is 4.25. The abnormal points have a certain aggregation in space. The larger the value, the more dispersed the points are, and the smaller the value, the more concentrated the points are. Through this value, the distribution characteristics of the abnormal points can be further analyzed, providing a reference for subsequent monitoring and early warning.

[0125] Please refer to Figure 4 , and the specific steps for obtaining the abnormal trend identification table are as follows:

[0126] S311: Based on the abnormal point fluctuation diagram, analyze multiple time series intervals of the monitoring data stream, extract the fluctuation amplitude, change rate, and data distribution characteristics within each interval, and calculate the mean, standard deviation, and change rate of the fluctuation amplitude. By comparing the mean differences and standard deviation change trends between adjacent intervals, obtain the fluctuation characteristic coefficient;

[0127] Analyze each time series interval of the monitoring data stream. In actual implementation, the temperature monitoring data stream of a certain device can be analyzed. Set the device to collect temperature data every 5 seconds during continuous production. There are 720 data points per hour, and the daily data points are about 17,280. Divide it into multiple time series intervals, for example, one hour as an interval, extract the fluctuation amplitude, change rate, and data distribution characteristics of each interval, calculate the difference between the maximum and minimum values of all data points within the one-hour interval to obtain the fluctuation amplitude, which can reflect the change range of the data within this time interval. Set the highest temperature of a certain hour to 75°C and the lowest temperature to 65°C, then the fluctuation amplitude is 10°C. Calculate the change rate between two adjacent time points, that is:

[0128] ;

[0129] Among them, represents the value of the current data point, represents the value of the previous data point, represents the sampling time interval;

[0130] If , , the sampling time interval , then:

[0131] ;

[0132] Calculate the mean and standard deviation of all change rates within this interval to evaluate the stability of data fluctuations. Calculate the standard deviation:

[0133] ;

[0134] Among them, is the number of data points within the interval, is a single data point, is the mean of this interval. Set the mean temperature of a certain hour to 70°C, which contains 720 data points, and calculate its standard deviation to obtain:

[0135] ;

[0136] Compare the mean difference and the standard deviation change trend of adjacent intervals. Set the mean temperature of the previous interval to 68°C and the mean of the current interval to 70°C. Then the mean change is 2°C. If the standard deviation of the previous interval is 2.8°C and the standard deviation of the current interval is 3.2°C, then the standard deviation change is 0.4°C. Based on the comprehensive calculation results, obtain the fluctuation characteristic coefficient.

[0137] S312: Use the fluctuation characteristic coefficient and combine the set dynamic anomaly threshold to analyze the deviation of real-time data. Use the formula:

[0138] ;

[0139] Calculate the data deviation degree. When exceeds the set threshold, mark the data point as an abnormal fluctuation point, and count the proportion of multiple abnormal fluctuation points in different time intervals to obtain the quantified deviation value;

[0140] Among them, represents the deviation degree of the real-time moment of the data, is the real-time data value, is the mean of the previous interval, is the standard deviation of the previous interval, is a constant, is the fluctuation weight factor;

[0141] Detailed explanation of the formula and the derivation process of the formula calculation:

[0142] represents the data value at the current moment . The data is sourced from a real-time temperature sensor, with the unit set to degrees Celsius. Set the temperature of a certain device at the current moment to 76.5°C, which is read by the temperature sensor and collected every 5 seconds;

[0143] represents the mean value of the data in the previous interval, that is, the average temperature in the past period of time (such as the past 60 minutes). It is assumed that the sampling frequency in the past 60 minutes is once every 5 seconds, the total number of data points is 720, and the average value is obtained after statistical analysis of the temperature data in the past 60 minutes ;

[0144] represents the standard deviation of the data in the previous interval, and the standard deviation = 2.8;

[0145] To prevent the denominator from being a zero-valued decimal, it is generally set to to avoid numerical instability;

[0146] is the fluctuation weight factor, which represents the degree of influence of data point fluctuations. This factor is set according to the fluctuation amplitude of the data. If the temperature fluctuation amplitude in the previous interval is large, a high weight value is set; if the fluctuation is small, a low weight value is set, and it is set to ;

[0147] Substitute the parameters into the formula:

[0148] ;

[0149] The result shows that the deviation degree of the real-time data point is 2.035. If this value exceeds the set abnormal threshold (such as 1.8), the current data point is determined as an abnormal fluctuation point. By counting the proportion of multiple abnormal fluctuation points, the change situation of the abnormal trend can be further identified.

[0150] S313: Combine the time distribution of abnormal fluctuation points, calculate the change rate of abnormal fluctuation trends in multiple time periods, and identify the deviation degree of abnormal distribution. Use the deviation degree quantization value and the change amplitude for evaluation to obtain the abnormal trend identification table;

[0151] Combine the time distribution of abnormal fluctuation points, calculate the change rate of abnormal fluctuation trends in each time period, and calculate the deviation degree of abnormal distribution. Set the deviation degree quantization values per hour within the past 24 hours to be , , ;

[0152] By calculating the change rate within adjacent time intervals:

[0153] ;

[0154] Among them, is the deviation degree quantization value of the current time period, is the deviation degree quantization value of the previous time period is the time interval;

[0155] Set hours, calculate the change rate in a certain period. For example, the rate from the 4th hour (7.2%) to the 5th hour (8.0%) is:

[0156] ;

[0157] Calculate the deviation degree of the abnormal distribution. Use the data of the past N periods to calculate the mean value and the standard deviation . If the deviation value exceeds , it is determined that the abnormal fluctuation deviation in this period is large. Set the mean value of the past 10 hours , the standard deviation , then the abnormal determination threshold is:

[0158] ;

[0159] If the quantization value of the deviation degree of a certain hour exceeds 7.8%, the abnormal fluctuation deviation of this hour is large. Count the number of occurrences and time distribution of abnormal deviations in all time intervals to obtain the abnormal trend identification table.

[0160] Please refer to Figure 5 , the steps for obtaining the data calibration set are specifically as follows:

[0161] S411: Use the abnormal trend identification table to extract the real-time data of the neighborhood points, perform continuity detection on the real-time data, screen out the abnormal data points, and identify the data fluctuation trend according to the data change range and time series to obtain the abnormal data point sequence;

[0162] Extract the real-time data of the neighborhood points. For the data extraction process, obtain the real-time data streams of all measurement points in the target monitoring area, classify them according to the spatial distribution of the measurement points and the data time series, ensure that all data have time stamps and measurement point numbers, perform continuity detection on the extracted data, calculate the fluctuation rate of the data of a single measurement point using the ratio of the time interval to the data change range. If the fluctuation rate exceeds the set threshold, it is determined that there is an abnormality in this data point and it is stored in the abnormal data point sequence. The set fluctuation rate threshold is calculated based on the change range of the monitoring data. If the normal data fluctuation range of the measurement points in a certain area is ±2.5°C and the data acquisition time interval is 5s, the normal change rate range can be calculated. If the temperature change rate of a certain measurement point exceeds 0.5, it is marked as an abnormal data point to obtain the abnormal data point sequence.

[0163] S412: Repair the abnormal points through the abnormal data point sequence, identify the error range of the repaired data, correct the data points whose deviation exceeds the error range, using the formula:

[0164] ;

[0165] Calculate the data correction deviation degree to obtain the data repair record;

[0166] Among them, represents the data correction deviation degree, represents the th repaired data value of the point, represents the th repaired data value of the point, represents the data mean value, represents the total number of data points;

[0167] Explanation of formula parameters:

[0168] represents the data correction deviation degree, measuring the fluctuation of the repaired data, and is used to judge whether the repaired data is stable;

[0169] represents the th repaired data value of the point, which is the data after interpolation or fitting of the abnormal point;

[0170] represents the th repaired data value of the point, used to calculate the change rate of adjacent points;

[0171] represents the data mean value, calculating the average value of all points in the repaired data set, indicating the central tendency of this data sequence;

[0172] Parameter assignment and calculation:

[0173] Set the repaired data point sequence of a certain measuring point as: 32.5°C, 33.0°C, 32.8°C, 33.2°C, 32.9°C, and calculate the data mean value:

[0174] ;

[0175] Calculate the cumulative deviation of the data change rate:

[0176] ;

[0177] Calculate the data variance:

[0178] ;

[0179] ;

[0180] ;

[0181] Calculate the data correction deviation degree:

[0182] ;

[0183] The result shows that the data correction deviation degree is 0.8. If the set deviation degree reference value is 0.75 °C, then the current repaired data still has slight fluctuations and needs further adjustment. Otherwise, it can be considered that the data repair is completed, and the next step of data compensation and cooling state marking can be entered.

[0184] S413: Use the data repair record to compensate for continuous abnormal points, calculate the data recovery rate, and mark the measurement and acquisition points in the cooling state to generate a data calibration set;

[0185] Compensate for continuous abnormal points, perform continuity analysis on the repaired data points, extract time series data, calculate the data recovery rate, and define the recovery rate as the absolute value of the data change per unit time. If the recovery rate remains stable or tends to decrease within multiple consecutive time steps, it indicates that the data has recovered to the normal state. If the recovery rate fluctuates violently or there are too many continuous abnormal points, the compensation method needs to be used to adjust the repair value. Set the recovery rate threshold based on the monitoring data and equipment operating conditions. For temperature data, the change rate range of all normal measurement points in the past hour can be statistically analyzed, and the mean plus twice the standard deviation is taken as the recovery threshold. If the mean change rate of the normal measurement point data is 0.3 °C / s and the standard deviation is 0.15 °C / s, then the recovery threshold can be set to 0.3 + 2×0.15 = 0.6. If the recovery rate of the repaired data exceeds 0.6 °C / s, it is still in an abnormal state, and the corrected data needs to be further adjusted, and the compensation value is calculated. Use the time series analysis method to perform trend fitting on the data and perform weighted calculation based on the real-time data. If the recovery rate is lower than the set threshold, it is determined that the data recovery is stable, mark the measurement and acquisition points in the cooling state, and generate a data calibration set.

[0186] Please refer to Figure 6 for the specific steps to obtain the intelligent processing results of the measurement data:

[0187] S511: Through the data calibration set, combine the ID, timestamp, and keyword fields of the acquisition points to extract the data fingerprint, calculate the hash value, compare with the cached fingerprint, eliminate duplicate data, and obtain the filtered data fingerprint set;

[0188] Based on the ID, timestamp, and key fields of the collection points, extract the original data fingerprints. Each data fingerprint is generated by calculating the hash of the unique identifier, timestamp, and key field values of the data entry. In an actual scenario, such as fingerprint extraction for industrial sensor data, each data collection point will include a timestamp, device ID, and measurement value. Suppose a sensor data includes ID001, time 2025-02-20 10:00:00, and measurement value 45.6. Then this data fingerprint can be expressed as hash(ID001 + 2025-02-20 10:00:00 + 45.6). After calculating the hash value, it is necessary to compare it with the fingerprint data in the cache to determine whether this data already exists in the cache fingerprint library. In the specific comparison process, the query can be accelerated through an index structure. For example, use a B+ tree index to store the hash value to reduce the computational amount of traversal search. There are already fingerprint hashes A1, A2, A3 in the cache fingerprint library. If the currently calculated hash value A2 matches the fingerprint in the cache library, then it is determined that this data is duplicate data and is excluded. In batch data processing, excluding duplicate data can effectively reduce the data storage burden and optimize the efficiency of data analysis and calculation, obtaining a set of filtered data fingerprints.

[0189] S512: Using the set of filtered data fingerprints, calculate the ratio of the total data volume before and after filtering to obtain the duplicate data ratio. Compare it with the set threshold to determine whether to adjust the sampling frequency. If it exceeds the set threshold, adjust the frequency and trigger device self-check, update the sampling parameters, and obtain the adjusted sampling frequency parameters;

[0190] Count the number of data entries before and after deletion and calculate the data duplication ratio. The duplication ratio can be obtained by calculating the ratio of the number of excluded data to the total number of data. Suppose the original number of data entries is 10,000 and 8,000 remain after exclusion. Then the duplication ratio is (10,000 - 8,000) / 10,000 = 20%. The data duplication ratio needs to be compared with the preset threshold to determine whether the current sampling frequency needs to be adjusted. If the threshold is set to 15%, then the current duplication ratio exceeds the threshold and the sampling frequency needs to be reduced. In the industrial sensing data scenario, if a device collects data 100 times per minute and the duplication ratio exceeds 15%, then adjust the sampling frequency to 80 times per minute and synchronously trigger device self-check. The self-check content includes sensor status inspection, data transmission stability detection, etc. After adjustment, it is necessary to re-store the updated sampling parameters, such as the new sampling interval, data storage period, etc., and obtain the adjusted sampling frequency parameters.

[0191] S513: Based on the adjusted sampling frequency parameters, combined with the device self-check record, identify the error range, analyze the error change, compare with the error threshold, screen the affected measurement data, and obtain the intelligent processing result of the measurement data;

[0192] Combined with the device self-check record, compare the measurement data at different sampling frequencies, calculate the measurement error range. The error calculation can analyze the data stability through statistical indicators such as the mean and variance of the data. Set the mean of the data measured at the original sampling frequency to 50.2, and the mean after adjusting the sampling frequency to 49.8. Then the error is |50.2 - 49.8| = 0.4. Further analyze the error trend. If the error is stable within 0.5 in a short period of time, it is determined that the error fluctuation is small. If it exceeds 1.0, there is a problem of sensor drift. After the error calculation, compare the error threshold. If the set error threshold is 0.8, then when the error exceeds 0.8, mark this data as abnormal data, and screen the affected data range. For example, data with an error greater than 0.8 within 10 minutes are regarded as abnormal data, and obtain the intelligent processing result of the measurement data.

[0193] The measurement abnormal data detection system based on the rule engine is used to execute the above-mentioned measurement abnormal data detection method based on the rule engine. The system includes:

[0194] The data flow analysis module obtains the source point number, transmission link number and status code of the measurement data packet, calculates the flow deviation degree, screens the data packets exceeding the flow deviation threshold, identifies the transmission time interval of adjacent nodes, calculates the time interval deviation value, and obtains the abnormal flow feature set;

[0195] The abnormal point monitoring module uses the abnormal flow feature set, counts the submission times and activity change rates of the measurement collection points, compares the data submission times of adjacent points, extracts the associated point information, and obtains the abnormal point fluctuation graph;

[0196] The data comparison module uses the abnormal point fluctuation graph, sets the dynamic abnormal threshold, calculates the data offset and compares it with the threshold, identifies the fluctuation amplitude and change situation, and obtains the abnormal trend identification table;

[0197] The offset degree identification module uses the abnormal trend identification table, extracts the real-time data of the neighborhood points, identifies the offset degree of the data distribution, analyzes the trend change, compensates for the continuous abnormal points, and obtains the data calibration set;

[0198] The data evaluation module passes through the data calibration set, compares the data fingerprint and the cache fingerprint, eliminates duplicate data, adjusts the sampling frequency, triggers the device self-check, and obtains the intelligent processing result of the measurement data.

[0199] The above is only the preferred embodiment of the present invention, and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A measurement abnormal data detection method based on a rule engine, characterized in that: The following steps are involved: S1: Based on the source point number, transmission link number and real-time status code of the measurement data packet, the transmission path and timestamp sequence are extracted, the node information of the measurement data collection, transmission and processing links is recorded, the data flow path is constructed, the data flow path is compared with the preset abnormal flow direction, the flow deviation is calculated, the deviation between the transmission time interval of adjacent nodes and the standard time interval is identified, the abnormal mode in which the path fitting degree exceeds the preset threshold is identified, and the abnormal flow feature set is obtained; S2: Based on the abnormal flow feature set, monitor the activity status of the measurement and collection points, identify the number of data submissions and changes in activity, analyze the synchronization of adjacent point data, extract related point information, identify abnormal points, and form an abnormal point fluctuation map; S3: using the abnormal point fluctuation graph, monitoring the data flow, setting a dynamic abnormal threshold, comparing the real-time data deviation, analyzing the fluctuation amplitude, and obtaining an abnormal trend identification table; S4: using the abnormal trend identification table, extracting the real-time data of the neighboring points, compensating the continuous abnormal points, marking the measurement and collection points of the cooling state, and generating a data calibration set; S5: extracting data fingerprints through the data calibration set, combining the ID, timestamp and key fields of the collection points, comparing the cache fingerprints, eliminating duplicate data, adjusting the sampling frequency, and obtaining the intelligent processing results of the measurement data; The steps for obtaining the abnormal flow feature set are specifically as follows: S111: extracting the transmission path and timestamp sequence based on the source point number, transmission link number and real-time status code of the measurement data packet, and establishing a data flow path; S112: using the data flow path, comparing with the preset abnormal flow direction, calculating the flow deviation, identifying the transmission time interval of adjacent nodes, and calculating the deviation with the standard time interval, extracting the data packet number and path exceeding the time threshold range, establishing the corresponding relationship between the path deviation degree and the time deviation, and obtaining the path deviation and time deviation set; S113: calling the path deviation and time deviation set, evaluating the correlation between the path deviation degree and the time deviation, using the formula: ; Calculate the path fitting degree, compare the threshold, filter out abnormal paths, extract the features of abnormal paths, and obtain the abnormal flow feature set; in, is the path fit, is the path deviation, and is the weight parameter, is the time deviation.

2. The measurement abnormal data detection method based on rule engine according to claim 1 is characterized in that: The abnormal flow feature set includes the flow deviation exceeding the limit path, the abnormal transmission time interval point, and the path matching abnormal mode; the abnormal point fluctuation diagram includes the active state fluctuation trend, the data submission frequency change trend, and the abnormal distribution of adjacent point data synchronization; the abnormal trend identification table includes the data fluctuation amplitude exceeding the limit point, the change rate abnormal point, and the data distribution offset exceeding the threshold point; the data calibration set includes the repair data record, the compensation abnormal point data, and the cooling state marking point; the measurement data intelligent processing results include repeated data ratio statistics, sampling frequency adjustment strategy, and equipment self-test trigger record.

3. The measurement abnormal data detection method based on rule engine according to claim 1 is characterized in that: The steps for obtaining the abnormal point fluctuation map are specifically as follows: S211: Based on the abnormal flow feature set, monitor the activity status of the measurement and collection points, count the number of data submissions of each measurement and collection point within a fixed time window, identify the submission frequency, calculate the change rate of the point submission frequency, call the set submission frequency change threshold, determine whether the point activity has a sudden change, and obtain the activity change record; S212: Analyze the synchronization of adjacent point data through the activity change record, compare the submission times of adjacent points in the same time window, calculate the time deviation value of adjacent points, and obtain the adjacent point time deviation identification result; S213: Using the adjacent point time deviation identification result, extract the associated point information, filter the point set whose time deviation value exceeds the set threshold, identify the distribution range of the measurement and collection points, and use the formula: ; Calculate the aggregation degree of abnormal points, filter the abnormal point set, and obtain the abnormal point fluctuation map; in, Represents the degree of abnormal point aggregation. Representative The spatial coordinates of the abnormal points, Represents the mean coordinates of the abnormal points, Represents the number of abnormal points.

4. The measurement abnormal data detection method based on rule engine according to claim 3 is characterized in that: The steps for obtaining the abnormal trend identification table are specifically as follows: S311: Based on the abnormal point fluctuation diagram, multiple time series intervals of the monitoring data stream are analyzed, the fluctuation amplitude, change rate and data distribution characteristics in each interval are extracted, and the mean, standard deviation and change rate of the fluctuation amplitude are calculated, and the fluctuation characteristic coefficient is obtained by comparing the mean difference and standard deviation change trend of adjacent intervals; S312: Using the fluctuation characteristic coefficient and the set dynamic abnormal threshold, analyze the deviation of real-time data, using the formula: ; Calculate the degree of data deviation. When the set threshold is exceeded, the data point is marked as an abnormal fluctuation point, and the proportion of multiple abnormal fluctuation points in the differentiated time interval is counted to obtain a quantitative value of the degree of deviation; in, Represents real time The degree of data deviation, is the real-time data value, is the mean of the previous interval, is the standard deviation of the previous interval, is a constant, is the volatility weight factor; S313: Calculate the rate of change of abnormal fluctuation trends in multiple time periods in combination with the time distribution of abnormal fluctuation points, identify the degree of deviation of the abnormal distribution, and use the quantified value of the degree of deviation and the amplitude of change to evaluate and obtain an abnormal trend identification table.

5. The measurement abnormal data detection method based on rule engine according to claim 4 is characterized in that: The steps for obtaining the data calibration set are specifically as follows: S411: using the abnormal trend identification table, extracting the real-time data of the neighborhood points, performing continuity detection on the real-time data, screening abnormal data points, and identifying the data fluctuation trend according to the data change amplitude and time series to obtain an abnormal data point sequence; S412: Repair the abnormal points through the abnormal data point sequence, identify the error range of the repaired data, and correct the data points whose deviations exceed the error range, using the formula: ; Calculate the data correction deviation and obtain the data repair record; in, Represents the data correction deviation, Representative The repair data value of points, Representative The repair data value of points, represents the data mean, represents the total number of data points; S413: Using the data repair record, compensating for the continuous abnormal points, calculating the data recovery rate, marking the measurement and collection points in the cooling state, and generating a data calibration set.

6. The measurement abnormal data detection method based on rule engine according to claim 5 is characterized in that: The steps for obtaining the measurement data intelligent processing result are specifically as follows: S511: extracting data fingerprints through the data calibration set, combining the ID, timestamp and key fields of the collection points, calculating hash values, comparing cache fingerprints, eliminating duplicate data, and obtaining a filtered data fingerprint set; S512: using the filtered data fingerprint set, calculating the ratio of the total amount of data before and after filtering, obtaining the repeated data ratio, comparing it with the set threshold, determining whether to adjust the sampling frequency, and if it exceeds the set threshold, adjusting the frequency, triggering the device self-check, updating the sampling parameters, and obtaining the adjusted sampling frequency parameters; S513: Based on the adjusted sampling frequency parameters and in combination with the equipment self-test record, the error range is identified, the error change is analyzed, the error threshold is compared, the affected measurement data is screened, and the measurement data intelligent processing result is obtained.

7. A measurement abnormal data detection system based on a rule engine, characterized in that: According to the rule engine-based measurement abnormal data detection method according to any one of claims 1 to 6, the system comprises: The data flow analysis module obtains the source point number, transmission link number and status code of the measured data packet, calculates the flow deviation, filters the data packets that exceed the flow deviation threshold, identifies the transmission time interval of adjacent nodes, calculates the time interval deviation value, and obtains the abnormal flow feature set; The abnormal point monitoring module uses the abnormal flow feature set to statistically measure the submission times and activity change rate of the collection points, compare the submission time of adjacent point data, extract the related point information, and obtain the abnormal point fluctuation map; The data comparison module uses the abnormal point fluctuation graph to set the dynamic abnormal threshold, calculate the data offset and compare the threshold, identify the fluctuation amplitude and change, and obtain the abnormal trend identification table; The deviation degree identification module uses the abnormal trend identification table to extract the real-time data of the neighborhood points, identify the deviation degree of data distribution, analyze the trend change, compensate for the continuous abnormal points, and obtain the data calibration set; The data evaluation module uses the data calibration set to compare data fingerprints and cache fingerprints, remove duplicate data, adjust the sampling frequency, trigger device self-check, and obtain intelligent processing results of measurement data.

Citation Information

Patent Citations

  • Defense system and dynamic defense method for WEB application automation attack behavior

    CN115065537A

  • Intelligent monitoring system of transformer substation channel cover plate based on Internet of Things

    CN119401658A