Data analysis method and system for power grid data exception
By monitoring the number of outliers and retransmissions in power grid data, and combining the analysis of repair delay and aggregation frequency, abnormal alarms are identified and output. This solves the problem of difficulty in identifying abnormal repair behavior in power grid data in existing technologies, and realizes observable and quantifiable analysis of the data repair process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG POWER GRID CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies lack systematic analysis methods for abnormal repair behavior of power grid data during cleaning, interpolation, or smoothing processes, which weakens or masks the original abnormal characteristics, making it difficult to effectively identify and quantify them.
By triggering a judgment strategy based on the repair mechanism based on the number of outliers and the number of data retransmissions, and combining a time-series analysis model of repair latency and repair aggregation frequency, interval aggregation features of repair types are extracted, and correlation detection is performed to identify abnormal repair behaviors.
It enables observable and quantifiable analysis of the power grid data repair process, identifies and outputs abnormal alarm signals, and ensures data credibility assessment and anomaly tracing.
Smart Images

Figure CN122019980A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and more specifically, to a data analysis method and system for power grid data anomalies. Background Technology
[0002] With the rapid development of new power systems and digital power grids, dispatch automation systems, distribution automation systems, electricity consumption information collection systems, and various online monitoring devices continuously generate massive amounts of power grid operation data. To meet the application needs of situational awareness, operation analysis, status assessment, and decision support, it is usually necessary to connect to a unified data processing platform to perform data cleaning, missing data completion, and other processing operations to improve data integrity and availability.
[0003] The existing technology has the following shortcomings: Currently, existing technologies focus on quality assessment and anomaly determination of raw or repaired power grid data. They lack systematic analysis methods for data processing processes such as data retransmission behavior, repair triggering mechanisms, repair type distribution, and time-series repair characteristics. Once the original anomalies are passively corrected during power grid data cleaning, interpolation, or smoothing, the original anomaly characteristics are weakened or even masked, and the anomaly repair behavior itself is difficult to effectively identify and quantify. Therefore, a data analysis method and system for power grid data anomalies is proposed.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a data analysis method and system for power grid data anomalies. This method utilizes a repair mechanism based on the number of anomalies and the number of data retransmissions to trigger a judgment strategy, and combines a time-series analysis model of repair delay and repair aggregation frequency, a repair type interval aggregation and repetitive feature extraction mechanism, and a correlation detection method between repair magnitude and power grid operation data to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a data analysis method for power grid data anomalies, comprising the following steps: Step S1: Monitor the power grid data, count the number of outliers in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of outliers, and collect the number of data retransmissions on the data processing platform before performing repair processing; Step S2: Select the triggering of the regular processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated. Step S3: Set up the feature analysis window, access the repair type information of each time series record in the feature analysis window, perform interval aggregation analysis on the repair type information and generate repair repetition features, and evaluate the data repair status of the time series record in combination with the repair aggregation frequency. Step S4: When the data repair status is abnormal, retrieve the repair amplitude data and power grid operation data, perform correlation detection on the repair amplitude data and power grid operation data to obtain the detection results, determine whether there is abnormal repair based on the detection results and output an abnormal alarm signal.
[0007] In a preferred embodiment, in step S1, the power grid data is continuously monitored, and the power grid data is the raw data collected in real time by the power grid side sensing device during operation. The power grid data forms a time-series data sequence arranged in chronological order according to a preset sampling period. Based on preset anomaly detection conditions, each data point in the time-series data sequence is compared one by one. When any data point meets the anomaly detection conditions, the data point is determined to be an anomaly. The number of abnormal values in the power grid data is obtained by cumulatively counting the data points that meet the anomaly detection criteria; When the number of outliers reaches or exceeds the preset outlier intervention threshold, it is determined that the degree of anomaly in the power grid data has exceeded the local processing capacity, and it is necessary to connect to the data processing platform and perform repair processing. When the number of outliers is less than the preset outlier intervention threshold, the outlier level of the current power grid data is determined to be within an acceptable range, and the data processing platform's repair process is not triggered.
[0008] In a preferred embodiment, in step S1, if it is determined that the data processing platform needs to be accessed and repair processing is to be performed, the number of data retransmissions in the current data transmission process corresponding to the data processing platform is collected before the power grid data is officially submitted to the data processing platform. The number of data retransmissions is the cumulative number of times data is repeatedly sent during the historical process of power grid data transmission to the data processing platform due to communication anomalies, transmission verification failures, or platform-side requests for retransmission.
[0009] In a preferred embodiment, in step S2, when the number of data retransmissions is less than a preset retransmission determination threshold, a conventional processing mechanism is triggered. When the number of data retransmissions reaches or exceeds the retransmission judgment threshold, the repair analysis mechanism is triggered. After the repair analysis mechanism is triggered, the third-party data processing platform completes the repair processing of the power grid data and sends the repaired power grid data back, and then obtains the time sequence record of the corresponding power grid data. The time-series record is a continuous data sequence indexed by timestamps, where each data point is associated with the original acquisition time and the repair completion and transmission time.
[0010] In a preferred embodiment, in step S2, based on the time sequence record, the repair delay is calculated for each data point. The repair delay is defined as the time difference between the repair completion and transmission time and the corresponding original acquisition time. Based on the repair delay, the repair completion event is time-located, and the repair completion event is aggregated and statistically analyzed within a preset time statistics window to obtain the number of repairs completed corresponding to the time statistics window. The number of repairs completed within a unit of time is used as the repair aggregation frequency.
[0011] In a preferred embodiment, in step S3, a feature analysis window is preset, and the repair type information corresponding to each time series record is obtained through the repair log of a third-party data processing platform within the feature analysis window; Repair type information refers to the repair type identifier used by a third-party data processing platform when performing repair processing on power grid data; Interval aggregation analysis is performed on the repair type information. Adjacent time-series records with the same repair type identifier within the feature analysis window are divided into the same repair type interval, and the continuous length and occurrence frequency of each repair type interval are counted. For the same repair type interval, the ratio of its occurrence frequency to its continuous length is taken as the repair type coverage coefficient, and the maximum value of each repair type coverage coefficient is taken as the repair repetition feature.
[0012] In a preferred embodiment, in step S3, the repair repetition feature and the repair aggregation frequency are standardized to obtain the repair repetition coefficient and the repair aggregation coefficient, respectively. The product of the repair repetition coefficient and the repair aggregation coefficient is used as the comprehensive repair index; If the comprehensive repair index is greater than the preset repair threshold, the data repair status of the time-series records is determined to be abnormal. Conversely, if the data repair status of the time-series record is not found to be normal, then the data repair status is determined to be normal.
[0013] In a preferred embodiment, in step S4, when the data repair status is abnormal, the repair magnitude data, including the repair magnitude value, is obtained through the repair log of the data processing platform. Power grid operation data refers to the operational status data collected by the power grid side monitoring unit within the feature analysis window, including the voltage change of the repaired power grid data. Among them, the repair magnitude value refers to the numerical difference between the data value after repair and the data value before repair when the third-party data processing platform performs repair processing on the power grid data; the voltage change refers to the numerical difference between the voltage data corresponding to adjacent time series records within the feature analysis window.
[0014] In a preferred embodiment, in step S4, if the repair amplitude value and voltage change amount have the same sign direction at the same time, it is determined that the repair behavior of the time sequence record is consistent with the change direction of the power grid operating state, and is recorded as consistent repair. Conversely, if the timing record's repair behavior is inconsistent with the direction of change in the power grid's operating status, it is considered an inconsistent repair. The number of time-series records that were corrected for inconsistencies and the total number of time-series records were counted, and the ratio of the number of time-series records that were corrected for inconsistencies to the total number of time-series records was used as the correlation detection result. If the associated detection result is greater than the preset detection result threshold, an abnormal alarm signal will be output. Conversely, if the condition is not met, then no abnormal alarm signal will be output.
[0015] A data analysis system for power grid data anomalies includes an anomaly monitoring module, a mechanism determination module, a status assessment module, and an anomaly alarm module. The functions of each module are as follows: The anomaly monitoring module is used to monitor power grid data, count the number of abnormal values in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of abnormal values, collect the number of data retransmissions on the data processing platform before performing repair processing, and pass the number of data retransmissions to the mechanism judgment module. The mechanism determination module selects to trigger the conventional processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated, and the repair aggregation frequency is transmitted to the status assessment module. The status assessment module sets up a feature analysis window, accesses the repair type information of each time series record in the feature analysis window, performs interval aggregation analysis on the repair type information and generates repair repetition features, evaluates the data repair status of the time series record in combination with the repair aggregation frequency, and passes the data repair status to the anomaly alarm module. When the data repair status is abnormal, the abnormal alarm module retrieves the repair magnitude data and power grid operation data, performs correlation detection on the repair magnitude data and power grid operation data, obtains the detection results, determines whether there is abnormal repair based on the detection results, and outputs an abnormal alarm signal.
[0016] The technical effects and advantages of this invention are as follows: This invention monitors the anomaly level and retransmission characteristics of power grid data, and determines whether to introduce a repair analysis mechanism by combining the number of anomalies and retransmission behavior. Under the repair analysis mechanism, the time-series records of repaired and returned data are analyzed to calculate the repair delay and repair aggregation frequency, thus characterizing the concentration of data repair behavior from a time dimension. Furthermore, within the feature analysis window, repair type information is aggregated in intervals to extract repair repetition features and comprehensively evaluate the data repair status. When the repair status is abnormal, correlation detection is performed by combining repair amplitude data and power grid operation data to identify abnormal repair behavior and output alarm results. Even when the original anomaly is masked by smoothing or interpolation, the observable and quantifiable analysis of the data repair process itself is achieved, providing a new technical means for power grid data reliability assessment and anomaly tracing. Attached Figure Description
[0017] Figure 1 This is a flowchart of a data analysis method for power grid data anomalies according to the present invention.
[0018] Figure 2 This is a schematic diagram of a data analysis system for power grid data anomalies according to the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] This invention monitors the anomaly level and retransmission characteristics of power grid data, and determines whether to introduce a repair analysis mechanism by combining the number of anomalies and retransmission behavior. Under the repair analysis mechanism, the time-series records of repaired and returned data are analyzed to calculate the repair delay and repair aggregation frequency, thus characterizing the concentration of data repair behavior from a time dimension. Furthermore, within the feature analysis window, repair type information is aggregated in intervals to extract repair repetition features and comprehensively evaluate the data repair status. When the repair status is abnormal, correlation detection is performed by combining repair amplitude data and power grid operation data to identify abnormal repair behavior and output alarm results. In cases where the original anomaly is masked by smoothing or interpolation, the observable and quantifiable analysis of the data repair process itself is achieved.
[0021] Example 1, such as Figure 1 As shown, a data analysis method for power grid data anomalies includes the following steps: Step S1: Monitor the power grid data, count the number of outliers in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of outliers, and collect the number of data retransmissions on the data processing platform before performing repair processing; Step S2: Select the triggering of the regular processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated. Step S3: Set up the feature analysis window, access the repair type information of each time series record in the feature analysis window, perform interval aggregation analysis on the repair type information and generate repair repetition features, and evaluate the data repair status of the time series record in combination with the repair aggregation frequency. Step S4: When the data repair status is abnormal, retrieve the repair amplitude data and power grid operation data, perform correlation detection on the repair amplitude data and power grid operation data to obtain the detection results, determine whether there is abnormal repair based on the detection results and output an abnormal alarm signal.
[0022] The specific implementation is as follows: In step S1, the power grid data is continuously monitored. The power grid data is the raw data collected in real time by the power grid side sensing devices during operation. The sensing devices include, but are not limited to, voltage acquisition devices, current acquisition devices, or power acquisition devices. The power grid data forms a time-series data sequence arranged in chronological order according to a preset sampling period.
[0023] When monitoring power grid data, each data point in the time series data sequence is compared one by one based on preset anomaly judgment conditions. When any data point meets the anomaly judgment conditions, the data point is judged as an anomaly value.
[0024] Anomaly detection criteria are a set of rules used to objectively characterize the deviation of power grid data from normal operating conditions, including criteria for missing data, criteria for sudden changes, or criteria for values exceeding reasonable ranges.
[0025] By cumulatively counting the data points that meet the anomaly detection criteria, the number of anomalies in the power grid data is obtained, which reflects the concentration and intensity of anomalies in the power grid data within the current statistical period.
[0026] The number of outliers is compared and analyzed with a preset anomaly intervention threshold. The anomaly intervention threshold is a threshold parameter set based on historical statistical results. It is used to indicate at what level of anomaly it is necessary to introduce a data processing platform to repair the power grid data. Specifically, based on the historical operation data of the power grid, a long-term statistical analysis is performed on the sequence of outliers of the target monitoring quantity. The mean and standard deviation of the number of outliers are calculated within a preset statistical period, and the sum of the mean and standard deviation is used as the anomaly intervention threshold.
[0027] When the number of outliers reaches or exceeds the outlier intervention threshold, it is determined that the degree of anomaly in the power grid data has exceeded the local processing capacity, and it is necessary to connect to the data processing platform and perform repair processing. When the number of outliers is less than the outlier intervention threshold, the anomaly level of the current power grid data is determined to be within an acceptable range, and the data processing platform's repair process is not triggered.
[0028] If it is determined that access to the data processing platform and repair processing are required, before the power grid data is officially submitted to the data processing platform, the number of data retransmissions during the current data transmission process of the data processing platform is collected. The number of data retransmissions is the cumulative value of the number of times data is repeatedly sent due to communication abnormalities, transmission verification failures, or platform-side requests for retransmission during the historical process of power grid data transmission to the data processing platform. It comes from the transmission records of the data processing platform and is used to reflect the stability of the data transmission link before the repair processing begins and the strength of the platform's access and confirmation behavior for power grid data.
[0029] In step S2, a mechanism is selected for the subsequent processing of power grid data based on the number of data retransmissions. Specifically, the number of data retransmissions is compared with a preset retransmission judgment threshold. The retransmission judgment threshold is a threshold parameter used to characterize whether there is abnormal retransmission behavior during data transmission and platform access. It is set based on historical data transmission statistics. Specifically, within a preset statistical period, the transmission process of the same type of power grid data under normal communication conditions is monitored, and the number of data retransmissions occurring in each data transmission period is counted to form a historical sample sequence of data retransmissions. Based on the historical sample sequence, statistical characteristic parameters of the number of data retransmissions are calculated, including the mean and standard deviation. The statistical upper limit reflecting the data retransmission level under normal communication conditions is used as the retransmission judgment threshold.
[0030] When the number of data retransmissions is less than the retransmission judgment threshold, it is determined that the third-party data processing platform has not experienced any significant abnormal retransmission behavior during the current data access process, and the normal processing mechanism is triggered. When the number of data retransmissions reaches or exceeds the retransmission judgment threshold, it is determined that the third-party data processing platform has a high frequency of data retransmission behavior during the current data access process, triggering the repair analysis mechanism.
[0031] It should be noted that the conventional processing mechanism refers to a method of processing power grid data based on conventional anomaly analysis procedures using data numerical characteristics, assuming that no significant abnormal retransmission behavior has occurred during the current data access process by the third-party data processing platform. In this mechanism, power grid data does not undergo a dedicated analysis process targeting data repair behavior. Instead, predetermined data verification, anomaly detection, or status assessment operations are performed directly on the returned power grid data, without separately modeling or analyzing the repair process, repair behavior, or repair traces of the third-party data processing platform. By adopting this conventional processing mechanism, additional analytical complexity can be avoided when platform access is stable and data repair behavior is not significant, thereby improving overall data processing efficiency.
[0032] The above method enables traffic control of processing paths under different data access states.
[0033] After the repair analysis mechanism is triggered, the third-party data processing platform completes the repair processing of the power grid data and sends the repaired power grid data back, and then obtains the time sequence record of the corresponding power grid data.
[0034] The time-series record is a continuous data sequence indexed by timestamps. Each data point is associated with the original acquisition time and the repair completion and feedback time. The repair completion and feedback time serves as a repair completion time marker, indicating the specific time when the data point completes the repair processing within the third-party data processing platform and is returned to the power grid side.
[0035] Based on time-series records, the repair delay is calculated for each data point. The repair delay is defined as the time difference between the time when the repair is completed and the corresponding original acquisition time. It is used to quantify the length of time a single data point takes to undergo repair processing and transmission within a third-party data processing platform.
[0036] The above calculations form a set of repair delays based on data points, reflecting the distribution of repair processing time for different power grid data points during the current data access phase.
[0037] After obtaining the repair latency of each data point, the repair completion event is time-located based on the repair latency, and the repair completion events are aggregated and statistically analyzed within a preset time statistics window. Specifically, the repair completion feedback time is used as the statistical base time. Data points whose repair completion feedback times fall within the same time statistics window are determined as data points that have completed repair processing within that time statistics window. The number of corresponding data points within that time statistics window is accumulated to obtain the number of repair completions corresponding to that time statistics window.
[0038] Furthermore, the number of repairs completed within a unit time period is used as the repair aggregation frequency. The repair aggregation frequency is used to characterize the density of data points that the third-party data processing platform completes repair processing within the corresponding time statistical window.
[0039] By determining the repair completion and feedback time through repair delay and performing aggregation counts within the time window accordingly, the repair aggregation frequency objectively quantifies the concentration level of repair processing behavior in the time dimension. This is used to characterize the overall intensity of the third-party data processing platform performing repair processing on power grid data at the current stage, and serves as a basic parameter for subsequent repair behavior feature analysis and anomaly judgment.
[0040] In step S3, a feature analysis window is preset, and the repair type information corresponding to each time series record is accessed in the feature analysis window. The repair type information refers to the specific repair type identifier used by the third-party data processing platform when performing repair processing on the power grid data. It comes from the repair log of the third-party data processing platform. After obtaining all repair type information within the feature analysis window, interval aggregation analysis is performed on the repair type information. Interval aggregation analysis refers to dividing adjacent time-series records with the same repair type identifier within the feature analysis window into the same repair type interval, and counting the continuous length and occurrence frequency of each repair type interval. Among them, the continuous length is the number of data points of the same repair type appearing consecutively on the time axis, which is used to reflect the persistence of a single repair method in a local time period; the number of occurrences is the number of intervals of the same repair type within the feature analysis window, which is used to reflect the repeated triggering of the repair method within the feature analysis window; For the same repair type interval, the ratio of its occurrence frequency to its continuous length is taken as the repair type coverage coefficient, and the maximum value of each repair type coverage coefficient is taken as the repair repetition feature. The higher the repetitive feature, the more likely the same repair type is to be triggered multiple times or exhibit a clear concentrated distribution characteristic within the feature analysis window. After standardizing the repair duplication feature and the repair aggregation frequency respectively, the repair duplication coefficient and the repair aggregation coefficient are obtained. The product of the repair repetition coefficient and the repair aggregation coefficient is used as the comprehensive repair index; The comprehensive repair index is compared with the preset repair threshold to assess the data repair status of the time-series records. If the comprehensive repair index is greater than the preset repair threshold, the data repair status of the time-series records is determined to be abnormal. Conversely, the data repair status of the time-series records is determined to be normal. When the frequency of data aggregation and the frequency of data duplication increase simultaneously, it indicates that a large amount of data using the same or similar repair methods are being processed in a concentrated manner within the feature analysis window, and the data repair behavior shows obvious characteristics of concentration and duplication.
[0041] It should be noted that the preset feature analysis window can be set according to the sampling period and repair behavior density of the power grid data; the standardization processing method includes, but is not limited to, standard linear transformation based on interval scaling, Z-Score standardization method based on statistics, or normalization method based on nonlinear mapping function. The application method of standardization processing will not be elaborated here; the preset repair threshold can be set according to the comprehensive repair index statistical results under the historical normal data repair state.
[0042] This step enables an observable and quantifiable state assessment of the data repair process from the perspective of the type distribution and temporal repeatability of the repair behavior itself, when the abnormal characteristics of power grid data may be weakened or masked by repair processing. This provides a reliable intermediate criterion for the identification of abnormal repair behavior.
[0043] In step S4, when the data repair status is abnormal, the repair magnitude data and power grid operation data are retrieved. The repair magnitude data refers to the correction amount information applied to the power grid data value by the third-party data processing platform during the repair processing of the power grid data. It comes from the repair log of the data processing platform. The repair magnitude data includes the repair magnitude value between the data value before repair and the data value after repair. Power grid operation data refers to the operational status data collected by the power grid side monitoring unit within the feature analysis window, including the voltage change of the repaired power grid data. Among them, the repair magnitude value refers to the numerical difference between the data value after repair and the data value before repair when the third-party data processing platform performs repair processing on the power grid data record; the voltage change refers to the numerical difference between the voltage data corresponding to adjacent time series records within the feature analysis window. After acquiring the repair magnitude data and power grid operation data, correlation detection is performed on the repair magnitude data and power grid operation data; If the repair amplitude value and voltage change amount have the same sign direction at the same time, it is determined that the repair behavior of the time sequence record is consistent with the change direction of the power grid operating state, and is recorded as consistent repair. Conversely, if the timing record's repair behavior is inconsistent with the direction of change in the power grid's operating status, it is recorded as inconsistent repair.
[0044] Within the feature analysis window, the number of time-series records that have undergone inconsistency repair and the total number of time-series records are counted. The ratio of the number of time-series records that have undergone inconsistency repair to the total number of time-series records is used as the correlation detection result. The associated detection results are compared with a preset detection result threshold to determine whether to output an abnormal alarm signal. If the associated detection result is greater than the preset detection result threshold, an abnormal alarm signal will be output. Conversely, if the condition is not met, then no abnormal alarm signal will be output.
[0045] It should be noted that the power grid side monitoring unit is a sensing and acquisition device used to collect real-time data on the power grid's operating status; the preset detection result threshold can be set based on the associated detection results during the historical operating cycle when the power grid is in a normal repair state.
[0046] This step enables cross-validation between the repair magnitude and the actual operating status of the power grid, based on the abnormal data repair status. This avoids misjudgment caused by directly triggering alarms based solely on repair behavior characteristics, ensuring that abnormal alarms are only triggered when the repair process is significantly inconsistent with the physical operating status of the power grid, thereby improving the reliability of abnormal repair identification results.
[0047] Example 2, as Figure 2 As shown, a data analysis system for power grid data anomalies is provided to implement a data analysis method for power grid data anomalies. The system includes an anomaly monitoring module, a mechanism determination module, a status assessment module, and an anomaly alarm module. The functions of each module are as follows: The anomaly monitoring module is used to monitor power grid data, count the number of abnormal values in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of abnormal values, collect the number of data retransmissions on the data processing platform before performing repair processing, and pass the number of data retransmissions to the mechanism judgment module. The mechanism determination module selects to trigger the conventional processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated, and the repair aggregation frequency is transmitted to the status assessment module. The status assessment module sets up a feature analysis window, accesses the repair type information of each time series record in the feature analysis window, performs interval aggregation analysis on the repair type information and generates repair repetition features, evaluates the data repair status of the time series record in combination with the repair aggregation frequency, and passes the data repair status to the anomaly alarm module. When the data repair status is abnormal, the abnormal alarm module retrieves the repair magnitude data and power grid operation data, performs correlation detection on the repair magnitude data and power grid operation data, obtains the detection results, determines whether there is abnormal repair based on the detection results, and outputs an abnormal alarm signal.
[0048] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0049] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0050] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0051] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0052] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0053] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0054] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0055] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0056] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0057] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0058] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0059] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above specification.
Claims
1. A data analysis method for power grid data anomalies, characterized in that: Includes the following steps: Step S1: Monitor the power grid data, count the number of outliers in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of outliers, and collect the number of data retransmissions on the data processing platform before performing repair processing; Step S2: Select the triggering of the regular processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated. Step S3: Set up the feature analysis window, access the repair type information of each time series record in the feature analysis window, perform interval aggregation analysis on the repair type information and generate repair repetition features, and evaluate the data repair status of the time series record in combination with the repair aggregation frequency. Step S4: When the data repair status is abnormal, retrieve the repair amplitude data and power grid operation data, perform correlation detection on the repair amplitude data and power grid operation data to obtain the detection results, determine whether there is abnormal repair based on the detection results and output an abnormal alarm signal.
2. The data analysis method for power grid data anomalies according to claim 1, characterized in that: In step S1, the power grid data is continuously monitored. The power grid data is the raw data collected in real time by the power grid-side sensing equipment during operation. The power grid data forms a time-series data sequence arranged in chronological order according to a preset sampling period. Based on preset anomaly detection conditions, each data point in the time-series data sequence is compared one by one. When any data point meets the anomaly detection conditions, the data point is determined to be an anomaly. The number of abnormal values in the power grid data is obtained by cumulatively counting the data points that meet the anomaly detection criteria; When the number of outliers reaches or exceeds the preset outlier intervention threshold, it is determined that the degree of anomaly in the power grid data has exceeded the local processing capacity, and it is necessary to connect to the data processing platform and perform repair processing. When the number of outliers is less than the preset outlier intervention threshold, the outlier level of the current power grid data is determined to be within an acceptable range, and the data processing platform's repair process is not triggered.
3. The data analysis method for power grid data anomalies according to claim 2, characterized in that: In step S1, if it is determined that the data processing platform needs to be accessed and repair processing needs to be performed, before the power grid data is officially submitted to the data processing platform, the number of data retransmissions in the current data transmission process corresponding to the data processing platform is collected. The number of data retransmissions is the cumulative number of times data is repeatedly sent during the historical process of power grid data transmission to the data processing platform due to communication anomalies, transmission verification failures, or platform-side requests for retransmission.
4. The data analysis method for power grid data anomalies according to claim 1, characterized in that: In step S2, when the number of data retransmissions is less than the preset retransmission judgment threshold, the normal processing mechanism is triggered. When the number of data retransmissions reaches or exceeds the retransmission judgment threshold, the repair analysis mechanism is triggered. After the repair analysis mechanism is triggered, the third-party data processing platform completes the repair processing of the power grid data and sends the repaired power grid data back, and then obtains the time sequence record of the corresponding power grid data. The time-series record is a continuous data sequence indexed by timestamps, where each data point is associated with the original acquisition time and the repair completion and transmission time.
5. A data analysis method for power grid data anomalies according to claim 4, characterized in that: In step S2, based on the time sequence record, the repair delay is calculated for each data point. The repair delay is defined as the time difference between the repair completion and transmission time and the corresponding original acquisition time. Based on the repair delay, the repair completion event is time-located, and the repair completion event is aggregated and statistically analyzed within a preset time statistics window to obtain the number of repairs completed corresponding to the time statistics window. The number of repairs completed within a unit of time is used as the repair aggregation frequency.
6. The data analysis method for power grid data anomalies according to claim 1, characterized in that: In step S3, a feature analysis window is preset, and the repair type information corresponding to each time series record is obtained through the repair log of the third-party data processing platform within the feature analysis window; Repair type information refers to the repair type identifier used by a third-party data processing platform when performing repair processing on power grid data; Interval aggregation analysis is performed on the repair type information. Adjacent time-series records with the same repair type identifier within the feature analysis window are divided into the same repair type interval, and the continuous length and occurrence frequency of each repair type interval are counted. For the same repair type interval, the ratio of its occurrence frequency to its continuous length is taken as the repair type coverage coefficient, and the maximum value of each repair type coverage coefficient is taken as the repair repetition feature.
7. The data analysis method for power grid data anomalies according to claim 1, characterized in that: In step S3, the repair repetition feature and the repair aggregation frequency are standardized to obtain the repair repetition coefficient and the repair aggregation coefficient, respectively. The product of the repair repetition coefficient and the repair aggregation coefficient is used as the comprehensive repair index; If the comprehensive repair index is greater than the preset repair threshold, the data repair status of the time-series records is determined to be abnormal. Conversely, if the data repair status of the time-series record is not found to be normal, then the data repair status is determined to be normal.
8. The data analysis method for power grid data anomalies according to claim 1, characterized in that: In step S4, when the data repair status is abnormal, the repair magnitude data, including the repair magnitude value, is obtained through the repair log of the data processing platform. Power grid operation data refers to the operational status data collected by the power grid side monitoring unit within the feature analysis window, including the voltage change of the repaired power grid data. Among them, the repair magnitude value refers to the numerical difference between the data value after repair and the data value before repair when the third-party data processing platform performs repair processing on the power grid data; the voltage change refers to the numerical difference between the voltage data corresponding to adjacent time series records within the feature analysis window.
9. A data analysis method for power grid data anomalies according to claim 8, characterized in that: In step S4, if the repair amplitude value and voltage change amount have the same sign direction at the same time, it is determined that the repair behavior of the time sequence record is consistent with the change direction of the power grid operating state, and is recorded as consistent repair. Conversely, if the timing record's repair behavior is inconsistent with the direction of change in the power grid's operating status, it is considered an inconsistent repair. The number of time-series records that were corrected for inconsistencies and the total number of time-series records were counted, and the ratio of the number of time-series records that were corrected for inconsistencies to the total number of time-series records was used as the correlation detection result. If the associated detection result is greater than the preset detection result threshold, an abnormal alarm signal will be output. Conversely, if the condition is not met, then no abnormal alarm signal will be output.
10. A data analysis system for power grid data anomalies, used to implement the data analysis method for power grid data anomalies as described in any one of claims 1-9, characterized in that: It includes an anomaly monitoring module, a mechanism determination module, a status assessment module, and an anomaly alarm module. The functions of each module are as follows: The anomaly monitoring module is used to monitor power grid data, count the number of abnormal values in the power grid data, determine whether to connect to the data processing platform and perform repair processing based on the number of abnormal values, collect the number of data retransmissions on the data processing platform before performing repair processing, and pass the number of data retransmissions to the mechanism judgment module. The mechanism determination module selects to trigger the conventional processing mechanism or the repair analysis mechanism based on the number of data retransmissions. In the repair analysis mechanism, after the power grid data repair is completed and transmitted back, the time sequence record of the power grid data is obtained, the repair delay of each time sequence record is counted and the repair aggregation frequency is calculated, and the repair aggregation frequency is transmitted to the status assessment module. The status assessment module sets up a feature analysis window, accesses the repair type information of each time series record in the feature analysis window, performs interval aggregation analysis on the repair type information and generates repair repetition features, evaluates the data repair status of the time series record in combination with the repair aggregation frequency, and passes the data repair status to the anomaly alarm module. When the data repair status is abnormal, the abnormal alarm module retrieves the repair magnitude data and power grid operation data, performs correlation detection on the repair magnitude data and power grid operation data, obtains the detection results, determines whether there is abnormal repair based on the detection results, and outputs an abnormal alarm signal.