A multi-source heterogeneous data quality and security guarantee method for unified control of power grids
Patent Information
- Application Number
- CN202511722422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-11-21
AI Technical Summary
[0004]上述方案存在的主要问题是:严重依赖于预先制定的规则
本发明通过设定多尺度时间窗口,针对归一化数据计算其在每个窗口下的质量特征和物理约束特征,能够从不同时间尺度上观察数据的特征,短期波动可能反映测量噪声,而长期趋势性异常可能预示着设备退化或潜在的网络攻击。这种多尺度分析避免了单一时间窗口的局限性,使质量评估更加全面和灵敏;通过多尺度特征和共识博弈,不依赖预定义的异常模式,而是通过比对当前数据状态与历史正常状态基元的相似度来判断。即使是一种从未见过的攻击或故障,只要其导致数据的多尺度特征偏离了正常模式,就能被信息熵指标捕捉到,提高了对未知异常的检测能力。
Smart Images

Figure CN121580048B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid data analysis technology, specifically to a method for ensuring the quality and security of multi-source heterogeneous data for unified power grid management. Background Technology
[0002] With the increasing intelligence of power grids, multi-source heterogeneous data is being used more and more widely in unified power grid management and control. However, due to the diverse sources, heterogeneous protocols and formats of power grid data, and the dynamic and complex operating environment, traditional data quality verification methods mainly rely on predefined static rules, making it difficult to effectively identify unknown anomalies or complex attack patterns. Existing technologies, such as rule-based data verification systems, can detect known anomalies, but they cannot perceive the dynamic behavior of data over time and lack sensitivity to short-term fluctuations or long-term trend deviations. This results in limitations in dealing with new types of network attacks and slow equipment degradation, severely restricting the ability to ensure the quality and security of power grid data.
[0003] In the prior art, CN107992519A discloses a multi-source heterogeneous data verification system and method for smart grid big data, including: generating a verification task and a configuration file corresponding to the verification task based on the acquired verification content and pre-defined verification rules; when the verification task needs to be executed, executing the configuration file corresponding to the verification task; the pre-defined verification rules include: setting verification rules based on verification functions.
[0004] The main problems with the above approach are: it heavily relies on pre-defined rules. It can only detect known anomalies that can be described by rule patterns. For novel and complex network attacks or unknown device failure patterns, the method cannot effectively identify them because the rules cannot be pre-defined; it focuses on examining the intrinsic quality of individual data points or static records. It cannot perceive the dynamic behavior of data over time and cannot effectively identify situations with abnormally drastic short-term fluctuations or slow shifts in long-term trends.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a method for ensuring the quality and security of multi-source heterogeneous data under unified power grid management, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for ensuring the quality and security of multi-source heterogeneous data under unified power grid management includes the following steps: Step 1: Starting from the current moment, backtrack through three different time windows of different scales and obtain the time series data of the same multi-source heterogeneous data within each time window; Step 2: Calculate the data integrity and data volatility for each time window based on the time series data, and generate a joint feature vector based on the data integrity and data volatility of time windows at different scales; Step 3: Calculate the consensus degree between time windows pairwise, construct the game matrix, calculate the weight of the joint feature vector of different time windows based on the game matrix, and then construct the consensus feature vector according to the joint feature vector of time windows of different scales and their corresponding weights. Step 4: Extract the multi-source heterogeneous data from Step 1 in the normal working state from the historical data using a clustering algorithm, calculate the joint feature vector of the normal working state data in different working states as its state primitive, calculate the Euclidean distance between the consensus feature vector and each state primitive, and generate the probability that the consensus feature vector belongs to the working state corresponding to different state primitives. Step 5: Calculate the information entropy based on the probability that the current consensus feature vector belongs to the working state corresponding to different state primitives, and judge the quality and safety status of the multi-source heterogeneous data corresponding to the current consensus feature vector based on the information entropy.
[0008] Furthermore, the three different time windows specifically include a short-term window, a medium-term window, and a long-term window. The short-term window is 30 minutes long, the medium-term window is 24 hours long, and the long-term window is 7 days long. The same number of data monitoring points are set in each time window. The multi-source heterogeneous data refers to power grid data from different business systems and physical devices, and has heterogeneous protocols and formats.
[0009] Furthermore, the principle for generating joint feature vectors is as follows: The formula for calculating data integrity is: ; in, This indicates that multi-source heterogeneous data is in the first... Data integrity within a time window The index representing the type of time window, and , These correspond to short-term, medium-term, and long-term windows, respectively. Indicates the first The number of data monitoring points that monitored multi-source heterogeneous data within a given time window. This indicates the total number of data monitoring points within the time window; The formula for calculating data volatility is: ; in, Indicates the first Fluctuations in data within a certain time window This represents the index of the data monitoring point within the time window, and , Indicates the first The number of data monitoring points within a given time window Indicates the first Each data monitoring point monitored the values of multi-source heterogeneous data; The joint feature vector is: ,in, Indicates the first The joint feature vector of the time window.
[0010] Furthermore, the principle for constructing the game matrix is as follows: The consensus between each pair of time windows is calculated using the following formula: ; ; ; in, This indicates the degree of consensus between the short-term and medium-term windows. This indicates the degree of consensus between the short-term and long-term windows. This indicates the degree of consensus between the medium-term and long-term windows. This represents the joint feature vector of the short-term window. This represents the joint feature vector of the intermediate window. Represents the joint feature vector of a long-term time window; The game matrix is: ; in, Represents the game matrix, This indicates the degree of consensus between the medium-term and short-term windows, and , This indicates the degree of consensus between the long-term and short-term windows, and , This indicates the degree of consensus between the long-term and medium-term windows, and .
[0011] Furthermore, the principle for constructing consensus feature vectors is as follows: The principle for calculating the weights of the joint feature vector across different time windows is as follows: For the game matrix and weight vector, the following conditions are met: ,in Represents the original feature weight vector. Represents the largest eigenvalue; The characteristic equation is expressed as: ,in, Represent a 3×3 identity matrix; solve for The maximum value is . ; Substitution The results are as follows: The original feature weight vector is obtained by solving the problem. ,in, These represent the original weights of the short-term, medium-term, and long-term windows, respectively. Normalization is performed to obtain the actual weights of the short-term window, medium-term window, and long-term window. The formula for constructing the consensus feature vector is: ; in, Represents the consensus feature vector. The actual weights of the short-term, medium-term, and long-term windows, respectively.
[0012] Furthermore, the principle for generating consensus feature vectors based on the probability of them belonging to the working states corresponding to different state primitives is as follows: For each combination of short-term, medium-term, and long-term windows in historical data, a joint feature vector is generated and K-means clustering is performed to divide the data into several clusters. The top 10% of clusters are selected, and the joint feature vector of their centroids is used as the state primitive of that cluster. The formula for calculating the Euclidean distance between the consensus feature vector and each state primitive is: ; in, Represents the consensus feature vector and the first The Euclidean distance between each state primitive. Indicates the first A state primitive The index represents the state primitive, and , Indicates the number of state primitives; The formula used to determine the probability of generating consensus feature vectors belonging to the working states corresponding to different state primitives is as follows: ; in, This indicates that the consensus feature vector belongs to the first... The probability of a working state corresponding to each state primitive.
[0013] Furthermore, the formula for calculating information entropy is: ; in, Represents information entropy; The higher the value, the less secure the current data is. An information entropy risk threshold is set based on an expert scoring method. ,when At that time, it was determined that the current data posed a risk.
[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention, by setting multi-scale time windows, calculates the quality and physical constraint characteristics of normalized data within each window. This allows for observation of data characteristics across different time scales; short-term fluctuations may reflect measurement noise, while long-term trend anomalies may indicate equipment degradation or potential cyberattacks. This multi-scale analysis avoids the limitations of a single time window, making quality assessment more comprehensive and sensitive. Through multi-scale features and consensus game theory, it does not rely on predefined anomaly patterns but rather judges by comparing the similarity between the current data state and historical normal state primitives. Even a previously unseen attack or fault, as long as it causes the multi-scale characteristics of the data to deviate from the normal pattern, can be captured by the information entropy index, improving the detection capability of unknown anomalies.
[0015] The normal operation of the power grid itself includes multiple modes. This invention constructs a set of normal states that includes multiple normal operating conditions, avoiding misjudging data with different states but all belonging to normal states as abnormal states, thus improving the accuracy of the assessment. This invention also introduces information entropy, an indicator for measuring uncertainty, to judge the quality and safety status of the data. The higher the information entropy, the more uncertain or abnormal the data state is, thereby realizing dynamic and quantitative assessment of the data quality and safety status and supporting real-time risk identification and alarm. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the method flow of an embodiment of the present invention; Figure 2 This is a scatter plot showing the change of information entropy with the maximum attribution probability in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0019] Example: Please see Figures 1 to 2 The present invention provides a technical solution: A method for ensuring the quality and security of multi-source heterogeneous data under unified power grid management includes the following steps: Step 1: Starting from the current moment, backtrack through three different time windows of different scales and obtain the time series data of the same multi-source heterogeneous data within each time window; In this embodiment, the three different time windows specifically include a short-term window, a medium-term window, and a long-term window. The short-term window has a length of 30 minutes, the medium-term window has a length of 24 hours, and the long-term window has a length of 7 days. The same number of data monitoring points are set in each time window. The multi-source heterogeneous data represents power grid data from different business systems and physical devices, and has heterogeneous protocols and formats.
[0020] Power grid data exhibits strong time correlation and periodicity, making it impossible to fully characterize its behavior on a single time scale. Furthermore, it is susceptible to accidental interference or local noise, leading to misjudgments. Multi-scale analysis is employed, considering both dynamic fluctuations and steady-state characteristics of the data. A short-term window of 30 minutes is set to capture instantaneous fluctuations, reflecting real-time noise, sudden faults, or instantaneous attacks on the power grid. A medium-term window of 24 hours reflects the daily periodic characteristics of the data, such as peak electricity consumption and equipment operating cycles, used to identify periodic anomalies. A long-term window of 7 days reflects trend changes in the data, such as equipment performance degradation, systemic attacks, or the impact of environmental changes. When dividing time windows, the process starts from the current moment and traces backward; that is, a long-term window contains a medium-term window, and a medium-term window contains a short-term window.
[0021] Multi-source heterogeneous data refers to data from different sources with inconsistent structures and formats. The data sources are primarily business systems and physical devices. Business systems include several subsystems, such as SCADA systems, power management systems, distribution network management systems, and electricity consumption information collection systems. Physical devices include smart meters, phasor measurement units, remote terminal units, and various sensors. Heterogeneity includes protocol heterogeneity and format heterogeneity. Protocol heterogeneity indicates that the rules and standards followed by the data during transmission and communication differ. Different equipment manufacturers may use proprietary or different standard protocols; commonly used protocols include IEC60870-5-101 / 104, IEC 61850, Modbus, MQTT / HTTP, etc. Format heterogeneity indicates that the structure, syntax, and type of the data differ during storage and representation.
[0022] Step 2: Calculate the data integrity and data volatility for each time window based on the time series data, and generate a joint feature vector based on the data integrity and data volatility of time windows at different scales; In this embodiment, the principle for generating the joint feature vector is as follows: The formula for calculating data integrity is: ; in, This indicates that multi-source heterogeneous data is in the first... Data integrity within a time window The index representing the type of time window, and , These correspond to short-term, medium-term, and long-term windows, respectively. Indicates the first The number of data monitoring points that monitored multi-source heterogeneous data within a given time window. This indicates the total number of data monitoring points within the time window; Data integrity reflects the ratio between the number of data points actually collected and the total number of data points that should have been collected within a certain time window. It is used to measure the completeness of the data. A value close to 1 indicates that the data collection was complete with minimal missing data. A smaller value indicates problems such as data loss, transmission interruption, or acquisition failure. A higher value indicates higher data quality. It is directly proportional to the number of times data is collected within the time window.
[0023] The formula for calculating data volatility is: ; in, Indicates the first Fluctuations in data within a certain time window This represents the index of the data monitoring point within the time window, and , Indicates the first The number of data monitoring points within a given time window Indicates the first Each data monitoring point monitored the values of multi-source heterogeneous data; Data volatility reflects the dispersion of data values within a given time window. Lower volatility indicates more stable data with less change within that time window, and higher data quality. Higher volatility indicates drastic data changes within that time window, potentially indicating noise, anomalies, or attacks. In a power grid environment, data from normally operating equipment should exhibit certain regularity; abnormal fluctuations may indicate equipment failure, sudden load changes, or external interference. Higher data volatility suggests potential equipment failure, network attacks, communication interruptions, or data acquisition errors, resulting in lower data quality. Lower data volatility indicates better data quality. Furthermore, when analyzing data across multiple time windows, data volatility can distinguish between transient anomalies and trend anomalies. Transient anomalies are sudden and unexpected. Short-duration anomalies, such as transient pulse interference from a sensor or a brief packet loss in communication, significantly increase the standard deviation of data within a short-term window, leading to a significant increase in short-term volatility. In a medium-term window, the impact of these anomalies is greatly diluted, and medium-term volatility may only change slightly. These anomalies have almost no impact on the long-term window. Trend anomalies, which are long-lasting, slowly developing, or systematic deviations, such as measurement deviations due to equipment aging or persistently high loads in a region, are relatively stable in the short-term window due to their slow changes, resulting in normal or slightly elevated short-term volatility. In a medium-term window, this slow trend begins to emerge. Data no longer fluctuates around a stable mean but exhibits an upward or downward trend, increasing the dispersion of data in the medium-term window and significantly increasing medium-term volatility. The impact of these trend anomalies is most pronounced in the long-term window. Data deviates from normal conditions throughout the entire week, with a significantly larger fluctuation range than normal, resulting in significantly increased long-term volatility.
[0024] The joint feature vector is: ,in, Indicates the first The joint feature vector of the time window.
[0025] The joint feature vector combines data integrity and data volatility. Data integrity measures data availability; excessive data loss indicates a serious quality problem, potentially stemming from communication interruptions, equipment malfunctions, or malicious attacks. Data volatility reflects the reliability of output transmission; high data loss coupled with high volatility may indicate an extremely unstable communication link, affecting both data acquisition and system control. High data loss with normal volatility may indicate a fault in the data acquisition section, but the physical system itself is operating smoothly. Normal data loss but high volatility may indicate a fault in the physical system itself or an attack involving sophisticated cyberattacks. A joint feature vector is generated at each of the three time scales. They all reflect the data quality at a specific time scale.
[0026] Step 3: Calculate the consensus degree between time windows pairwise, construct the game matrix, calculate the weight of the joint feature vector of different time windows based on the game matrix, and then construct the consensus feature vector according to the joint feature vector of time windows of different scales and their corresponding weights. In this embodiment, the principle for constructing the game matrix is as follows: The consensus between each pair of time windows is calculated using the following formula: ; ; ; in, This indicates the degree of consensus between the short-term and medium-term windows. This indicates the degree of consensus between the short-term and long-term windows. This indicates the degree of consensus between the medium-term and long-term windows. This represents the joint feature vector of the short-term window. This represents the joint feature vector of the intermediate window. Represents the joint feature vector of a long-term time window; Consensus degree is used to quantify the degree of consistency among data quality characteristics at different time scales. Specifically, short-term, medium-term, and long-term windows reflect the instantaneous fluctuations, periodic patterns, and trend changes of data, respectively. The higher the consensus degree, the more consistent the data quality characteristics are at different time scales, and the more stable and reliable the data behavior is. If the consensus degree of a certain time window is significantly lower than that of other time windows, it indicates that the data at that time scale may be abnormal or interfered with. Cosine similarity measures the degree of closeness between two vectors. Cosine similarity reflects the difference in direction between the two vectors, which means that even if the data integrity and data volatility values of two time windows are very different, as long as they show the same trend (such as both showing high data integrity and low data volatility), their consensus degree is still very high.
[0027] The game matrix is: ; in, Represents the game matrix, This indicates the degree of consensus between the medium-term and short-term windows, and , This indicates the degree of consensus between the long-term and short-term windows, and , This indicates the degree of consensus between the long-term and medium-term windows, and .
[0028] The idea behind constructing a game theory matrix is to treat time windows at different time scales as participants in a cooperative game, ultimately determining their respective weights in the final judgment. Each participant's strategy is its joint eigenvector, and the consistency of strategies between two participants is represented by consensus degree. In the game theory matrix, the diagonal line has a value of 1 because its elements represent the consensus degree between each window and itself, which is constant at 1. The game theory matrix is used to determine the weight of each time window. The state of power grid data is dynamically changing. In some cases, short-term data may be unreliable due to noise; in other cases, long-term trends may lag due to slow system changes. By solving the game theory matrix and eigenvalues, the system can reward windows with higher consensus degrees with other windows and penalize windows with lower consensus degrees. If a time window has high consensus degrees with both other time windows, it indicates that its information is reliable, and its weight is higher. If a time window has low consensus degrees with other time windows, it indicates that its information contradicts the observations of other time-scale windows, and its weight is reduced. The mechanism of the game theory matrix inherently possesses the ability to denoise and resist interference. If the data in a time window becomes unreliable due to noise, transient failures, or targeted attacks, its feature vector will differ significantly from other relatively normal windows, leading to a decrease in its consensus with other windows. In subsequent weight calculations, this unreliable window will be automatically assigned a lower weight, thereby mitigating the negative impact of anomalous data on the final consensus feature vector.
[0029] The principle of constructing consensus feature vectors is as follows: The principle for calculating the weights of the joint feature vector across different time windows is as follows: For the game matrix and weight vector, the following conditions are met: ,in Represents the original feature weight vector. Represents the largest eigenvalue; The characteristic equation is expressed as: ,in, Represent a 3×3 identity matrix; solve for The maximum value is . ; The purpose of using the eigenvector method is to find a weighting scheme that keeps the relative importance of each time window consistent and stable. The largest eigenvalue represents the overall strength of consensus among each time window, and its corresponding eigenvector represents the relative importance of each time window in the consensus-building process. Will Substitution The results are as follows: The original feature weight vector is obtained by solving the problem. ,in, These represent the original weights of the short-term, medium-term, and long-term windows, respectively. Normalization is performed to obtain the actual weights of the short-term window, medium-term window, and long-term window. The formula for constructing the consensus feature vector is: ; in, Represents the consensus feature vector. The actual weights of the short-term, medium-term, and long-term windows, respectively.
[0030] Consensus feature vectors reflect the consistency of data quality across multiple scales, as well as the overall credibility and stability of the data. By weighting and fusing feature vectors from different time scales, a unified feature representation is obtained, capturing the overall performance of the data in the three dimensions of instantaneous, periodic, and trend. If the features of a certain time window are highly consistent with those of other windows, it indicates that the data quality of that window is reliable and its weight is high. If the features of a certain window differ significantly from those of other windows, its weight is low, indicating that the window may contain noise, anomalies, or attacks.
[0031] Step 4: Extract the multi-source heterogeneous data from Step 1 in the normal working state from the historical data using a clustering algorithm, calculate the joint feature vector of the normal working state data in different working states as its state primitive, calculate the Euclidean distance between the consensus feature vector and each state primitive, and generate the probability that the consensus feature vector belongs to the working state corresponding to different state primitives. In this embodiment, the principle for generating consensus feature vectors based on the probability of belonging to different working states of primitives is as follows: For each combination of short-term, medium-term, and long-term windows in historical data, a joint feature vector is generated and K-means clustering is performed to divide the data into several clusters. The top 10% of clusters are selected, and the joint feature vector of their centroids is used as the state primitive of that cluster. Under normal operating conditions, power grid data exhibits certain regularities and clustering characteristics. Similar operating states cluster together in the feature space. The K-means clustering algorithm automatically identifies these typical state primitives from historical data, representing the characteristic vectors of normal operating states. Each state primitive is a feature vector, representing the data quality characteristics under a typical operating state. During power grid operation, it is in a normal and stable state most of the time; therefore, the cluster containing the most feature vectors indicates a normal operating state, reflecting the data characteristics of the power grid under fault-free, attack-free, and normal communication conditions. Clusters in the top 10% of the size are selected because normal power grid operation itself includes various different operating states, such as periodic load changes (peak load during the day, low load at night). The data volatility and baseline during peak periods are completely different from those during low periods. Furthermore, there are significant differences in electricity consumption patterns and load levels between weekdays and weekends, and electricity consumption characteristics also vary across seasons. Therefore, multiple clusters should be selected to cover these normal operating states. The formula for calculating the Euclidean distance between the consensus feature vector and each state primitive is: ; in, Represents the consensus feature vector and the first The Euclidean distance between each state primitive. Indicates the first A state primitive The index represents the state primitive, and , Indicates the number of state primitives; Used to quantify the difference between the current data state and the historical normal working state. The smaller the value, the closer the current data state is to a certain historical normal working state. The larger the value, the more the current data state deviates from this historical normal working state.
[0032] The formula used to determine the probability of generating consensus feature vectors belonging to the working states corresponding to different state primitives is as follows: ; in, This indicates that the consensus feature vector belongs to the first... The probability of a working state corresponding to each state primitive.
[0033] The probability of the consensus feature vector belonging to the corresponding working state of different state primitives is calculated. This is done by comparing the current data's consensus feature vector with state primitives extracted from historical normal states to determine if the current data conforms to a known normal operating mode. The Euclidean distance between the consensus feature vector and each state primitive is transformed into a probability distribution, representing the likelihood of the current data belonging to each normal state, thus providing input for subsequent information entropy calculation. If a certain... A value close to 1 and others close to 0 indicates that the current state clearly belongs to a known normal working state, resulting in higher data quality and security. If all... If the values are close, it means that the current state cannot be clearly matched with the known normal working state, and the quality and safety of the data are abnormal. The calculation formula reflects that the smaller the Euclidean distance, The higher the value, the smaller the Euclidean distance, indicating greater similarity between the two states, and thus a higher probability of belonging to that state. The influence of the Euclidean distance is amplified through an exponential function. When I was very young, It will be very big, when When it is very large, A sharp decrease in value, approaching zero, helps to widen the probability gap between normal and abnormal states, making the judgment clearer.
[0034] Step 5: Calculate the information entropy based on the probability that the current consensus feature vector belongs to the working state corresponding to different state primitives, and judge the quality and security status of the multi-source heterogeneous data corresponding to the current consensus feature vector based on the information entropy.
[0035] In this embodiment, the formula for calculating information entropy is: ; in, Represents information entropy; The higher the value, the less secure the current data is. An information entropy risk threshold is set based on an expert scoring method. ,when At that time, it was determined that the current data posed a risk.
[0036] Information entropy is used to measure system uncertainty, reflecting the degree of match between the current data state and historical normal operating states, and also reflecting the degree to which power grid data can be understood and matched; if If one value is close to 1, and the others are close to 0, it indicates that the current data state is highly certain and belongs to a certain normal state. The value is lower; if all The similar values indicate that the current data state cannot be clearly categorized. A higher value indicates that the data is abnormal or has been interfered with; information entropy is used to quantify whether the probability distribution is uniform or biased. The more uniform the probability distribution, the better. The higher the value; Multiple typical data working scenarios are constructed, including normal steady state, noise interference, communication interruption, etc., and the information entropy of each scenario is calculated. The boundary between normal working state and abnormal working state is evaluated based on expert scoring, and the corresponding information entropy is determined as the information entropy risk threshold. Table 1 reflects the change of information entropy with the maximum attribution probability. The higher the maximum attribution probability, the more likely the current data is to belong to a known normal working state. As the maximum attribution probability increases, the information entropy generally shows a downward trend. The information entropy risk threshold is set to 0.7 to judge the data quality and safety status corresponding to different information entropies. Table 1. Information entropy changes with maximum attribution probability.
[0037] Once the system determines that the current data poses a risk, it issues an alert, isolates the data deemed risky to prevent it from entering subsequent processes, and takes measures such as data imputation, outlier correction, and discarding for the isolated data.
[0038] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0039] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0040] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0041] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for ensuring the quality and security of multi-source heterogeneous data under unified power grid management, characterized in that, The specific steps include: Step 1: Starting from the current moment, backtrack through three different time windows of different scales and obtain the time series data of the same multi-source heterogeneous data within each time window; Step 2: Calculate the data integrity and data volatility for each time window based on the time series data, and generate a joint feature vector based on the data integrity and data volatility of time windows at different scales; Step 3: Calculate the consensus degree between time windows pairwise, construct the game matrix, calculate the weight of the joint feature vector of different time windows based on the game matrix, and then construct the consensus feature vector according to the joint feature vector of time windows of different scales and their corresponding weights. Step 4: Extract the multi-source heterogeneous data from Step 1 in the normal working state from the historical data using a clustering algorithm, calculate the joint feature vector of the normal working state data in different working states as its state primitive, calculate the Euclidean distance between the consensus feature vector and each state primitive, and generate the probability that the consensus feature vector belongs to the working state corresponding to different state primitives. Step 5: Calculate the information entropy based on the probability that the current consensus feature vector belongs to the working state corresponding to different state primitives, and judge the quality and safety status of the multi-source heterogeneous data corresponding to the current consensus feature vector based on the information entropy; The principle for generating joint feature vectors is as follows: The formula for calculating data integrity is: in, This indicates that multi-source heterogeneous data is in the first... Data integrity within a time window The index representing the type of time window, and , These correspond to short-term, medium-term, and long-term windows, respectively. Indicates the first The number of data monitoring points that monitored multi-source heterogeneous data within a given time window. This indicates the total number of data monitoring points within the time window; The formula for calculating data volatility is: in, Indicates the first Fluctuations in data within a certain time window This represents the index of the data monitoring point within the time window, and , Indicates the first The number of data monitoring points within a given time window Indicates the first Each data monitoring point monitored the values of multi-source heterogeneous data; The joint feature vector is: ,in, Indicates the first The joint feature vector of the time windows; The principle of constructing the game matrix is as follows: The consensus between each pair of time windows is calculated using the following formula: in, This indicates the degree of consensus between the short-term and medium-term windows. This indicates the degree of consensus between the short-term and long-term windows. This indicates the degree of consensus between the medium-term and long-term windows. This represents the joint feature vector of the short-term window. This represents the joint feature vector of the intermediate window. Represents the joint feature vector of a long-term time window; The game matrix is: in, Represents the game matrix, This indicates the degree of consensus between the medium-term and short-term windows, and , This indicates the degree of consensus between the long-term and short-term windows, and , This indicates the degree of consensus between the long-term and medium-term windows, and ; The principle of constructing consensus feature vectors is as follows: The principle for calculating the weights of the joint feature vector across different time windows is as follows: For the game matrix and weight vector, the following conditions are met: ,in Represents the original feature weight vector. Represents the largest eigenvalue; The characteristic equation is expressed as: ,in, Represent a 3×3 identity matrix; solve for The maximum value is . ; Substitution The results are as follows: The original feature weight vector is obtained by solving the problem. ,in, These represent the original weights of the short-term, medium-term, and long-term windows, respectively. Normalization is performed to obtain the actual weights of the short-term window, medium-term window, and long-term window. The formula for constructing the consensus feature vector is: in, Represents the consensus feature vector. The actual weights of the short-term, medium-term, and long-term windows, respectively.
2. The method for ensuring the quality and security of multi-source heterogeneous data for unified power grid management according to claim 1, characterized in that: The three different time windows in step 1 specifically include a short-term window, a medium-term window, and a long-term window. The short-term window is 30 minutes long, the medium-term window is 24 hours long, and the long-term window is 7 days long. The same number of data monitoring points are set in each time window. The multi-source heterogeneous data refers to power grid data from different business systems and physical devices, and has heterogeneous protocols and formats.
3. The method for ensuring the quality and security of multi-source heterogeneous data for unified power grid management as described in claim 1, characterized in that: The principle behind generating consensus feature vectors and assigning them to the working states corresponding to different state primitives is as follows: For each combination of short-term, medium-term, and long-term windows in historical data, a joint feature vector is generated and K-means clustering is performed to divide the data into several clusters. The top 10% of clusters are selected, and the joint feature vector of their centroids is used as the state primitive of that cluster. The formula for calculating the Euclidean distance between the consensus feature vector and each state primitive is: in, Represents the consensus feature vector and the first The Euclidean distance between each state primitive. Indicates the first A state primitive The index represents the state primitive, and , Indicates the number of state primitives; The formula used to determine the probability of generating consensus feature vectors belonging to the working states corresponding to different state primitives is as follows: in, This indicates that the consensus feature vector belongs to the first... The probability of a working state corresponding to each state primitive.
4. The method for ensuring the quality and security of multi-source heterogeneous data for unified power grid management according to claim 3, characterized in that, The formula for calculating information entropy in step 5 is: in, Represents information entropy; The higher the value, the less secure the current data is. An information entropy risk threshold is set based on an expert scoring method. ,when At that time, it was determined that the current data posed a risk.
Citation Information
Patent Citations
Multi-source heterogeneous data verification system and method for intelligent power grid big data
CN107992519A
Incremental heterogeneous graph clustering method on basis of game theories
CN108399268A
Network security detection method and device for power Internet of Things, and electronic equipment
CN119172135A