Power system sensitive data identification method based on data inference relation
By applying a sensitive data identification method based on a probability graph model in the power system, combined with encryption and access control, the problem of sensitive data identification and protection in the power system is solved, and efficient and accurate identification and security protection of sensitive data are achieved.
Patent Information
- Application Number
- CN202510226471.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively identify and protect sensitive data in power systems, especially when inferred relationships between data exist, which can easily lead to confidential data leakage.
Using a probability graph model and a circular belief propagation algorithm based on data inference relationships, sensitive data in the power system is identified through preprocessing, hierarchical analysis, principal component analysis, linear mixed effect model and probability graph model construction, and is protected in combination with encryption and access control mechanisms.
It realizes intelligent identification of sensitive data in the power system, improves identification efficiency and accuracy, ensures the security of data during the sharing and disclosure process, and prevents the leakage and abuse of sensitive information.
Smart Images

Figure CN120162820A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power grid data security, and specifically relates to a sensitive data identification method for confidential data leakage caused by inference relationships between data. Background Art
[0002] In the current power data security framework, the definition of sensitive data mainly relies on manual classification methods. This method is not only time-consuming and laborious, but also prone to omissions. There is still a lack of intelligent sniffing technology to detect whether the data to be publicly disclosed or shared contains sensitive information.
[0003] Existing sensitive data identification methods mainly focus on the medical and financial fields, lacking identification methods for the power field. Most of the data content targeted by existing sensitive data identification methods involves personal privacy information and has obvious identification features, such as specific text information like identity ID, gender, address, and occupation. Usually, these methods mainly rely on natural language processing technology and use text analysis to identify sensitive information. For structured numerical data, especially in power data, the applicability of these methods is insufficient. Power data often exists in numerical form, such as power consumption records, equipment status data, etc. The identification of the sensitivity level of these data requires different methods and technologies.
[0004] At the same time, there are a large number of explicit or implicit inference relationships between power data, further increasing the security risk and identification difficulty of power data. Some data that are manually classified as directly publicly shareable may be used to infer non - publicly available confidential data. For example, different types of electrical appliances and equipment in the power system have specific voltage and current characteristics when working. For example, an electric motor has a large starting current when starting, while the current waveform of the load is relatively stable. By analyzing the characteristics of voltage and current data, the type and quantity of the load can be inferred. Although this information seems to be technical data individually, after comprehensive analysis, it can be used to identify the operating conditions and working status of specific equipment.
[0005] In summary, the current methods are difficult to meet the requirements of power data security, and there is an urgent need to propose new sensitive data identification methods to address these privacy and security challenges. Summary of the Invention
[0006] In view of the above - mentioned drawbacks and problems existing in the prior art, the present invention proposes a method for identifying sensitive data in a power system based on a probabilistic graph model of data inference relationships, which solves the problem of confidential data leakage caused by inference relationships and improves the security of sensitive data.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0008] A method for identifying sensitive data in a power system based on a data inference relationship probability graph model, comprising the following steps:
[0009] Step 1, preprocess the power grid data in the dataset, and determine the confidential data in the dataset according to industry standards and laws and regulations.
[0010] Step 2, in accordance with national standards (such as "Data Security Technology - Rules for Data Classification and Grading", GB / T 43697-2024), based on the analytic hierarchy process algorithm, and further combined with expert prior knowledge, formulate an evaluation standard for the sensitivity of power system data, and complete the initial sensitivity calibration of power data.
[0011] Step 3, based on the principal component analysis algorithm, calculate the correlation index between data, and obtain the similarity vector between data. Specifically:
[0012] Step 3.1, directly solve the similarity vector for data with the same dimension;
[0013] Step 3.2, use the principal component analysis technology to reduce the dimension of data with different dimensions to the same dimension;
[0014] Step 3.3, calculate the similarity vector of the data after dimension reduction, and supplement the vacancy of the similarity vector in Step 3.1;
[0015] Step 3.4, for different - dimensional data that cannot be dimension - reduced, introduce an intermediate dataset to indirectly calculate the similarity vector.
[0016] Step 4, based on the linear mixed - effect model, calculate the correlation metric index between data. Specifically:
[0017] Step 4.1, construct a linear mixed - effect model and initialize the model parameters;
[0018] Step 4.2, calibrate the training dataset according to the inferred relationship of the physical mechanism between data;
[0019] Step 4.3, define a loss function and use the gradient descent method to optimize the model parameters
[0020] Step 4.4, use the model to predict the correlation metric index of all data according to the similarity vector between data.
[0021] Step 5, combine the inferred relationship of the physical mechanism existing between data to construct a probability graph model of power system data.
[0022] Step 6, implement the loop belief propagation algorithm on the probability graph model, and determine whether each data is sensitive data according to the result. Specifically:
[0023] Step 6.1, constructing a potential matrix of class association between nodes based on the weight of edge association measurement in the probability graph;
[0024] Step 6.2, initializing all node information;
[0025] Step 6.3, selecting nodes and performing iterative calculations until convergence;
[0026] Step 6.4, obtaining the final sensitivity of all nodes, and setting a threshold to obtain the identification result of sensitive data.
[0027] Furthermore, the power data in the dataset mentioned in Step 1 are the basic system information (node data, branch data, etc.) and power flow calculation data of a multi-node power system at different times, including confidential data stipulated by industry standards and laws and regulations and low-secrecy-level data considered directly public.
[0028] Furthermore, the process of preprocessing the grid data in the dataset and determining the confidential data in the dataset mentioned in Step 1 includes: removing missing values, outliers, and duplicate data to ensure the integrity and consistency of the data, and determining the confidential data in the dataset according to regulations such as the "Basic Rules for the Disclosure of Power Market Information" and the "Power Market Information Disclosure Mechanism of PJM in the United States".
[0029] Furthermore, based on the analytic hierarchy process algorithm mentioned in Step 2, formulating an evaluation criterion for the sensitivity of power system data, including: constructing a hierarchical structure model, conducting expert scoring, determining its relative importance, and performing a consistency test on the scoring results, calculating the weights of various factors according to the judgment matrix passing the consistency test, and obtaining the initial calibration result of the sensitivity of power data according to the expert scoring results.
[0030] Furthermore, the similarity vector mentioned in Step 3 is a vector describing the similarity between data. Each component in the similarity vector should have a clear meaning and be able to explain a certain relationship or characteristic between the observed values. The Spearman correlation coefficient, Pearson correlation coefficient, and cosine similarity can be selected to construct the similarity vector.
[0031] Further, the construction of the probabilistic graphical model of power system data by inferring the relationships existing between the combined data based on physical mechanisms mentioned in step 5 means sorting out the data in the dataset to obtain the inferential relationships derived from physical mechanisms between the data, and representing them in a graphical model; representing various power data as a standard weighted graph G=(V, E), where the vertex set V contains different data types in the power system, each node represents a specific power data, and each directed edge in the edge set E represents the inferential relationship between two power data, and the correlation metric represents the strength of the inferential relationship. Taking the maximum value of the correlation metric between data that are not directly related physically as the threshold, for the edges with the correlation metric higher than the set threshold, add bidirectional edges to represent the data-driven inferential relationship, and construct a complete inferential relationship probabilistic graphical model.
[0032] Further, the Loopy Belief Propagation algorithm mentioned in step 6 is an approximate inference algorithm for probabilistic graphical models. It can run in a graph with loops and approximately calculate the marginal probability distribution of each node by iteratively propagating local information in the graph. The Loopy Belief Propagation algorithm has the advantages of high computational efficiency and easy implementation, and is widely used in complex inference tasks.
[0033] Further, the present invention performs data encryption processing on the determined confidential data and sensitive data. Specifically, it means that before the confidential data and the identified sensitive data are stored in the database and file system, the symmetric encryption algorithm AES-128 is selected for data storage encryption. The encryption key of AES-128 adopts a strict life cycle management strategy, is updated regularly and destroyed when not in use.
[0034] Further, a role-based access control mechanism can also be used to manage system data. Define corresponding roles (system administrator, data analyst, power dispatcher, etc.) for different users in the system, and define different permission levels for different user roles. Confidential data and sensitive data are only open to authenticated administrators or specific authorized users, and other users have no access rights.
[0035] Compared with the prior art, the present invention provides a method for identifying sensitive data based on data inference relationships, and has the following beneficial effects:
[0036] By introducing a probabilistic graphical model and the Loopy Belief Propagation algorithm, the present invention realizes the intelligent identification of sensitive data in the power system, overcomes the deficiency of relying on manual identification in traditional methods, and greatly improves the identification efficiency and accuracy.
[0037] Using the analytic hierarchy process algorithm and the principal component analysis algorithm, calibrate the sensitivity of the data and calculate the similarity, and construct a correlation metric in combination with the linear mixed effects model to ensure the high efficiency of model construction and calculation, and can quickly process large-scale power data.
[0038] During the process of data sharing and disclosure, this method can intelligently identify and protect sensitive data, ensuring that while achieving data openness and transparency, it prevents the leakage and abuse of sensitive information, providing security guarantees for the digital transformation of the power industry.
[0039] The present invention is applicable to various power system data scenarios, can discover potential sensitive data through inferential relationships, has broad application prospects, and improves the existing power system security system. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will describe in detail the embodiments of the present invention in conjunction with the drawings. As Figure 1 shown, the embodiments of the present invention provide the following specific steps:
[0043] Step 1: Preprocess the grid data in the dataset, and determine the confidential data in the dataset according to industry standards and laws and regulations;
[0044] Step 2: Based on the analytic hierarchy process algorithm, formulate an evaluation standard for the sensitivity of power system data, and complete the initial sensitivity calibration of power data;
[0045] Step 3: Based on the principal component analysis algorithm, calculate the correlation index between data, and obtain the similarity vector between data;
[0046] Step 4: Based on the linear mixed effects model, calculate the association metric between data,
[0047] Step 5: Combine the inferential relationships of the physical mechanisms existing between data, and construct a probabilistic graphical model of power system data;
[0048] Step 6: Implement the loopy belief propagation algorithm on the probabilistic graphical model, and determine whether each data is sensitive data according to the results.
[0049] Step 7: Perform data encryption and access control processing on the determined confidential data and sensitive data.
[0050] Specifically, in step 1, the power grid data in the dataset is preprocessed. The process of determining the confidential data in the dataset according to industry standards and laws and regulations includes: using Matpower to collect IEEE multi-node power system data to construct a dataset. In this example, the IEEE 30, 57, and 14-node power system data is used for subsequent implementation process demonstration. The missing values, outliers, and duplicate data in the dataset are cleaned. For missing values, interpolation method is used to fill them; for the data with constant values in the dataset, we retain or delete them according to whether they have actual physical meanings; for duplicate data, the duplicate items are directly deleted to ensure the integrity and consistency of the data.
[0051] According to the regulations on confidential data in the "Basic Rules for Power Market Information Disclosure" and the "Power Market Information Disclosure Mechanism of PJM in the United States", the generator operating status, voltage phase angle, load data, etc. involved in the dataset are marked as confidential data.
[0052] Specifically, in step 2, based on the analytic hierarchy process algorithm, the evaluation criteria for the sensitivity of power system data are formulated. The process of completing the initial sensitivity calibration of power data includes:
[0053] Formulate the evaluation criteria for the sensitivity of power system data based on the security attributes (confidentiality, integrity, availability) of the data. According to the characteristics of power grid data and business scenarios, the specific criteria include six aspects: leakage risk, access control, impact of data tampering, data consistency, system dependence, and impact of data loss.
[0054] Construct a hierarchical structure model. Each expert scores pairwise comparisons of each criterion and index on a scale of 1-9 to form a judgment matrix, calculate the maximum eigenvalue and the corresponding eigenvector of the judgment matrix, calculate the consistency index (CI) and the random consistency index (RI) according to the judgment matrix, and calculate the consistency ratio (CR):
[0055] where λ max is the maximum eigenvalue of the judgment matrix, and n is the order of the judgment matrix.
[0056] If CR < 0.1, the judgment matrix is consistent and passes the consistency test; otherwise, the scoring needs to be adjusted and recalculated.
[0057] For the judgment matrix that passes the consistency test, calculate the weights of each criterion and index. Use the eigenvector corresponding to the maximum eigenvalue for normalization processing to obtain the relative weights of each factor.
[0058] According to the weights of each factor, conduct preliminary sensitivity calibration on the power data. Multiply the sensitivity scores evaluated by experts for each data item by the corresponding weights and perform weighted summation to obtain the preliminary sensitivity value of each data item.
[0059] Specifically, step 3 is based on the principal component analysis algorithm. The process of calculating the correlation index between data and obtaining the similarity vector between data includes:
[0060] In order to better utilize the similarity between data to solve the correlation metric index between data, the construction of the similarity vector is very important. Each component in the similarity vector should have a clear meaning and be able to explain a certain relationship or characteristic between the observed values. At the same time, the similarity vector should describe the similarity between data as comprehensively as possible. Select the Spearman correlation coefficient, Pearson correlation coefficient, and cosine similarity to construct the similarity vector. That is, X = [Spearman correlation coefficient, Pearson correlation coefficient, cosine similarity].
[0061] First, calculate the similarity vector for the data with the same dimension in the dataset, and a similarity vector matrix with many vacancies can be obtained, denoted as similarity matrix one; for the data for which the similarity vector cannot be directly solved, use the principal component analysis technique to uniformly reduce the dimension of these data to the same dimension. Then, solve the similarity vector for the reduced-dimension data to obtain similarity matrix two; use similarity matrix two to fill the vacancies in similarity matrix one to obtain similarity matrix three;
[0062] At this time, there are still a small number of vacancies in similarity matrix three due to the data that cannot be reduced in dimension because multiple groups of different data cannot be obtained. For the problem that the similarity vector cannot be solved for two different-dimensional data A and B using the PCA technique, the data that can be directly or can be reduced in dimension through PCA and solve the similarity vector with data A and B at the same time is denoted as data C. Find all data C, and through data C as a bridge, indirectly infer the similarity between data A and data B. Multiply the similarity coefficient between data A and data C by the similarity coefficient between data B and data C, and select the maximum value as the similarity coefficient between data A and B, so as to construct the similarity vector between data A and B and obtain a complete similarity matrix.
[0063] Specifically, step 4 is based on the linear mixed effects model. The process of calculating the correlation metric index between data includes:
[0064] Construct a linear mixed effects model:
[0065] Y = μ + XW + S + ∈,
[0066]
[0067] where Y is the correlation metric to be solved, the final numerical value of the correlation metric between every two data, and the weight of the edge in the probabilistic graphical model, μ is the baseline of the correlation degree between data. Here, X = [sim1, sim2,...] is the similarity vector between any two data, and the elements in the vector are various correlation metrics between the data; W = [w1, w2,...] is the vector of the fixed-effect measure for the similarity metrics in the similarity vector, S is the random effect in the process of solving the correlation degree between data, and ∈ is the noise term. The distributions of S and ∈ respectively follow the normal distributions with mean zero and variances σ s , σ ε . Intuitively, this model evaluates the importance of the similarity attributes between data for measuring the correlation metric between two data, and calculates the final value of the correlation metric based on these similarity attributes.
[0068] Calibrate the correlation metric between some data according to the physical mechanism relationship between data to obtain training data for the estimation of parameters μ, W, σ s , σ ε . Use the gradient descent method to solve the optimization problem of the cost function in the model parameter solution. The loss function is defined as:
[0069] where is the output of the linear mixed effect model for the final correlation metric value of the i-th pair of two data, Y i is the expected correlation metric value of the two data, provided by the training set data, and N is the number of instances in the training data of the linear mixed effect model. By minimizing the loss function, the estimates of these parameters can be easily obtained based on the gradient descent algorithm. Then, substituting the solved similarity vector into the model, the similarity metrics between all data can be obtained.
[0070] Specifically, step 5 combines the physical mechanism inference relationship existing between data. The process of constructing the probabilistic graphical model of power system data includes:
[0071] Represent multiple power data as a standard weighted graph G = (V, E). In this graph, the set V represents the set of nodes, and each element in the set corresponds to a kind of power data. The set E represents the set of edges. Each edge connects two nodes and has a weight, and the weight represents the inference relationship between these two nodes.
[0072] Specifically, the point set V contains different data types in the power system, such as bus voltages, bus currents, etc. Each node represents a specific power data. This representation helps to clearly identify and distinguish different data types. Each edge in the edge set E represents the relationship between two power data, and the weight on the edge quantitatively reflects the strength of this relationship, represented by the calculated correlation metric.
[0073] Sort out the inference relationships based on the physical mechanisms between data to construct directed edges, and clarify the inference relationships between data by constructing directed edges, further elaborating on the data inference process under different physical mechanisms.
[0074] Specifically, the inference relationships can be divided into strong inference relationships and weak inference relationships according to their nature. A strong inference relationship means that there is a clear physical derivation relationship between two types of data. Based on the basic physical laws in the power system, such as Kirchhoff's laws and Ohm's law, the derivation calculation of two types of data can be realized. The strong inference relationships and their corresponding directed edge construction methods include:
[0075] In the power system, there is a clear relationship between the active power and reactive power of the bus load and the voltage amplitude and phase angle of the bus. According to the power flow equation, the magnitude and direction of power are closely related to the phase difference between the voltage amplitude and current, and the product of voltage and current determines the strength of the power flow. Thus, directed edges can be obtained from the bus voltage amplitude and phase angle to derive the active power and reactive power of the bus load; there is a direct derivation relationship between the active power and reactive power of the generator connected to the power grid and the operating voltage of the generator. The power output by the generator is jointly determined by its voltage and the system load, usually described by the power equation of the synchronous generator. Thus, directed edges can be obtained from the operating voltage of the generator to derive the active power and reactive power of the generator; the resistance and reactance of the branch determine the power loss through the branch. According to Ohm's law and the power loss model, the power loss is the superposition of the Joule loss when current passes through the resistance and the reactive power loss caused by reactance. Thus, directed edges can be obtained from the resistance and reactance of the branch to derive the power loss.
[0076] A weak inference relationship means that there is no direct physical derivation formula between data, but they can affect each other under certain constraints. The construction of such directed edges includes: for example, the operating range of the bus voltage limits the actual voltage. Although there is no direct derivation relationship, this range strongly constrains the actual voltage. Therefore, directed edges are constructed: maximum bus voltage → actual bus voltage, minimum bus voltage → actual bus voltage; although there is no direct derivation relationship between the operating state of the generator and the output power, the operating state limits the power output. For example, power cannot be provided in the shutdown state. Therefore, directed edges are constructed: generator state → generator active power, generator state → generator reactive power.
[0077] Strong inference relationships provide an actionable computational path for power flow based on clear physical laws, while weak inference relationships reveal potential operating constraints in the system. By constructing and analyzing directed edges, the accuracy of the probabilistic graph model can be improved.
[0078] Then, taking the maximum correlation metric value between data that is not directly physically related as the threshold, bidirectional edges are added between data nodes with higher correlation metrics to represent data-driven inference relationships, and a complete probabilistic graph model of inference relationships is constructed.
[0079] Specifically, step 6 implements the Loopy Belief Propagation algorithm on the probabilistic graph model. The process of determining whether each piece of data is sensitive data based on the results includes:
[0080] Construct the potential matrix of category associations between nodes, and the process is as follows:
[0081] 1) Construct the original potential matrix:
[0082] The initial potential matrix ψ of category associations between nodes is constructed based on the correlation metric weights of the edges in the probabilistic graph.
[0083] 2) Calculate the degree matrix:
[0084] The degree matrix D is a diagonal matrix, where the diagonal element D ii represents the degree of node i, that is, the sum of the edge weights connected to node i in the system:
[0085]
[0086] In the formula, ψ ij represents the probability that its adjacent data node j is a sensitive data node under the condition that data node i is a sensitive data node.
[0087] 3) Perform symmetric normalization:
[0088] Use the degree matrix to perform symmetric normalization on the potential matrix to obtain the normalized potential matrix ψ norm :
[0089]
[0090] In addition, φ i (Y i ) represents the prior probability that node i is of category Y, that is, the prior probability that a certain node in the probabilistic graph is a sensitive data node. This prior probability can be obtained from the results of the initial sensitivity calibration of the data nodes. The initial sensitivity of each data node is the prior probability that it is a sensitive data node.
[0091] Next, iterative calculations of Loopy Belief Propagation can be performed, and the specific steps are as follows:
[0092] 1) Initialize all node information:
[0093] Define that the sensitive data node belongs to category Y. Then, for each node i, initialize its belief as φ i (Y i ). Initialize the message m i→j (Y i ) to be uniformly distributed, representing the initial information from node i to its adjacent node j.
[0094] 2) Randomly select nodes and perform iterative calculations until convergence:
[0095] For each node i, calculate the message it receives from the adjacent node j:
[0096]
[0097] In the formula, denotes the summation over all possible categories Y k , ψ norm (Y j , Y i ) represents the probability that the adjacent node i is of category Y j under the condition that node j is of category Y i ; ∏ k∈邻居(j)\i m k→j (Y k ) denotes the product of the messages ∏ k∈邻居(j)\i m k→j (Y k ) from all neighbor nodes k except node i.
[0098] Normalize the calculated message m j→i (Y j ) to make it a valid probability distribution. Subsequently, update the belief φ i (Y i )
[0099]
[0100] In the formula, is the prior probability of node i, representing the initial sensitivity that node i is a sensitive data node. ∝ represents a proportional relationship, that is, the result of the right - hand expression needs to be normalized so that the sum of the belief values for all possible categories is equal to 1.
[0101] During the iterative process, determine whether the change in the message and the belief is less than the set threshold or whether the maximum number of iterations is reached. If either condition is met, it is considered that the algorithm has converged.
[0102] 3) After convergence, the belief of any node "i" being a sensitive data node is φ i (Y i ):
[0103] After the iterative calculation converges, the final belief φ of node i i (Y i ) represents the probability that this node is a sensitive data node. This probability combines the prior probability and the information propagated from adjacent nodes, reflecting the sensitivity of node i in the global network.
[0104] Specifically, the process of step 7 for implementing data encryption and access control processing on the determined confidential data and sensitive data includes:
[0105] Before the confidential data and the identified sensitive data are stored in the database and file system, the symmetric encryption algorithm AES-128 is selected for data storage encryption. The encryption key of AES-128 adopts a strict life cycle management strategy, is updated regularly and destroyed when not in use.
[0106] The role-based access control mechanism is adopted to manage the system data. Corresponding roles (system administrator, data analyst, power dispatcher, etc.) are defined for different users in the system, and different permission levels are defined for different user roles. The confidential data and sensitive data are only open to the authenticated administrator or specific authorized users, and other users have no access rights.
[0107] Through the above steps, the loop belief propagation algorithm based on the probabilistic graph model is implemented, which can effectively infer the probability that each node is a sensitive data node.
[0108] Multiple experiments are respectively carried out in the IEEE30 node system, the IEEE57 node system, and the IEEE14 node system to obtain the final sensitivity of each data node. The data nodes with the final sensitivity higher than 0.7 are recorded as the sensitive nodes identified by the present invention
[0109] In the analysis of the experimental results under different power systems and the same power system, the identified sensitive data nodes show high stability and consistency, verifying the applicability and robustness of the method. In order to verify the effectiveness of the identification results; a linear regression model is established, and the confidential data is inferred by using the identified sensitive data, and the calculated accuracy rate can exceed 70%, showing high effectiveness and reliability.
[0110] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying sensitive data in a power system based on data inference relationship, characterized in that: The following steps are involved: Step 1: pre-process the power grid data in the data set and determine the confidential data in the data set according to industry standards and laws and regulations; Step 2: According to national standards, the initial sensitivity calibration of power data is completed based on the hierarchical analysis algorithm; Step 3, based on the principal component analysis algorithm, calculate the correlation index between the data and obtain the similarity vector between the data; Step 4, based on the linear mixed effects model, calculate the correlation metric between the data; Step 5: Combine the inference relationship of the physical mechanism between the data to build a probabilistic graphical model of the power system data; Step 6: Implement the cyclic belief propagation algorithm on the probabilistic graphical model and determine whether each type of data is sensitive data based on the results.
2. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: In step 1, the power data in the data set includes basic system information and power flow calculation data of the multi-node power system at different times, including confidential data specified by industry standards and laws and regulations and low-level data that is considered to be directly public; The process of preprocessing the power grid data in the data set and determining the confidential data in the data set includes: removing missing values, abnormal values and duplicate data, ensuring the integrity and consistency of the data, and determining the confidential data in the data set, including motor operating status, voltage phase angle and load data.
3. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1 is characterized in that: In step 2, based on the hierarchical analysis algorithm, an evaluation standard for the sensitivity of power system data is formulated, including: According to the leakage risk, access control, impact of data tampering, data consistency, system dependency and impact of data loss, a hierarchical model is constructed, and experts score them to determine their relative importance. The scoring results are then checked for consistency. The relative weight of each factor is calculated based on the judgment matrix that passes the consistency test, and the initial calibration result of the power data sensitivity is obtained based on the expert scoring results.
4. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: In step 3, the Spearman correlation coefficient, the Pearson correlation coefficient and the cosine similarity are selected to construct a similarity vector, and each component in the similarity vector has a clear meaning and can explain a certain relationship or characteristic between the observations.
5. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: The step 3 specifically includes: Step 3.1, directly solve the similarity vector of data with the same dimension; Step 3.2, using principal component analysis technology to reduce the different dimensional data to the same dimension; Step 3.3, calculate the similarity vector of the data after dimensionality reduction, and fill in the gaps in the similarity vector in step 3.1; In step 3.4, for data of different dimensions that cannot be reduced in dimension, an intermediate data set is introduced to indirectly calculate the similarity vector.
6. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: In step 4, based on the linear mixed effects model, the calculation of the correlation measurement index between data specifically includes: Step 4.1, construct a linear mixed effects model and initialize model parameters; Step 4.2, calibrate the training data set based on the inferred relationship between the physical mechanisms of the data; Step 4.3, define the loss function and use the gradient descent method to optimize the model parameters; Step 4.4, use the model to predict the correlation measurement index of all data based on the similarity vectors between the data.
7. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: In step 5, a probabilistic graphical model of power system data is constructed by combining the physical mechanism inference relationship between the data, including: The data in the dataset are sorted out to obtain the inferred relationship between the data based on the physical mechanism, and represented by a graph model; various power data are represented as a standard weighted graph G=(V,E), where the point set V contains different data types in the power system, and each node represents a specific power data. Each directed edge in the edge set E represents the inferred relationship between two power data, and the correlation measurement index represents the strength of the inferred relationship. The maximum correlation measurement index value between data that have no direct association in the physical sense is used as the threshold. Bidirectional edges are added to the edges whose correlation measurement index is higher than the set threshold to represent the data-driven inferred relationship, and a complete inferred relationship probabilistic graph model is constructed.
8. The method for identifying sensitive data in a power system based on data inference relationship according to claim 7, characterized in that: dividing the inferred relationship into a strong inferred relationship and a weak inferred relationship; A strong inference relationship means that there is a clear physical inference relationship between two types of power data. Based on the basic physical laws in the power system, the inference calculation of the two types of power data can be realized. The corresponding directed edge construction methods include: According to the power flow equation, the directed edges for deriving the active power and reactive power of the bus load from the bus voltage amplitude and phase angle are obtained; According to the power equation of the synchronous generator, the directed edge of the active power and reactive power of the generator is derived from the working voltage of the generator. According to Ohm's law and the power loss model, the directed edge for deriving the power loss from the resistance and reactance of the branch is obtained; A weak inference relationship means that there is no direct physical derivation formula between the two types of power data, but they can affect each other under certain constraints. The corresponding directed edge construction includes: The maximum bus voltage → the actual bus voltage, the minimum bus voltage → the actual bus voltage; Generator status → generator active power, generator status → generator reactive power.
9. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: The step 6 specifically includes: Step 6.1, constructing a potential matrix of node-category associations based on the association metric weights of the edges in the probability graph; Step 6.2, initialize all node information; Step 6.3, select nodes and perform iterative calculation until convergence; In step 6.4, the final sensitivity of all nodes is obtained and the threshold is set to obtain the sensitive data identification result.
10. The method for identifying sensitive data in a power system based on data inference relationship according to claim 1, characterized in that: The following methods are used to encrypt and control access to confidential and sensitive data: Before storing confidential data and identified sensitive data in the database and file system, the symmetric encryption algorithm AES-128 is selected for data storage encryption; A role-based access control mechanism is used to manage system data, corresponding roles are defined for different users in the system, and different permission levels are defined for different user roles.