Node feature data correction method based on causal inference and sensitivity analysis
By applying causal inference and sensitivity analysis methods in the power grid system, the causal relationship and sensitivity between the characteristic data of the power grid nodes is solved, and the problem of limited correction effect caused by ignoring data dependencies in the prior art is solved, and more efficient and reliable data correction is achieved.
Patent Information
- Application Number
- CN202510265085.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art ignores the complex dependencies between data when correcting the characteristic data of power grid nodes, resulting in limited correction effect.
Using a method based on causal inference and sensitivity analysis, we collect and preprocess the characteristic data of the power grid system, identify the causal relationship between the data, build a causal graph, calculate the sensitivity of key data, sort and focus correction, and optimize the correction results through optimization algorithms.
By fully capturing the complex dependencies between data, more accurate and effective node feature data correction can be achieved, and the stability and reliability of power grid system operation are improved.
Smart Images

Figure CN120144923A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of power technology, and more specifically, relates to a method for correcting node feature data based on causal inference and sensitivity analysis. Background Art
[0002] With the increasing complexity and intelligence of power systems, the operation and management of power grids rely on a large amount of node feature data, including voltage, current, power, frequency, etc. However, due to problems such as the accuracy and stability of data acquisition devices, the collected data may have errors or inaccuracies, and these data errors may affect the accuracy and reliability of key operations such as power grid operation control, fault diagnosis, and optimization adjustment. Therefore, how to effectively correct the node feature data of the power grid and ensure the accuracy and reliability of the data is a technical problem that urgently needs to be solved.
[0003] Currently, traditional data correction methods include mean correction, linear regression correction, Kalman filtering, etc. These methods mainly correct data errors through statistical means or mathematical models. Although they can improve the accuracy of data to a certain extent, they often ignore the complex dependence relationships between data, especially causal relationships; and causal relationships are crucial for understanding the dependence and influence mechanisms between data, so the correction effect is limited. Summary of the Invention
[0004] The present invention provides a method for correcting node feature data based on causal inference and sensitivity analysis, aiming to solve the technical problem that the current method ignores the complex dependence relationships between data, resulting in limited correction effect.
[0005] The method for correcting node feature data based on causal inference and sensitivity analysis includes the following steps:
[0006] Step 1: Collect the feature data of relevant nodes in the power grid system and preprocess the collected feature data;
[0007] Step 2: Use a causal inference algorithm to identify the causal relationships between the node feature data of the power grid for the preprocessed data, and construct a causal graph based on the identified causal relationships;
[0008] Step 3: Use the global sensitivity analysis method to calculate the sensitivity of the key node feature data to the output of the causal model. According to the sensitivity analysis results, sort the node feature data to determine the data that needs to be corrected with emphasis;
[0009] Step 4: Conduct a preliminary correction on the data that needs to be corrected with emphasis, and perform a further adjustment on the preliminarily corrected data using historical data;
[0010] Step 5: Use an optimization algorithm to optimize the correction results in Step 4.
[0011] The method of the present invention preprocesses by collecting the characteristic data of relevant nodes in the power grid system to ensure the quality and consistency of the data; secondly, a causal inference algorithm is used to identify the causal relationships between the node characteristic data and construct a causal graph, thereby clarifying the dependence paths between the data; then, a global sensitivity analysis method is used to calculate the sensitivity of the key node characteristic data to the output of the causal model, and the data that needs to be corrected with emphasis is determined based on the sensitivity ranking; next, these key data are initially corrected and further adjusted in combination with historical data to improve the accuracy of the correction; finally, an optimization algorithm is used to optimize the correction result to ensure that the effect of the data correction reaches the best. Through causal relationship identification and sensitivity analysis, this method can comprehensively capture the complex dependencies between the data, thereby realizing more accurate and effective correction of the node characteristic data and improving the stability and reliability of the operation of the power grid system.
[0012] Preferably, step 2 includes the following steps:
[0013] Define the output variable: The output variable Y is defined as the performance index or state variable in the power grid system;
[0014] Construct a fully connected graph: In the initial state, a fully connected graph is constructed, and the nodes X 1 , X 2 , …, X n in the graph represent the respective characteristic data in the power grid system; Y represents the output variable; in the fully connected graph, there is an edge between each pair of nodes, including between the characteristic data nodes and between the characteristic data nodes and the output variable Y;
[0015] Gradually delete edges: Using conditional independence tests, gradually delete the edges that are independent under given conditions, where the conditional independence test is as follows:
[0016] P(X i ⊥ X j | S) = P(X i | S)P(X j | S);
[0017] In the formula: X i and X j respectively represent characteristic data i and characteristic data j; S represents the conditional set;
[0018] If the independence test passes, that is, X i and X j are independent under the given S condition, then delete the edge between X i and X j ;
[0019] Direction determination: Use the method of directional rules to determine the edges to construct a directed acyclic causal graph.
[0020] Preferably, the orientation rule includes:
[0021] If there exists a path X i →X k →X j , and there is no direct edge then add an edge X i →X j ;
[0022] If there exists a path X i →X j , and there is no other path between X i and X j , then keep the direction of the edge X i →X j .
[0023] Preferably, step 3 includes the following steps:
[0024] Calculate the sensitivity index based on the input variables and output variables in step 2, where the input variables are the characteristic data of each node in the power grid system obtained through the causal diagram construction in step 2;
[0025]
[0026] In the formula: V i represents the individual contribution of the input variable X i ; V ij represents the interactive contribution of the input variables X i and X j , and so on; Var(Y) represents the variance of the output variable Y;
[0027]
[0028] In the formula: S i represents the first-order sensitivity index, indicating the individual contribution of the input variable X i to Y;
[0029]
[0030] In the formula: Var(Y -i ) represents the conditional variance of the output Y when all input variables other than the fixed input X i are fixed; ST i represents the total effect sensitivity index;
[0031] Sort the input variables according to the total effect sensitivity index. The higher the ST i value, the greater the influence of the input variable on the output variable, and the higher the priority of correcting its data.
[0032] Preferably, step 4 includes the following steps:
[0033] Preliminary calibration: Smooth the time series data, and identify and remove or correct outliers in the data using the 3σ principle;
[0034] Further calibration using historical data: Calculate the mean and standard deviation of the historical data, construct a normal distribution range, and detect and correct the current data:
[0035]
[0036] In the formula: μ represents the mean of the historical data; represents the standard deviation of the historical data; X ′ represents the calibrated data; X represents the data after preliminary calibration; N represents the number of historical data samples.
[0037] Preferably, the objective function and constraint conditions of the optimization algorithm in step 5 are as follows:
[0038]
[0039] In the formula: MinimizeJ(X) represents minimizing the objective function; X represents the feature data set to be optimized; Y i represents the actual output value measured on the i-th side; represents the output value calculated based on the calibrated feature data; w i represents the weight of the i-th measurement;
[0040] Constraint conditions:
[0041]
[0042] Among them: and respectively represent the minimum and maximum values of the i-th feature data;
[0043] g k (X) ≤ 0, k = 1, 2,..., p;
[0044] In the formula: g k (X) represents the k-th system operation safety constraint function.
[0045] Preferably, the ion swarm algorithm is used to optimize the calibration result.
[0046] The beneficial effects of the present invention include:
[0047] The method of the present invention collects the characteristic data of relevant nodes in the power grid system and preprocesses it to ensure the quality and consistency of the data. Secondly, a causal inference algorithm is used to identify the causal relationships between the node characteristic data and construct a causal graph, thereby clarifying the dependence paths between the data. Then, the global sensitivity analysis method is used to calculate the sensitivity of the key node characteristic data to the output of the causal model, and the data that needs to be corrected with emphasis is determined based on the sensitivity ranking. Next, these key data are preliminarily corrected and further adjusted in combination with historical data to improve the accuracy of the correction. Finally, an optimization algorithm is used to optimize the correction results to ensure that the effect of data correction reaches the best. Through causal relationship identification and sensitivity analysis, this method can comprehensively capture the complex dependencies between data, thereby realizing more accurate and effective correction of node characteristic data and improving the stability and reliability of the operation of the power grid system. Brief Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a block diagram of the overall steps provided by the embodiment of the present invention. Detailed Embodiments
[0050] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] Refer to Figure 1 As shown, a further description of the optimal embodiment of the present invention is made;
[0052] A method for correcting node characteristic data based on causal inference and sensitivity analysis includes the following steps:
[0053] Step 1: Collect the characteristic data of relevant nodes in the power grid system and preprocess the collected characteristic data;
[0054] Collect the characteristic data of each node in the power grid system, including voltage, current, power, etc.
[0055] Preprocess the data, such as removing noise, handling missing values, etc.
[0056] Step 2: Use the causal inference algorithm to identify the causal relationships between the power grid node feature data for the preprocessed data, and construct a causal graph based on the identified causal relationships;
[0057] The said Step 2 includes the following steps:
[0058] Define the output variable: The output variable Y is defined as a performance index or a state variable in the power grid system;
[0059] Construct a fully connected graph: In the initial state, construct a fully connected graph, where the nodes X 1 , X 2 , …, X n represent the respective feature data in the power grid system; Y represents the output variable; in the fully connected graph, there is an edge between every pair of nodes, including between the feature data nodes and between the feature data nodes and the output variable Y;
[0060] Gradually delete edges: Use conditional independence tests to gradually delete the edges that are independent under given conditions, where the conditional independence test is as follows:
[0061] P(X i ⊥ X j | S) = P(X i | S)P(X j | S);
[0062] In the formula: X i and X j represent feature data i and feature data j respectively; S represents the conditional set;
[0063] If the independence test passes, that is, X i and X j are independent under the given condition S, then delete the edge between X i and X j ;
[0064] Direction determination: Use the method of directional rules to determine the edges, and construct a directed acyclic causal graph;
[0065] The said directional rules include:
[0066] If there exists a path X i → X k → X j , and there is no direct edge then add the edge X i → X j ;
[0067] If there exists a path X i → X j , and X i and X jIf there is no other path between them, then edge X is maintained i →X j in the direction of.
[0068] The input variables are the characteristic data of each node in the power grid system, and these data are the basis for creating the causal graph and conducting causal analysis. Typical input variables include: the voltage of each node, the current of each node, the active power of each node, the reactive power of each node, the frequency of each node, the power grid topology information, the load of each node, the power generation of each generating node;
[0069] The output variables are the results of causal analysis and simulation inference, usually the characteristic data of one or more target nodes or the overall performance indicators of the system. For example: the voltage level of the target node, the current of the target node, the system stability index, the fault probability;
[0070] The structure in the causal graph determines which variables are the direct or indirect causes of Y. The variance of each cause variable and its influence intensity on Y will affect the total variance of Y.
[0071] In this embodiment, by defining the output variables, constructing a fully connected graph, gradually deleting edges, and determining the direction, a directed acyclic causal graph is gradually constructed, which has the following beneficial effects: First, by defining the output variables as the performance indicators or state variables in the power grid system, the goal of causal inference is clarified, providing a clear direction for subsequent analysis; Second, by constructing a fully connected graph, all possible causal relationships are ensured to be taken into account, avoiding missing potential important relationships; Using conditional independence tests to gradually delete irrelevant edges significantly reduces the complexity of the graph, improves the calculation efficiency, and at the same time ensures that the causal relationships represented by the remaining edges are more credible; Finally, the direction of the edges is determined using the orientation rules to construct a directed acyclic graph, accurately reflecting the causal relationships between the characteristic data. This method not only improves the accuracy and reliability of causal relationship identification, but also improves the efficiency and effect of subsequent sensitivity analysis and optimization correction by simplifying the graph structure, making the correction of the power grid node characteristic data more scientific, reasonable, and efficient.
[0072] Step 3: Calculate the sensitivity of the key node characteristic data to the output of the causal model using the global sensitivity analysis method, and according to the sensitivity analysis results, sort the node characteristic data to determine the data that needs to be corrected with emphasis;
[0073] As a possible implementation manner of this embodiment, the step 3 includes the following steps:
[0074] Calculate the sensitivity index based on the input variables and output variables in step 2, where the input variables are the characteristic data of each node in the power grid system obtained through the causal graph construction in step 2;
[0075]
[0076] Where: V i represents the individual contribution of the input variable X i ; V ij represents the interaction contribution of the input variables X i and X j , and so on; Var(Y) represents the variance of the output variable Y;
[0077]
[0078] Where: X i represents the first-order sensitivity index, indicating the individual contribution of the input variable X i to Y;
[0079]
[0080] Where: Var(Y -i ) represents the conditional variance of the output Y when all input variables other than the fixed input X i are held constant; ST i represents the total-effect sensitivity index;
[0081] Sort the input variables according to the total-effect sensitivity index. The higher the ST i value, the greater the impact of the input variable on the output variable, and the higher the priority of correcting its data.
[0082] When determining the data that needs to be corrected with priority, give priority to the input variables with higher total-effect sensitivity indices because they have the greatest impact on the output variable, including direct effects and interaction effects with other variables; however, the first-order sensitivity index provides important information when analyzing the direct effects of variables. Therefore, data sorting and correction are based primarily on the total-effect sensitivity index, but the first-order sensitivity index is not completely ignored;
[0083] In this embodiment, the sensitivity of the key node feature data to the output of the causal model is calculated by the global sensitivity analysis method, and the node feature data is sorted according to the sensitivity analysis results to determine the data that needs to be corrected with priority, accurately identify the key variables that have the greatest impact on the system output, improve the correction efficiency, enhance the stability and reliability of the power grid system; optimize resource allocation to ensure that correction resources are used where they are most needed; provide scientific decision-making support to make the correction process more transparent and interpretable; and flexibly adapt to different scenarios to ensure the effectiveness of the method in various power grid systems. Through these measures, the efficiency and effectiveness of system data correction are significantly improved.
[0084] Step 4: Perform preliminary correction on the data that needs to be corrected with priority, and further adjust the preliminarily corrected data using historical data;
[0085] Step 4 includes the following steps:
[0086] Preliminary calibration: Smooth the time series data, and identify and remove or correct outliers in the data using the 3σ principle;
[0087] Further calibration using historical data: Calculate the mean and standard deviation of the historical data, construct the normal distribution range, and detect and correct the current data:
[0088]
[0089] where: μ represents the mean of the historical data; σ represents the standard deviation of the historical data; X ′ represents the calibrated data; X represents the preliminarily calibrated data; N represents the number of historical data samples.
[0090] In this embodiment, by performing preliminary calibration on the data that needs to be calibrated with emphasis, and further adjusting the preliminarily calibrated data using historical data, the accuracy and reliability of the node characteristic data of the power grid system can be significantly improved. First, the preliminary calibration step can quickly correct obvious abnormal data, initially improving the data quality; then, using historical data for further adjustment can make full use of existing experience and data patterns to make up for possible deficiencies in the preliminary calibration, so that the calibration result is closer to the actual situation. This dual-calibration process can effectively reduce the system performance deviation caused by data errors, enhance the stability and prediction ability of the power grid system. In addition, this method can also improve the adaptability and universality of the calibration method by combining historical data for calibration, enabling it to achieve good results in different application scenarios. Finally, the implementation of this step helps to improve the efficiency and effect of the overall calibration process, ensure that the data in the power grid system during operation is more accurate and reliable, and provide a solid foundation for subsequent optimization and adjustment.
[0091] Step 5: Optimize the calibration result in Step 4 using an optimization algorithm.
[0092] The objective function and constraint conditions of the optimization algorithm in Step 5 are as follows:
[0093]
[0094] where: Minimize J(X) represents minimizing the objective function; X represents the feature data set to be optimized, representing the variable to be optimized, such as node voltage, phase angle, etc.; Y i represents the actual output value measured on the i-th side, such as the actually measured voltage value; represents the output value calculated based on the calibrated feature data, such as the predicted voltage value calculated based on the prediction model; wi denotes the weight of the \(i\)-th measurement;
[0095] Select the characteristic data variable to be optimized, denoted as \(X = [X 1 , X 2 , …, X m \), which is the key node characteristic data determined by sensitivity analysis in step 4. For example, node voltage, phase angle, etc.;
[0096] Constraint conditions:
[0097]
[0098] Among them: and respectively represent the minimum and maximum values of the \(i\)-th characteristic data; for example, the range of node voltage may be between 0.95 and 1.05 times the nominal voltage;
[0099] g k (X) ≤ 0, k = 1, 2, …, p;
[0100] In the formula: g k (X) represents the \(k\)-th system operation safety constraint function. For example, the system load cannot exceed its rated capacity;
[0101] Optimize the correction result using the particle swarm algorithm;
[0102] Initialize the particle swarm: Initialize the particle swarm position and velocity. The position of each particle represents a potential solution;
[0103] X i (0), V i (0) for i = 1, 2, …, N;
[0104] In the formula: X i (0) represents the initial position of the \(i\)-th particle, representing the initial characteristic data set; V i (0) represents the initial velocity of the \(i\)-th particle; N represents the number of particles;
[0105] Evaluate the fitness function: Calculate the value of each fitness function, that is, the value of the objective function;
[0106] Update the particle velocity and position:
[0107] V i (t + 1) = w·V i (t) + c 1 ·r 1 ·(P i (t) - X i (t)) + c 2 ·r 2·(G(t) - X i (t));
[0108] X i (t + 1) = X i (t) + V i (t + 1);
[0109] Where: w is the inertia weight, which controls the continuity of the particle velocity; c 1 and c 2 represent the learning factors, which control the velocities of the particle moving towards its own best position and the global best position; r 1 and r 2 represent random numbers, which follow a uniform distribution in [0, 1]; P i (t) represents the historical best position of the i-th particle; G(t) represents the global best position;
[0110] Update the historical best position of each particle and the global best position:
[0111]
[0112] G(t + 1) = min i J(P i (t + 1));
[0113] If the preset number of iterations is reached or the change in the objective function value is less than the preset threshold, stop the iteration; otherwise, continue to execute the steps from evaluating the fitness function to updating the historical best position of each particle and the global best position;
[0114] In this embodiment, the correction result is further optimized by an optimization algorithm, and its objective function and constraint conditions ensure the accuracy of the correction data and the safe operation of the system. Specifically, this method maximizes the accuracy and effectiveness of data correction, adopts an objective function of minimizing errors, enables the corrected characteristic data to be as close as possible to the actual measured value, and reduces the negative impact brought by data errors. At the same time, considering the complexity of the power grid system and the requirements for safe operation, upper and lower limit constraints of characteristic data and system operation safety constraints are set to ensure that the corrected data fluctuates within a reasonable range and meets the safe operation standards of the power grid system. The beneficial effects of this method are as follows: the accuracy and reliability of data correction are effectively improved through the optimization algorithm, and unnecessary errors are reduced; at the same time, the stability and safety of the power grid system are ensured, providing a solid data foundation for system operation; in addition, the importance of different measured values can be flexibly adjusted through the weight factor, making the optimization process more targeted and practical. In short, this technical solution not only improves the efficiency and effect of data correction, but also guarantees the stable and safe operation of the power grid system, with significant practical value.
[0115] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A node feature data correction method based on causal inference and sensitivity analysis, characterized in that: The following steps are involved: Step 1: Collect characteristic data of relevant nodes in the power grid system and pre-process the collected characteristic data; Step 2: Using a causal inference algorithm on the preprocessed data to identify the causal relationship between the characteristic data of the power grid nodes, and constructing a causal graph based on the identified causal relationship; Step 3: Use the global sensitivity analysis method to calculate the sensitivity of key node feature data to the output of the causal model. According to the sensitivity analysis results, sort the node feature data and determine the data that needs to be corrected; Step 4: Perform preliminary correction on the data that need to be corrected, and use historical data to make further adjustments to the data after preliminary correction; Step 5: Use an optimization algorithm to optimize the correction result in step 4.
2. The node feature data correction method based on causal inference and sensitivity analysis according to claim 1 is characterized in that: The step 2 comprises the following steps: Defining output variables: The output variable Y is defined as the performance indicator or state variable in the power grid system; Construct a fully connected graph: In the initial state, construct a fully connected graph with nodes X1, X2, …, X n represents each characteristic data in the power grid system; Y represents the output variable; in the fully connected graph, there is an edge between each pair of nodes, including between characteristic data nodes and between characteristic data nodes and output variable Y; Gradually delete edges: Use the conditional independence test to gradually delete edges that are independent under given conditions. The conditional independence test is as follows: P(X i ⊥X j |S)=P(X i |S)P(X j |S); Where: X i and X j They represent feature data i and feature data j respectively; S represents the condition set; If the independence test passes, X i and X j Given S, if we are independent, then we can delete X. i and X j The edge between Direction determination: Use directional rules to determine the edges and construct a directed acyclic causal graph.
3. The node feature data correction method based on causal inference and sensitivity analysis according to claim 2 is characterized in that: The targeting rules include: If there is a path X i →X k →X j , and there is no direct edge Then add edge X i →X j ; If there is a path X i →X j , and X i and X j If there is no other path between them, keep edge X i →X j direction.
4. The node feature data correction method based on causal inference and sensitivity analysis according to claim 3 is characterized in that: The step 3 comprises the following steps: Calculate the sensitivity index based on the input variables and output variables in step 2, wherein the input variables are characteristic data of each node in the power grid system obtained by constructing the causal graph in step 2; Where: V i Represents the input variable X i The individual contribution of V ij Represents the input variable X i and X j The interactive contribution of, and so on; Var(Y) represents the variance of the output variable Y; Where: S i represents the first-order sensitivity index, which represents the input variable X i Contribution to Y alone; Where: Var(Y -i ) indicates that at a fixed input X i The conditional variance of the output Y when all input variables except ST i represents the total effect sensitivity index; The input variables were ranked according to the overall effect sensitivity index, ST i The higher the value, the greater the impact of the input variable on the output variable, and the higher the priority of correcting its data.
5. The node feature data correction method based on causal inference and sensitivity analysis according to claim 1 is characterized in that: The step 4 comprises the following steps: Preliminary correction: smooth the time series data and use the 3σ principle to identify and remove or correct outliers in the data; Further correction using historical data: Calculate the mean and standard deviation of historical data, construct a normal distribution range, and detect and correct current data: In the formula: μ represents the mean of historical data; σ represents the standard deviation of historical data; X ′ represents the corrected data; X represents the preliminary corrected data; N represents the number of historical data samples.
6. The node feature data correction method based on causal inference and sensitivity analysis according to claim 1 is characterized in that: The objective function and constraints of the optimization algorithm in step 5 are as follows: Where: Minimize J(X) represents the minimization objective function; X represents the feature data set to be optimized; Y i represents the actual output value measured on the i-th side; represents the output value calculated based on the corrected feature data; w i represents the weight of the i-th measurement; Constraints: in: and Respectively represent the minimum and maximum values of the i-th feature data; g k (X)≤0,k=1,2,…,p; Where: g k (X) represents the k-th system operation safety constraint function.
7. The node feature data correction method based on causal inference and sensitivity analysis according to claim 6 is characterized in that: The correction results were optimized using an ion cluster algorithm.