Low-voltage distribution network user branch connection identification method and system based on data driving method

By using a combination of stepwise regression algorithm and t-test in low-voltage distribution networks, iteratively updates the multilinear regression model and significance threshold, solving the problems of poor identification performance and hidden error effects in specific conditions in the existing technology, achieving higher recognition accuracy and recall.

CN120124006APending Publication Date: 2025-06-10POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182355.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing topology identification method for low-voltage distribution networks has poor performance under short electrical distances, light loads or balanced three-phase loads, and there are problems that hidden errors affect the recognition accuracy.

Method used

A data-driven method is adopted to build a multilinear regression model through a stepwise regression algorithm, combining t-test and hierarchical stepwise regression algorithms, and iteratively updates the model and significance threshold parameters to improve the accuracy of the identification results.

Benefits of technology

In the case of small hidden error, the recall and accuracy of branch connection recognition of low-voltage distribution network user is improved, and the uncertainty of hidden errors on the regression coefficient estimation is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124006A_ABST
    Figure CN120124006A_ABST
Patent Text Reader

Abstract

The invention discloses a low-voltage power distribution network user branch connection identification method based on a data driving method. The method comprises the following steps: S1, obtaining measurement data of stem leaf nodes of a low-voltage power distribution network; constructing a multi-linear regression model based on a stepwise regression algorithm; s2, obtaining a stem leaf node subset according to the measurement data and a multi-linear regression model; s3, obtaining a significance factor subset of the stem and leaf node subset based on a linear regression principle and t test; s4, correcting the stem and leaf node subset according to the significance factor subset to obtain a corrected stem and leaf node subset; s5, iteratively updating the multi-linear regression model according to the corrected stem and leaf node subset in combination with a hierarchical stepwise regression algorithm to obtain a final regression model; and S6, performing low-voltage distribution network user branch connection identification according to the final regression model. The invention further discloses a low-voltage distribution network user branch connection identification system based on the data driving method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of active low-voltage distribution network topology identification, and particularly to a method and system for identifying user branch connections in a low-voltage distribution network based on a data-driven method. Background Art

[0002] The LVDN (low voltage distribution network) is expected to meet the requirements of flexible configuration and coordinated operation of future source-load storage resources. However, the flexible configuration and coordinated operation of future source-load storage resources require a relatively high level of digitization of the LVDN, and also require the integrity and accuracy of the LVDN topology information. The existing LVDN topology measurement devices have a limited coverage range, and the network structure is complex and changes frequently, making the network topology measurement information of the LVDN usually unavailable or inaccurate for a long time.

[0003] The LVTI (Low Voltage Topology Identification) method is mainly used in the power system to identify and analyze the structure and connection mode of the low-voltage distribution network. By identifying the power network topology, the LVTI can provide a detailed view of the network structure, device status, and their connection relationships, which is important for improving the management, fault diagnosis, and recovery of the power system. In recent years, with the comprehensive coverage of advanced measurement systems in the LVDN, related research on the LVTI has gradually attracted wide attention. The existing data-driven LVTI methods can be divided into two categories: one is based on voltage correlation, and the other is based on energy conservation. However, the LVTI method based on voltage correlation has poor performance in the case of short electrical distance, light load, or balanced three-phase load. The LVTI method based on energy conservation has lower limitations than the LVTI method based on voltage correlation, does not have the problem of poor performance in individual cases, and considers hidden errors, and improves the accuracy of the consumer stage identification algorithm through the SR (Stepwise Regression) algorithm. However, obtaining a higher accuracy through the SR algorithm will lead to a reduction in the recall rate, and although hidden errors are considered, the lack of a precise definition of hidden errors will affect the accuracy of topology identification to a certain extent.

[0004] Therefore, there is an urgent need for a new technical solution to solve the technical problem of how to design a method for identifying user branch connections in a low-voltage distribution network that considers hidden errors and improves the existing identification algorithm. Summary of the Invention

[0005] The present invention provides a method and system for identifying user branch connections in a low-voltage distribution network based on a data-driven method, so as to solve the technical problem of how to design a method for identifying user branch connections in a low-voltage distribution network that considers hidden errors and improves the existing identification algorithm.

[0006] To achieve the above object, the present invention provides a method for identifying user branch connections in a low-voltage distribution network based on a data-driven method, including the following steps:

[0007] S1. Obtain the measurement data of the stem-leaf nodes of the low-voltage distribution network; construct a multiple linear regression model based on the stepwise regression algorithm;

[0008] S2. Obtain the stem-leaf node subset according to the measurement data and the multiple linear regression model;

[0009] S3. Obtain the significance factor subset of the stem-leaf node subset based on the principle of linear regression and t-test;

[0010] S4. Modify the stem-leaf node subset according to the significance factor subset to obtain the modified stem-leaf node subset;

[0011] S5. Iteratively update the multiple linear regression model according to the modified stem-leaf node subset combined with the hierarchical stepwise regression algorithm to obtain the final regression model;

[0012] S6. Identify the user branch connections in the low-voltage distribution network according to the final regression model.

[0013] Preferably, constructing the multiple linear regression model based on the stepwise regression algorithm includes:

[0014] Define Y as the vector of the current size measurement of the stem node, X as the design matrix of the current amplitude measurement values of the leaf nodes, β as the regression coefficient vector, and e as the error vector, construct the multiple linear regression model Y = Xβ + e, and set the significance threshold;

[0015] Setting the significance threshold includes:

[0016] Set the threshold λ for significance introduction entry ; set the threshold λ for significance removal remove .

[0017] Preferably, S2 includes:

[0018] S21. Sequentially introduce variables into the linear regression model. If the j-th independent variable x j satisfies the introduction criterion P 0j = λ entry according to its significance, then introduce this new variable; λ entry is the preset threshold for significance introduction;

[0019] S22. Each time a new variable is introduced, the old variables in the selected equation are tested one by one. If the non-significant exclusion condition P 0i = λ remove is satisfied, then remove the i-th independent variable xi , to ensure that all variables in the leaf and stem node subset X φ are significant; λ remove is the preset threshold for significance removal;

[0020] S23. Repeat S11 and S12 until no new variables can be introduced.

[0021] Preferably, S3 includes:

[0022] For a linear regression model based on the stepwise regression algorithm, if the error e satisfies a normal distribution, i.e., e ∼ N(0, σ 2 I), the least squares estimate of the regression is:

[0023] β * ∼ N(β, σ 2 (X T X) -1 )

[0024]

[0025] where β * is the least squares estimate of the regression coefficient; σ is the standard deviation of the residuals of a fitted linear regression model; X is the design matrix for the current amplitude measurements of the leaf nodes; c jj is the diagonal element of the matrix (X T X) -1 ; X T is the transpose of X; is the least squares estimate of the regression coefficient β * for the j-th independent variable in β j ; β

[0026] β * is an unbiased estimate of a regression coefficient and can be interpreted as having no systematic bias; thus, in different sample spaces, the estimated value of the regression coefficient may be very large or very small, and the bias can be positive or negative and statistically averages to zero;

[0027] The error of the regression coefficient estimate can be expressed as:

[0028]

[0029] where is the error of the regression coefficient estimate, reflecting the range of variation in different sample spaces; σ * is the residual of the unbiased estimate;

[0030] Since the value of σ is usually unknown, the unbiased estimate σ * is used as a substitute, i.e.:

[0031]

[0032] S SE = Y T Y - β *T X T Y

[0033] where S SE is the sum of squared residuals, and the magnitude of S SE reflects the goodness of fit between the actual data and the theoretical model in Y = Xβ + e; S SE The smaller the value, the better the fitting effect of the data and the model;

[0034] If is small, it can be considered that the least squares estimate of the regression coefficient takes a smaller and more accurate value; therefore, based on the estimate of the regression coefficient, the connectivity relationship between multiple SNs and multiple LNs can be determined, where SN is the stem node in the LVDN and LN is the leaf node in the LVDN; however, due to the existence of hidden errors, it will lead to a large S SE ; in the case of using the regression coefficient estimation method, its standard deviation may be very large, so the estimation of the regression coefficient has great uncertainty, and the accuracy of the traditional method based on the regression coefficient estimation is low, and optimization is required on this basis;

[0035] The estimated value of the regression coefficient corresponding to Y = Xβ + e should be significantly different from 0 and close to 1; which is equivalent to testing whether the first hypothesis holds:

[0036] H 0 : β j = 0, j ∈ C

[0037] where H 0 is the hypothesis event; j represents the j-th element; C is the index set of the leaf nodes;

[0038] When the first hypothesis holds, there is:

[0039]

[0040] In linear regression, the t-statistic is the test statistic of the analysis of variance method to test the significance of each component in the model; the t-statistic can be calculated as:

[0041]

[0042] where t T-N follows a t-distribution with T - N degrees of freedom;

[0043] At this time, the probability that the hypothesis H 0 holds is:

[0044] P0 = P(β j = 0) = P(t T-N >|t j-0 |)

[0045] where P 0 is the P-value of the t-test for β j = 0; similarly, the calculated P 1 is the P-value of the t-test for β j = 1; thus, the smaller the values of P 0 and P 1 , the lower the probabilities of β j = 0 and β j = 1;

[0046] According to the principle of linear correlation, for consumers with small errors and large loads, the expected value of the regression coefficient estimate should be close to 1, and the variance should be close to 0; at this time, it indicates that the probability of β j = 1 is relatively large, and the probability of β j = 0 is small; therefore, the significance factor ξ can be defined as:

[0047] ξ = ln(P 1 / p 0 )

[0048] where the value of ξ is in the range of (-∞, +∞). If P 1 > p 0 , then ξ > 0; in the extreme case, if P 1 = 1 and P 0 = 0, the probability of β j = 1 is relatively large; if P 1 < P 0 , then ξ < 0; in the extreme case, if P 1 = 0 and P 0 = 1, the probability of β j = 0 is relatively large; if P 1 = P 0 , then ξ = 0; the possibilities of β j = 0 or 1 will be the same;

[0049] By calculating the significance factors of each independent variable in the linear regression model, a subset ξ φ of the significance factors of X φ is obtained:

[0050]

[0051] where X φ is a subset of stem-and-leaf nodes, used to reflect the connectivity relationship between SN and multiple LNs; n φ is the number of significant independent variables in the subset X φ ; xφ (n φ ) is the n-th significant independent variable corresponding to the leaf node in subset X φ ; x φ (j) is the j-th significant independent variable corresponding to the leaf node in subset X φ ; ξ φ (j) is the significance factor corresponding to x φ (j); ξ φ (n φ ) is the significance factor corresponding to x φ (n φ (n φ ).

[0052] Preferably, S4 includes:

[0053] Define X # as the intersection of all stem-leaf node subsets:

[0054] X # = X 1 ∩X 2 ∩... ∩X M

[0055]

[0056] ξ # (j) = max{ξ|ξ φ (k), x φ (k) = x # (j), Φ ∈ J, k ∈ C Φ , j ∈ C #}}

[0057] where x # (j) is the LN corresponding to the j-th independent variable in the intersection X # ; ξ #max is the set of maximum values of the significance factors related to the independent variables in the intersection X # ; ξ # (j) is the maximum significance factor x φ of different X # ; X M is the intersection of the leaf node subsets attached to the M-th stem node; x # (j) is the j-th significant independent variable corresponding to the leaf node in subset X φ ; x # (n # ) is the n-th significant independent variable corresponding to the leaf node in subset X # ; ξ # (n # (n # ) is x # (n# ) corresponding significance factor; C # is the intersection X # index set of the middle leaf nodes; ξ φ (k) is x φ (k) corresponding significance factor; x φ (k) is the leaf node corresponding to the subset X φ the k-th significant independent variable in; Φ represents the Φ-th; k represents the k-th;

[0058] To improve the accuracy of the recognition result of the stepwise regression algorithm, the recognition result should be corrected according to ξ φ and ξ #max and based on the following rules:

[0059] For any X φ (j) ∈ X φ , if ξ φ (j) < 0, it indicates that the confidence of LN belonging to the Φ-th SN is low and should be removed from the stem-leaf node subset X φ ;

[0060] For any X φ (j) ∈ X φ , if ξ φ (j) = ξ # (j), it indicates that the confidence of LN belonging to the Φ-th SN is the highest and should be removed from other subsets of the stem-leaf nodes;

[0061] The stem-leaf node subset is corrected to obtain the corrected stem-leaf node subset X RΦ , which can be expressed as:

[0062]

[0063] where, x Rφ (j) is the LN corresponding to the j-th significant independent variable in X Rφ ; x Rφ (n Rφ ) is the n Rφ -th significant independent variable corresponding to the leaf node in the corrected subset X φ ; C RΦ is the index set of the leaf nodes in the corrected subset.

[0064] Preferably, S5 includes:

[0065] Define the updated model as Y’ = X’β + e, where Y’ is the observation vector of the dependent variable:

[0066]

[0067] where, I φThe injection current magnitude vector for the Φth stem node; I Dj The injection current magnitude vector for the jth leaf node;

[0068] X’ is the updated design matrix obtained according to the specified stem-leaf node dependency:

[0069]

[0070] where T is the number of measurement current data; C Φall is the set of LNs specifying the stem-leaf connection relationship:

[0071]

[0072] where X Rφ is the subset of corrected leaf nodes attached to the φth stem node; X RM is the subset of corrected leaf nodes attached to the Mth stem node;

[0073] In the hierarchical stepwise regression algorithm, when the connection relationships of the remaining LNs between the stem and leaves cannot be determined by the stepwise regression algorithm, that is it is necessary to increase the significance threshold to ensure that more topological information of the LNs can be determined; if the significant threshold increment is set to Δλ, then when the increment reaches a certain level, that is λ remove > λ max , pause the hierarchical iteration process and output the final recognition result; for the remaining and undetermined LNs, use voltage correlation analysis or on-site investigation methods to determine the stem-leaf connection relationship;

[0074] To fully evaluate the performance of the hierarchical stepwise regression algorithm, define the index precision Ω p and recall rate Ω r to measure the accuracy of the algorithm:

[0075]

[0076] where N output is the total number of LNs for which the hierarchical stepwise regression algorithm can determine the stem-leaf connection relationship; N correct is the total number of LNs for which the stem-leaf connection relationship is correctly determined in the output;

[0077] Perform iterative updates until the significance threshold reaches the maximum to obtain the final regression model.

[0078] The present invention also provides a low-voltage distribution network user branch connection recognition system based on a data-driven method for the method of the present invention. The system includes a first module, a second module, a third module, a fourth module, a fifth module, and a sixth module;

[0079] The first module is used to obtain the measurement data of the stem-leaf nodes of the low-voltage distribution network; and construct a multiple linear regression model based on the stepwise regression algorithm;

[0080] The second module is used to obtain the stem-leaf node subset according to the measurement data and the multiple linear regression model;

[0081] The third module is used to obtain the significance factor subset of the stem-leaf node subset based on the principle of linear regression and t-test;

[0082] The fourth module is used to correct the stem-leaf node subset according to the significance factor subset to obtain the corrected stem-leaf node subset;

[0083] The fifth module is used to iteratively update the multiple linear regression model according to the corrected stem-leaf node subset combined with the hierarchical stepwise regression algorithm to obtain the final regression model;

[0084] The sixth module is used to identify the user branch connection of the low-voltage distribution network according to the final regression model.

[0085] The present invention has the following beneficial effects:

[0086] The method for identifying the user branch connection of the low-voltage distribution network based on the data-driven method of the present invention, on the basis of the stepwise regression algorithm, maximizes the recall rate by iteratively updating the stepwise regression model and the significance threshold parameter, and also has high accuracy when the proportion of hidden errors is small. The significance factor based on the t-test considers the hidden error to correct the recognition result, reduces the uncertainty of the regression coefficient estimation caused by the hidden error, and thus improves the accuracy of the recognition result.

[0087] The system for identifying the user branch connection of the low-voltage distribution network based on the data-driven method of the present invention, which is used for the method of the present invention, has the same beneficial effects as the method of the present invention.

[0088] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The following will refer to the drawings to further elaborate on the present invention in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0090] Figure 1 is a schematic flowchart of a preferred embodiment of the present invention.

[0091] Figure 2 is a specific flowchart of the method of a preferred embodiment of the present invention.

[0092] Figure 3It is a schematic diagram of the topology of the LVDN test system according to the preferred embodiment of the present invention. Detailed implementation manners

[0093] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.

[0094] See Figures 1 to 2 , in the preferred embodiment of the present invention, a method for identifying user branch connections in a low-voltage distribution network based on a data-driven method is provided, including the following steps:

[0095] S1. Obtain the measurement data of the stem and leaf nodes of the low-voltage distribution network; construct a multiple linear regression model based on the stepwise regression algorithm.

[0096] Constructing a multiple linear regression model based on the stepwise regression algorithm includes:

[0097] Define Y as the vector for measuring the current size of the stem node, X as the design matrix for the current amplitude measurement values of the leaf nodes, β as the regression coefficient vector, and e as the error vector. Construct the multiple linear regression model Y = Xβ + e and set the significance threshold;

[0098] Setting the significance threshold includes:

[0099] Set the threshold λ for significance introduction entry ; set the threshold λ for significance removal remove .

[0100] S2. Obtain the stem and leaf node subsets according to the measurement data and the multiple linear regression model.

[0101] S2 includes:

[0102] S21. Sequentially introduce variables into the linear regression model. If the j-th independent variable x j satisfies the introduction criterion P 0j = λ entry according to its significance, then introduce this new variable; λ entry is the pre-set threshold for significance introduction;

[0103] S22. Each time a new variable is introduced, the old variables in the selected equation are tested one by one. If the non-significant exclusion condition P 0i = λ remove is satisfied, then remove the i-th independent variable x i to ensure that all variables in the stem and leaf node subset X φ are significant; λ remove is the pre-set threshold for significance removal;

[0104] S23. Repeat S11 and S12 until no new variables can be introduced.

[0105] S3. Obtain a subset of significance factors for the stem - leaf node subset based on the principle of linear regression and t - test.

[0106] S3 includes:

[0107] For a linear regression model based on the stepwise regression algorithm, if the error e satisfies a normal distribution, i.e., e ∼ N(0, σ 2 I), the least - squares estimate of the regression is:

[0108] β * ∼ N(β, σ 2 (X T X) -1 )

[0109]

[0110] where β * is the least - squares estimate of the regression coefficient; σ is the standard deviation of the residuals of a fitted linear regression model; X is the design matrix for the current amplitude measurements of the leaf nodes; c jj is the diagonal element of the matrix (X T X) -1 ; X T is the transpose of X; is the least - squares estimate of the regression coefficient β * for the j - th independent variable in β j ; β

[0111] β * is an unbiased estimate of a regression coefficient, which can be interpreted as having no systematic bias; thus, in different sample spaces, the estimated values of the regression coefficients may be very large or very small, and the bias can be positive or negative, and statistically averages to zero;

[0112] The error of the estimated value of the regression coefficient can be expressed as:

[0113]

[0114] where, is the error of the estimated value of the regression coefficient, reflecting the range of variation in different sample spaces; σ * is the residual of the unbiased estimate;

[0115] Since the value of σ is usually unknown, the unbiased estimate σ * is used as a substitute, i.e.:

[0116]

[0117] S SE = Y T Y - β *T X T Y

[0118] Among them, S SE is the sum of squared residuals, and the magnitude of S SE reflects the fitting degree between the actual data and the theoretical model in Y = Xβ + e; the smaller the value of S SE , the better the fitting effect between the data and the model;

[0119] If is small, it can be considered that the least squares estimate of the regression coefficient takes a smaller and more accurate value; therefore, based on the estimate of the regression coefficient, the connectivity relationship between multiple SNs and multiple LNs can be determined. SN is the stem node in the LVDN, and LN is the leaf node in the LVDN; however, due to the existence of hidden errors, a large S SE will be caused; in the case of using the regression coefficient estimation method, its standard deviation may be very large, so the estimation of the regression coefficient has great uncertainty, and the accuracy of the traditional method based on the regression coefficient estimation is low, and optimization is required on this basis;

[0120] The estimated value of the regression coefficient corresponding to Y = Xβ + e should be significantly different from 0 and close to 1; which is equivalent to testing whether the first hypothesis holds:

[0121] H 0 : β j = 0, j ∈ C

[0122] Among them, H 0 is the hypothesis event; j represents the jth element; C is the index set of the leaf nodes;

[0123] When the first hypothesis holds, there is:

[0124]

[0125] In linear regression, the t-statistic is the test statistic of the analysis of variance method to test the significance of each component in the model; the t-statistic can be calculated as:

[0126]

[0127] Among them, t T-N follows a t-distribution with T - N degrees of freedom;

[0128] At this time, the probability that the hypothesis H 0 holds is:

[0129] P 0 = P(β j = 0) = P(tT-N >|t j-0 |)

[0130] Among them, P 0 is the P-value of the t-test for β j = 0; Similarly, P 1 is the P-value of the t-test for β j = 1; Therefore, the smaller the values of P 0 and P 1 , the lower the probabilities of β j = 0 and β j = 1;

[0131] According to the principle of linear correlation, for consumers with small errors and large loads, the expected value of the regression coefficient estimate should be close to 1, and the variance should be close to 0; at this time, it indicates that the probability of β j = 1 is relatively large, and the probability of β j = 0 is small; Therefore, the significance factor ξ can be defined as:

[0132] ξ = ln(p 1 / P 0 )

[0133] Among them, the value of ξ is in the range of (-∞, +∞). If p 1 > P 0 , then ξ > 0; In the extreme case, if P 1 = 1 and P 0 = 0, the probability of β j = 1 is relatively large; If P 1 < P 0 , then ξ < 0; In the extreme case, if P 1 = 0 and P 0 = 1, the probability of β j = 0 is relatively large; If P 1 = P 0 , then ξ = 0; The possibility of β j = 0 or 1 will be the same;

[0134] By calculating the significance factor of each independent variable in the linear regression model, the significance factor subset ξ φ of X φ is obtained:

[0135]

[0136] Among them, X φ is the stem-and-leaf node subset, which is used to reflect the connectivity relationship between SN and multiple LNs; n φ is the number of significant independent variables in the subset X φ ; x φ (n φ ) is the leaf node corresponding to the subset Xφ the nth φ significant independent variable in; x φ (j) is the leaf node corresponding to the subset X φ the jth significant independent variable in; ξ φ (j) is the significance factor corresponding to x φ (j); ξ φ (n φ ) is the significance factor corresponding to x φ (n φ ).

[0137] S4. Modify the stem - leaf node subset according to the significance factor subset to obtain the modified stem - leaf node subset.

[0138] S4 includes:

[0139] Define X # as the intersection of all stem - leaf node subsets:

[0140] X # = X 1 ∩X 2 ∩... ∩X M

[0141]

[0142] ξ # (j) = max{ξ|ξ φ (k), x φ (k) = x # (j), Φ ∈ J, k ∈ C Φ , j ∈ C #}

[0143] where x # (j) is the LN corresponding to the jth independent variable in the intersection X # ; ξ #max is the set of maximum values of the significance factors related to the independent variables in the intersection X # ; ξ # (j) is the maximum significance factor x φ (j) of different X # ; X M is the intersection of the leaf node subsets attached to the Mth stem node; x # (j) is the leaf node corresponding to the subset X φ the jth significant independent variable in; x # (n # ) is the leaf node corresponding to the subset X # the nth # significant independent variable in; ξ # (n # ) is for x# (n # ) corresponding significance factor; C # is the intersection X # the index set of the leaf nodes in the middle; ξ φ (k) is x φ (k) corresponding significance factor; x φ (k) is the leaf node corresponding to the subset X φ the k-th significant independent variable in; Φ represents the Φ-th; k represents the k-th;

[0144] To improve the accuracy of the recognition result of the stepwise regression algorithm, the recognition result should be based on ξ φ and ξ #max be corrected based on the following rules:

[0145] For any X φ (j) ∈ X φ , if ξ φ (j) < 0, it indicates that the confidence of LN belonging to the Φ-th SN is low and should be removed from the stem-leaf node subset X φ ;

[0146] For any X φ (j) ∈ X φ , if ξ φ (j) = ξ # (j), it indicates that the confidence of LN belonging to the Φ-th SN is the highest and should be removed from other subsets of the stem-leaf nodes;

[0147] The stem-leaf node subset is corrected to obtain the corrected stem-leaf node subset X RΦ , which can be expressed as:

[0148]

[0149] where x Rφ (j) is the LN corresponding to the j-th significant independent variable in X Rφ ; x Rφ (n Rφ ) is the n Rφ -th significant independent variable corresponding to the leaf node in the corrected subset X φ ; C RΦ is the index set of the leaf nodes in the corrected subset.

[0150] S5. Iteratively update the multiple linear regression model according to the corrected stem-leaf node subset combined with the hierarchical stepwise regression algorithm to obtain the final regression model.

[0151] S5 includes:

[0152] Define the updated model as Y’ = X’β + e, where Y’ is the observation vector of the dependent variable:

[0153]

[0154] Among them, I φ is the injection current magnitude vector of the Φth stem node; I Dj is the injection current magnitude vector of the jth leaf node;

[0155] X’ is the updated design matrix obtained according to the specified stem-leaf node dependency:

[0156]

[0157] Among them, T is the number of measured current data; C Φall is the set of LNs specifying the stem-leaf connection relationship:

[0158]

[0159] Among them, X Rφ is the subset of corrected leaf nodes attached to the φth stem node; X RM is the subset of corrected leaf nodes attached to the Mth stem node;

[0160] In the hierarchical stepwise regression algorithm, when the connection relationships of the remaining LNs of the stem and leaf cannot be determined by the stepwise regression algorithm, that is it is necessary to increase the significance threshold to ensure that more topological information of the LNs can be determined; if the significant threshold increment is set to Δλ, then when the increment reaches a certain level, that is, λ remove > λ max , pause the hierarchical iteration process and output the final recognition result; for the remaining and undetermined LNs, use voltage correlation analysis or on-site investigation methods to determine the stem-leaf connection relationship;

[0161] To fully evaluate the performance of the hierarchical stepwise regression algorithm, define the index precision Ω p and recall rate Ω r to measure the accuracy of the algorithm:

[0162]

[0163] Among them, N output is the total number of LNs for which the hierarchical stepwise regression algorithm can determine the stem-leaf connection relationship; N correct is the total number of LNs for which the stem-leaf connection relationship is correctly determined in the output;

[0164] Perform iterative updates until the significance threshold reaches the maximum to obtain the final regression model.

[0165] After obtaining the final regression model, analyze the influence of the hidden error:

[0166] In the LVDN, serious problems such as power theft, PLC crosstalk, and communication interruption may lead to serious distortion of measurement data. The errors caused by these problems are usually unpredictable and hidden, which are called "hidden errors". The hidden error can be expressed as:

[0167] e h = e hq + e hz + e hk

[0168] where, e hq is the electricity theft error (ETEs); e hz is the PLC crosstalk error (PCEs); e hk is the communication error (CMEs); e h is the total hidden error.

[0169] (1) ETEs refer to the hidden errors introduced by power theft, usually by modifying the smart meter records or bypassing the electricity consumption so that the measured load current of the customer is zero or much less than the actual load current. During a certain time period S, the magnitude e hq and rate ε hq of ETEs can be expressed as:

[0170]

[0171] where, Q is the set of power theft consumers; I DQij is the measured load current of the j-th power theft consumer at the i-th moment; I Dij is the true value of the load current amplitude of the j-th LN at the i-th moment. Considering that the phase angle data cannot be measured by sensors or smart meters, the actual current amplitude I Dij is used to replace the current phasor

[0172] (2) PCE refers to the hidden error introduced by PLC crosstalk between adjacent LVDN segments, which causes the PLC sub-network information to reflect an incorrect root-leaf node connection relationship in this LVDN network segment. This relationship is manifested as the mixing of non-users in this network segment. The load current data collected by non-users in this segment seriously interferes with the identification of the stem-leaf node dependency relationship in this LVDN segment. During a certain time period S, the magnitude e hz and rate ε hq of PCEs can be expressed as:

[0173]

[0174] where, Z is the set of non-segmented consumers.

[0175] (3) CMEs refer to the hidden errors introduced by external interference or relay anomalies during the consumer load data acquisition process, which cause some consumers to have missing or zero load data at certain moments. Within a certain time period S, the size e of CMEs hk and the rate ε hk can be expressed respectively as:

[0176]

[0177] where K is the set of consumers with communication anomalies.

[0178] When considering the influence of different types of hidden errors, different methods M1 (Least Squares (LS)), M2 (Integer Quadratic Programming (IQP)), M3 (Least Absolute Shrinkage and Selection Operator lasso regression), and M4 (Hierarchical Stepwise Regression Algorithm (LSR)) are set, and their recognition results are compared. Several hidden error scenarios will be discussed below.

[0179] (1) Single type of hidden error: There is only one single category of hidden error, such as ETEs, PCEs, or CMEs, and the hidden error is added to the measurement data. For example, when there are only ETEs in LVDN, the current measurement values of some users will be reduced to 10% of the true measurement values. When the hidden error rate ε hq 、ε hz or ε hk varies between 0 and 10% respectively, calculate the accuracy rate of the model output results under methods M1 to M4 and the recall rate of method M4.

[0180] (2) Multiple types of hidden errors existing simultaneously: If multiple types of hidden errors exist simultaneously in LVDN, such as ETEs, PCEs, and CMEs. Therefore, when the hidden error rate ε h ranges from 0 to 25%, calculate and display the accuracy rates of methods M1 to M4, and the recall rate of method M4.

[0181] To compare the effectiveness of the proposed method M4 in identifying the stem - leaf node dependence at different levels, the following two situations need to be compared:

[0182] (1) Situation 1: Select the injected current in the secondary SN as the dependent variable, and select the outflow currents of all LNs as the independent variables.

[0183] (2) Situation 2: Select the injected current in the primary SN as the dependent variable, and select the outflow currents of all LNs as the independent variables.

[0184] Calculate the accuracy rates of methods M1 to M4, and the recall rate of method M4 in the second situation. Finally, it is obtained that:

[0185] If there are multiple types of hidden errors present simultaneously, such as ETEs, PCEs, and CMEs, when the error rate ε h varies from 0 to 25%, the precision and recall are calculated. The precision and recall of the algorithm in Case 1 are both better than those in Case 2. At the same time, if the granularity of topology recognition is finer, the difference in the significance test of the regression coefficient is greater. Therefore, in Case 2, M4 has better algorithm performance.

[0186] S6. Identify the user branch connections in the low-voltage distribution network according to the final regression model.

[0187] The method for identifying user branch connections in a low-voltage distribution network based on a data-driven approach in the present invention, based on the stepwise regression algorithm, maximizes the recall by iteratively updating the stepwise regression model and the significance threshold parameter, and also has a high accuracy when the proportion of hidden errors is small. The significance factor based on the t-test takes into account the hidden errors to correct the recognition results, reducing the uncertainty of the regression coefficient estimation caused by the hidden errors, thereby improving the accuracy of the recognition results.

[0188] In a preferred embodiment of the present invention, there is also provided a system for identifying user branch connections in a low-voltage distribution network based on a data-driven approach, for the method of the present invention. The system includes a first module, a second module, a third module, a fourth module, a fifth module, and a sixth module;

[0189] The first module is used to obtain the measurement data of the stem-and-leaf nodes of the low-voltage distribution network; construct a multiple linear regression model based on the stepwise regression algorithm;

[0190] The second module is used to obtain the stem-and-leaf node subset according to the measurement data and the multiple linear regression model;

[0191] The third module is used to obtain the subset of significance factors of the stem-and-leaf node subset based on the principle of linear regression and the t-test;

[0192] The fourth module is used to correct the stem-and-leaf node subset according to the subset of significance factors to obtain the corrected stem-and-leaf node subset;

[0193] The fifth module is used to iteratively update the multiple linear regression model according to the corrected stem-and-leaf node subset in combination with the hierarchical stepwise regression algorithm to obtain the final regression model;

[0194] The sixth module is used to identify the user branch connections in the low-voltage distribution network according to the final regression model.

[0195] The system for identifying user branch connections in a low-voltage distribution network based on a data-driven approach in the present invention, for the method of the present invention, has the same beneficial effects as the method of the present invention.

[0196] See Figure 3, in the preferred embodiment of the present invention, a LVDN test system with a total of 63 users is set up. In Figure 3 , C1 to C63 are users, that is, leaf nodes LN; S11 to S32 are meter boxes that can detect the current of multiple LNs; the feeder columns S1 to S3 are three-phase branches, that is, stem nodes SN; the distribution transformer is the root node RN. The measurement data of the stem and leaf nodes of the low-voltage distribution network is obtained by detecting the current data of the user load. Then, a multi-linear regression model is constructed based on the SR algorithm, and the stem and leaf node subsets are obtained according to the measurement data and the multi-linear regression model. Subsequently, the significance factor subset of the stem and leaf node subsets is obtained based on the principle of linear regression and the t-test. Then, the stem and leaf node subsets are corrected according to the significance factor subset to obtain the corrected stem and leaf node subsets, and the multi-linear regression model is iteratively updated in combination with the LSR algorithm according to the corrected stem and leaf node subsets to obtain the final regression model of the LVDN test system. The user branch connection of the low-voltage distribution network is identified according to the final regression model to determine the stem and leaf nodes to which each user belongs.

[0197] Based on the SR algorithm, the present invention adds a significance correction link to continuously iterate and layer to propose an optimized LSR algorithm, which is tested through the current data obtained by LVDN detection, so as to reduce the influence of hidden errors on the regression coefficients of the model and improve the recall rate of the user branch connection recognition method on the basis of maintaining the accuracy of the SR algorithm.

[0198] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method, characterized in that: The following steps are involved: S1, obtaining the measurement data of the stem-leaf nodes of the low-voltage distribution network; Construct multiple linear regression models based on stepwise regression algorithm; S2. Obtaining a stem-leaf node subset according to the measurement data and the multi-linear regression model; S3. Obtaining a subset of significant factors of the stem-leaf node subset based on the linear regression principle and t-test; S4, modifying the stem-leaf node subset according to the significance factor subset to obtain a modified stem-leaf node subset; S5, iteratively updating the multi-linear regression model according to the modified stem-leaf node subset combined with a hierarchical stepwise regression algorithm to obtain a final regression model; S6. Identify user branch connections in the low-voltage power distribution network according to the final regression model.

2. The method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method according to claim 1, characterized in that: The multi-linear regression model constructed based on the stepwise regression algorithm includes: Define Y as the vector of current size measurement of stem nodes, X as the design matrix of current amplitude measurement values ​​of leaf nodes, β as the regression coefficient vector, e as the error vector, construct the multi-linear regression model Y=Xβ+e, and set the significance threshold; The setting of the significance threshold comprises: Set the threshold λ for significant introduction entry ; Set the threshold λ for significant removal remove .

3. The method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method according to claim 2, characterized in that: The S2 includes: S21. Introduce variables into the linear regression model one by one. If the jth independent variable x j According to its significance, it meets the introduction standard P 0j =λ entry , then introduce the new variable; λ entry Thresholds introduced for pre-set significance; S22. Each time a new variable is introduced, the old variables in the selected equation are tested one by one. If the non-significant exclusion condition P is met, 0i =λ remove , then remove the i-th independent variable x i To ensure that the stem-leaf node subset X φ All variables in are significant; remove is the pre-set threshold for significance removal; S23. Repeat S11 and S12 until no new variables can be introduced.

4. The method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method according to claim 3, characterized in that: The S3 includes: For the linear regression model based on the stepwise regression algorithm, if the error e satisfies the normal distribution, that is, e~N(0,σ 2 I), the least squares estimate of the regression is: b * ~N(β,σ 2 (X T X) -1 ) Among them, β * is the least squares estimate of the regression coefficient; σ is the standard deviation of the residual of a fitted linear regression model; X is the design matrix of the current amplitude measurement values ​​for the leaf nodes; c jj is the matrix (X T X) -1 The diagonal elements of T is the transpose of X; is the least squares estimate of the regression coefficient β * The regression coefficient of the j-th independent variable in ; β j is the regression coefficient of the jth independent variable in the regression coefficient vector β; β * It is an unbiased estimate of the regression coefficient, which can be interpreted as having no systematic bias; therefore, in different sample spaces, the estimated value of the regression coefficient may be large or small, the bias may be positive or negative, and statistically the average is zero; The error in the estimated regression coefficient can be expressed as: in, is the error of the estimated value of the regression coefficient, reflecting the range of variation in different sample spaces; σ * is the residual of the unbiased estimate; Since the value of σ is usually unknown, an unbiased estimate of σ is used. * As an alternative, that is: S SE =And T Y-β *T X T AND Among them, S SE is the residual sum of squares, S SE The size reflects the degree of fit between the actual data and the theoretical model in Y = Xβ + e; S SE The smaller the value, the better the fit between the data and the model; like If the least squares estimate of the regression coefficient is smaller, it can be considered that the least squares estimate of the regression coefficient is smaller and more accurate; therefore, according to the estimation of the regression coefficient, the connectivity relationship between multiple SNs and multiple LNs can be determined, SN is the stem node in the LVDN, and LN is the leaf node in the LVDN; however, due to the existence of hidden errors, a larger S SE ; When the regression coefficient estimation method is used, its standard deviation may be large, so the estimation of the regression coefficient has great uncertainty. The accuracy of the traditional method based on regression coefficient estimation is low and needs to be optimized on this basis; The estimated value of the regression coefficient corresponding to Y=Xβ+e should be significantly different from 0 and close to 1; this is equivalent to testing whether the first hypothesis is true: H0:b j =0,j∈C Among them, H0 is the hypothesized event; j represents the jth element; C is the indicator set of the leaf node; But when the first assumption is established, there exists: In linear regression, the t-statistic is the test statistic of the ANOVA method to test the significance of each component in the model; the t-statistic can be calculated as: Among them, t T-N is a t-distribution with TN degrees of freedom; At this time, the probability that H0 is true is: P0=P(β j =0)=P(t T-N >|t j-0 |) Among them, P0 is β j =0; similarly, P1 is calculated to be β j =1; therefore, the smaller the values ​​of P0 and P1, the greater the j =0 and β j =1, the lower the probability; According to the linear correlation principle, for consumers with small errors and large loads, the expected value of the regression coefficient estimate should be close to 1 and the variance should be close to 0; this means that β j =1 is more likely, β j =0 is small; therefore, the significance factor ξ can be defined as: ξ=ln(P1 / P0) Among them, the value of ξ ranges from (-∞, +∞). If P1 > P0, then ξ > 0; in the extreme case, if P1 = 1 and P0 = 0, the probability of β j = 1 is relatively high; if P1 < P0, then ξ < 0; in the extreme case, if P1 = 0 and P0 = 1, the probability of β j = 0 is relatively high; if P1 = P0, then ξ = 0; the possibility of β j = 0 or 1 will be the same; By calculating the significance factor of each independent variable in the linear regression model, we can get X φ The significant factor subset ξ φ : Among them, X φ It is a subset of stem-leaf nodes, which is used to reflect the connectivity relationship between SN and multiple LNs; n φ For subset X φ The number of significant independent variables in x φ (n φ ) is a leaf node corresponding to the subset X φ The nth φ significant independent variable; x φ (j) is the leaf node corresponding to the subset X φ The jth significant independent variable in ξ φ (j) is x φ (j) the corresponding significance factor; ξ φ (n φ ) is x φ (n φ ) corresponding to the significance factor.

5. The method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method according to claim 4, characterized in that: The S4 includes: Define X # is the intersection of all stem-leaf node subsets: X # =X1∩X2∩…∩X M x # (j)=max{ξ|ξ φ (k),x φ (k)=x # (j),Φ∈J,k∈C Φ ,j∈C # } Among them, x # (j) is the intersection X # LN corresponding to the j-th independent variable in ξ; #max is the intersection with X # The maximum value set of significant factors related to the independent variables in ξ; # (j) is different X φ The maximum significance factor x # (j); X M is the intersection of the leaf node subsets attached to the Mth stem node; x # (j) is the leaf node corresponding to the subset X φ The jth significant independent variable in # (n # ) is a leaf node corresponding to the subset X # The nth # significant independent variable; # (n # ) is x # (n # ) corresponding to the significant factor; C # is the intersection X # The index set of the leaf nodes; ξ φ (k) is x φ (k) The corresponding significance factor; x φ (k) is the leaf node corresponding to the subset X φ The kth significant independent variable in ; Φ refers to the Φth; k refers to the kth; In order to improve the accuracy of the recognition results of the stepwise regression algorithm, the recognition results should be based on ξ φ and #max Make corrections based on the following rules: For any X φ (j)∈X φ , if φ (j) < 0, indicating that LN has a low confidence in the Φth SN and should be selected from the stem-leaf node subset X φ Removed; For any X φ (j)∈X φ , if φ (j) = ξ # (j), indicating that LN has the highest confidence in belonging to the Φth SN and should be removed from the other subsets of stem-leaf nodes; The stem-leaf node subset is modified to obtain a modified stem-leaf node subset X RΦ , which can be expressed as: Among them, x Rφ (j) is X Rφ LN corresponding to the jth significant independent variable in x Rφ (n Rφ ) is a leaf node corresponding to the modified subset X Rφ The nth φ significant independent variable; C RΦ is the index set of the leaf nodes in the modified subset.

6. The method for identifying user branch connections in a low-voltage power distribution network based on a data-driven method according to claim 5, characterized in that: The S5 includes: Define the updated model as Y'=X'β+e, where Y' is the observation vector of the dependent variable: Among them, I φ is the injection current magnitude vector of the Φth stem node; I Dj is the injection current magnitude vector of the jth leaf node; X' is the updated design matrix obtained according to the specified stem-leaf node dependencies: Where, T is the number of measured current data; C Φall For a LN set with a specified stem-leaf connection relationship: Among them, X Rφ is the modified subset of leaf nodes attached to the φth stem node; X RM is the modified subset of leaf nodes attached to the Mth stem node; In the hierarchical stepwise regression algorithm, when the connection relationship between the stem and the remaining LNs cannot be determined by the stepwise regression algorithm, that is, The significance threshold needs to be increased to ensure that more topological information of LN can be determined; if the significance threshold increment is set to Δλ, then when the increment reaches a certain level, that is, λ remove >λ max , suspend the hierarchical iteration process and output the final identification results; for the remaining and undetermined LNs, voltage correlation analysis or field survey methods are used to determine the stem-leaf connection relationship; In order to fully evaluate the performance of the hierarchical stepwise regression algorithm, the indicator accuracy Ω is defined p and recall Ω r To measure the accuracy of the algorithm: Among them, N output The total number of LNs that can determine the stem-leaf connection relationship using the hierarchical stepwise regression algorithm; N correct The total number of LNs with correctly determined stem-leaf connection relationships in the output; Iterative updating is performed until the significance threshold reaches the maximum limit, thereby obtaining the final regression model.

7. A low-voltage power distribution network user branch connection identification system based on a data-driven method, used in the method according to any one of claims 1 to 6, characterized in that: The system comprises a first module, a second module, a third module, a fourth module, a fifth module and a sixth module; The first module is used to obtain the measurement data of the stem-leaf nodes of the low-voltage distribution network; construct a multi-linear regression model based on the stepwise regression algorithm; The second module is used to obtain a stem-leaf node subset according to the measurement data and the multi-linear regression model; The third module is used to obtain the significant factor subset of the stem-leaf node subset based on the linear regression principle and t-test; The fourth module is used to modify the stem-leaf node subset according to the significance factor subset to obtain a modified stem-leaf node subset; The fifth module is used to iteratively update the multi-linear regression model according to the modified stem-leaf node subset combined with the hierarchical stepwise regression algorithm to obtain a final regression model; The sixth module is used to identify user branch connections in the low-voltage power distribution network according to the final regression model.