Line fault risk spatio-temporal distribution prediction method under bias-heterogeneous environmental factors
By using the method of using condition-related mode identification and probability fuzzy inference system in the prediction of fault risk of transmission line, the prediction inaccurate problems caused by data bias and heterogeneous environment in the prior art are solved, and a more accurate and flexible fault risk assessment is achieved.
Patent Information
- Application Number
- CN202510090299.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, when predicting the risk of transmission line failure, the prediction results are inaccurate due to data bias and heterogeneous environmental characteristics.
The spatial and temporal distribution prediction method for line failure risk under bias-heterogeneous environmental factors is used to obtain and normalize the line historical fault records, combine the condition-related pattern identification and probability fuzzy inference system, and learn discrete and continuous features in parallel, distinguish common and rare factors, and evaluate the risk of high-risk factors through fuzzy support and conditional fuzzy support.
It improves the accuracy and adaptability of fault risk prediction, can be flexibly applied in different environments and conditions, and avoids the excessive dependence of traditional methods on prior knowledge.
Smart Images

Figure CN120105226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a method for predicting the spatiotemporal distribution of line fault risks under bias-heterogeneous environmental factors. Background Art
[0002] Transmission lines are an important part of the power system, and their stability is directly related to the reliability and safety of power supply. With the continuous expansion of the scale of the power grid and the growth of global electricity demand, transmission lines are facing more and more operating pressure and potential failure risks. Among them, environmental changes such as extreme weather and natural disasters have posed higher challenges to the stability and safety of transmission lines. For example, natural events such as storms, lightning, and floods directly affect the physical structure and function of transmission lines, increasing the possibility of failures. At present, most line inspections still rely on manual work, which requires high inspection experience for operators, and it is difficult to carry out inspections in harsh environmental areas, and the operation risks are high. Therefore, based on fault history data, the high-precision positioning function of satellite Internet is used to provide accurate real-time location information, and the fault risk of transmission lines, especially its temporal and spatial distribution, is predicted in advance, which empowers the stability and reliability of new power system lines and becomes the key to improving power grid management efficiency and responsiveness.
[0003] At present, research shows that faced with the multi-source, heterogeneous, and equally distributed characteristics of fault data, some traditional models will produce a certain degree of error and fluctuation in the prediction results and prediction objects, resulting in inaccurate predictions. Summary of the invention
[0004] In view of this, the present invention provides a method for predicting the spatiotemporal distribution of line fault risks under biased-heterogeneous environmental factors, so as to at least solve the problem of inaccurate model prediction results in the prior art.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] A method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors comprises the following steps:
[0007] S1. Obtain a set of historical line fault records, normalize the historical line fault records and corresponding characteristic factors, construct an input mapping space between the historical line fault records and the corresponding fault characteristics, and obtain an evaluation database A;
[0008] S2. For data bias-heterogeneous environment, complete the parallel learning of discrete features and continuous features in fault features to obtain common factors and rare factors, including: learning discrete features through conditional correlation pattern identification CCPI and distinguishing common factors and rare factors that cause corresponding faults; learning continuous features through probabilistic fuzzy inference system PFIS and distinguishing common factors and rare factors that cause corresponding faults;
[0009] S3. A fuzzy conditional correlation pattern recognition model FCCPI was established. The high-risk HR factors and rare high-risk RHR factors were extracted from common factors and rare factors respectively through fuzzy support and conditional fuzzy support. The fuzzy support and conditional fuzzy support were normalized and then integrated to obtain the support score. The risk level was quantified through the support score.
[0010] Preferably, the specific contents of constructing the input mapping space in S1 include:
[0011] The collection of m historical fault records is denoted as D = {t 1 ,…,t i ,…,t m}, where i = 1,…,m,t i is the i-th historical fault record;
[0012] The feature set is denoted as F = {f 1 ,…,f j ,…,f n}, where j = 1,…,n, n is the number of features, f j represents the jth feature, each feature is composed of a set of factors, denoted as f j ={c 1,j ,…,c i,j ,…,c m,j}∈F,c i,j Represents feature f j A factor in
[0013] will be and When X occurs, the relevant pattern that Y must occur is expressed as X→Y, where X is the conditional factor set, which is a subset of the feature set F, and Y is the target factor set;
[0014] The target factor set is denoted as Y = {y 1 ,…,y i ,…,y m}, where y i is the fault handling result, indicating a fault record t iThe result of whether the current fault is successfully handled is recorded in the fault handling result, which includes: successful fault handling Y(S), barely successful Y(P) or failed Y(F). Therefore, y i =Y(r)∈{Y(S),Y(P),Y(F)};
[0015] Then the evaluation database A is:
[0016]
[0017] Preferably, in S2, the discrete features are learned by conditional correlation pattern identification CCPI, and the specific contents of distinguishing common factors and rare factors causing the corresponding faults include:
[0018] The scores are calculated according to the association rules and compared with the preset thresholds. Then, the conditional factor set X is decomposed to obtain common factor subsets and rare factor subsets corresponding to the discrete features.
[0019] Preferably, in S2, the continuous features are learned by the probabilistic fuzzy inference system PFIS, and the specific contents of distinguishing the common factors and rare factors causing the corresponding faults include:
[0020] The continuous features are learned through the probabilistic fuzzy inference system PFIS, and the probability distribution function PDF of each continuous feature is divided into rare degrees to obtain common factors and rare factors.
[0021] Preferably, the specific content of S3 includes:
[0022] S31. Expand the correlation pattern X→Y to obtain the expanded correlation pattern:
[0023] X,P→Y,Q
[0024] Where P = {p 1,1 ,p 1,2 ,…,p i,j ,…,p m,n} and Q = {q 1 ,q 2 ,…,q i ,…,q m} are fuzzy sets corresponding to X and Y, respectively, where p i,j ,q i are the factors in the fuzzy set of P and Q respectively; the expanded correlation model means that if X is related to P, then Y is considered to be related to Q;<X,P> Represents related feature-fuzzy set pairs;
[0025] S32. Input common factors and their fuzzy sets into the fuzzy support S c In the model, the corresponding fuzzy support is calculated, where the fuzzy support is greater than the fuzzy support voting threshold Ths The common factor is the HR factor, and the fuzzy support of the HR factor After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score;
[0026] Input the rare factors and their fuzzy sets into the conditional fuzzy support S r In the model, the corresponding conditional fuzzy support is calculated, where the conditional fuzzy support is greater than the fuzzy support voting threshold Th s The rare factors are RHR factors, and the fuzzy support of RHR factors is After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score;
[0027] The support score ranges from (0 to 1), where the closer it is to 1, the higher the risk level of failure.
[0028] Preferably, the fuzzy support model S in S32 c The specific contents include:
[0029] Fuzzy support model S c for:
[0030]
[0031] Where: B() represents the cardinality of fault records after preprocessing, Th s The voting threshold indicating support, t i (c i,j ) indicates t i Medium i,j The value of Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree of for Abbreviation of .
[0032] Preferably, the conditional fuzzy support model S in S32 r The specific contents include:
[0033] Conditional fuzzy support model S r for:
[0034]
[0035] In the formula, X r is a rare factor set, is the coefficient:
[0036]
[0037] in, Represents the fuzzy set p i,j The corresponding fuzzy weight of .
[0038] Preferably, S3 also includes S33:
[0039] Verify the high-risk HR factors and rare high-risk RHR factors screened out in S32:
[0040] Let the characteristic fuzzy set pair<X,P> and<Y,Q> Another characteristic fuzzy set<Z,L> A subset of And Z = X ∪ Y, And L=P∪Q, then Z={z 1 ,z 2 ,…,z i ,…,z m}, L = {l 1 ,l 2 ,…,l i ,…,l m} is a fuzzy set associated with Z, where z i ,l i are the i-th related factors in Z and L respectively;
[0041] For the HR factor, the three cases of fault handling results Y(r) are evaluated by the fuzzy certainty evaluation model Ce c and fuzzy correlation evaluation model Co c Find two characteristic fuzzy set pairs<<X,P> ,<Y,Q> >The fuzzy certainty and fuzzy relevance of the fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the fuzzy correlation is greater than the correlation evaluation threshold Th co Then the input factor is determined to be the HR factor;
[0042] At the same time, for the RHR factor, the three cases of fault handling results Y(r) are respectively analyzed by the conditional fuzzy certainty model Ce r and fuzzy correlation evaluation model Co r Find two characteristic fuzzy set pairs<<X,P> ,<Y,Q> >The conditional fuzzy certainty and conditional fuzzy relevance, the conditional fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the conditional fuzzy relevance is greater than the relevance evaluation threshold Th co Then the input factor is determined to be the RHR factor.
[0043] Preferably, the specific content of S33 includes:
[0044] Fuzzy Certainty Evaluation Model Ce c for:
[0045]
[0046] Where: Th ce represents the voting threshold of the certainty measure, t i (z i ) is z i In t i The value in or Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree, M li or M li [t i (z i )] represents the fuzzified membership of the input membership function to the fuzzy set li;
[0047] Conditional Fuzzy Certainty Model Ce r for:
[0048]
[0049] Where: X r and Z r denote the rare factor set and its associated fuzzy set, respectively. and The two coefficients are:
[0050]
[0051] in, and Respectively represent the fuzzy set p i,j and fuzzy set l i The corresponding fuzzy weight of
[0052] Fuzzy correlation evaluation model Co c for:
[0053]
[0054] Where: Cov c (X,Y) is the covariance, Var c (X) and Var c (Y) is the variance;
[0055] Cov c (X,Y)=E c [<Z,L> ]-E c [<X,P> ]×E c [<Y,Q> ]
[0056] Among them, E[…] represents the expectation:
[0057]
[0058] and The two coefficients are:
[0059]
[0060] Var c (X) and Var c (Y) are:
[0061] Var c (X) = E c [<X,P> 2 ]-E c [<X,P> ] 2
[0062] Var c (Y) = E c [<Y,Q> 2 ]-E c [<Y,Q> ] 2
[0063]
[0064]
[0065] Conditional fuzzy correlation evaluation model Co r for:
[0066]
[0067] In the formula, Cov r (X,Y) is the covariance, Var r (X) and Var r (Y) is the variance;
[0068] Cov r (X,Y)=E r [<Z,L> ]-E r [<X,P> ]×E r [<Y,Q> ]
[0069] Among them, E[…] represents the expectation:
[0070]
[0071]
[0072] Var r (X) and Var r (Y) are:
[0073] Var r (X) = Er [<X,P> 2 ]-E r [<X,P> ] 2
[0074] Var r (Y) = E r [<Y,Q> 2 ]-E r [<Y,Q> ] 2
[0075]
[0076] It can be seen from the above technical solution that, compared with the prior art, the present invention discloses a method for predicting the spatiotemporal distribution of line fault risks under biased-heterogeneous environmental factors, which has the following beneficial effects:
[0077] 1. The FCPIie model provided by the present invention combines the internal operation and external environmental characteristics of the line as input, and constructs an evaluation database A by acquiring and normalizing the historical fault records of the line and the external environmental characteristics related to the fault. The risk of line failure is predicted by combining the internal operation status of the line and the external environmental characteristics. The model input includes historical fault records and external environmental factors. By effectively combining multiple characteristics (such as line operation status, external environment, etc.), the model can be applied in different scenarios. Whether facing different types of line configurations or various sudden environmental changes, effective predictions can be made, meeting the application requirements of the model for a wide range of applications and multiple scenarios; and the construction of the evaluation database A does not depend on the topological structure of the line, nor does it require the input data to have a specific mathematical distribution, which can ensure that the method can be flexibly applied in different environments and conditions, avoiding the excessive reliance of traditional methods on prior knowledge.
[0078] Traditional fault prediction methods usually require complex system topology, equipment configuration and specific mathematical models, while the present invention adopts the factor-risk action model, which can effectively avoid these limitations and ensure flexible response in uncertain environments;
[0079] 2. The present invention also uses conditional correlation pattern identification (CCPI) to learn discrete features, and uses probabilistic fuzzy inference system (PFIS) to learn continuous features, thereby realizing parallel processing of multi-type heterogeneous input features, which not only improves the processing capability of the model, but also effectively distinguishes the impact of different features on fault risks, especially in the processing of multi-type heterogeneous data, showing strong adaptability and efficiency;
[0080] 3. The present invention also uses fuzzy support and conditional fuzzy support to evaluate common factors and rare factors respectively, and quantifies the risks of high-risk factors and rare high-risk factors through normalization processing, so that the system can take into account the ambiguity and uncertainty of different data when dealing with rare high-risk factors. These fuzzy methods help to accurately identify and process biased data, thereby ensuring the comprehensiveness and accuracy of risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0082] Figure 1 It is a flow chart of a method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors disclosed in the present invention;
[0083] Figure 2 is the probability distribution function of the average wind speed and average humidity disclosed in the embodiment of the present invention, Figure 2 (a) is the probability distribution function of the average wind speed, Figure 2 (b) is the probability distribution function of average humidity;
[0084] Figure 3 The input membership function of the average wind speed and average humidity disclosed in the embodiment of the present invention;
[0085] Figure 4 The output membership function of the average wind speed and average humidity disclosed in the embodiment of the present invention;
[0086] Figure 5 A schematic diagram of probabilistic fuzzy risk obtained after defuzzification of the aggregated risk area in the output membership function disclosed in the embodiment of the present invention;
[0087] Figure 6 A schematic diagram of the working principle of the FCPIie model disclosed in an embodiment of the present invention;
[0088] Figure 7 The embodiment of the present invention discloses the use of the FCPIie model to predict the fault risk of a certain transmission line and visualize the prediction results;
[0089] Figure 8 The ROC curve diagram of the performance analysis results of the FCPIie model, PFIS model and CCPI model disclosed in the embodiment of the present invention for prediction in three fault processing result scenarios, Figure 8(a) is the scenario where the fault is successfully handled. Figure 8 (b) In the scenario where the fault handling is barely successful, Figure 8 (c) is the scenario where the fault handling fails;
[0090] Fig. 9 The PR curve diagram of the performance analysis results of the FCPIie model, PFIS model and CCPI model disclosed in the embodiment of the present invention under three fault processing result scenarios is shown in FIG. Fig. 9 (a) is the scenario where the fault is successfully handled. Fig. 9 (b) In the scenario where the fault handling is barely successful, Fig. 9 (c) is the scenario where the fault handling fails;
[0091] Fig.10 KS curve diagram of the performance analysis results of the FCPIie model, PFIS model and CCPI model disclosed in the embodiment of the present invention under three fault processing result scenarios, Fig.10 (a) is the scenario where the fault is successfully handled. Fig.10 (b) In the scenario where the fault handling is barely successful, Fig.10 (c) is the scenario where fault handling fails. DETAILED DESCRIPTION
[0092] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0093] The present invention provides a method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors, such as Figure 1 As shown, the following steps are included:
[0094] S1. Obtain a set of historical line fault records, normalize the historical line fault records and corresponding characteristic factors, construct an input mapping space between the historical line fault records and the corresponding fault characteristics, and obtain an evaluation database A;
[0095] S2. For data bias-heterogeneous environment, complete the parallel learning of discrete features and continuous features in fault features to obtain common factors and rare factors, including: learning discrete features through conditional correlation pattern identification CCPI and distinguishing common factors and rare factors that cause corresponding faults; learning continuous features through probabilistic fuzzy inference system PFIS and distinguishing common factors and rare factors that cause corresponding faults;
[0096] S3. A fuzzy conditional correlation pattern recognition model FCCPI was established. The high-risk HR factors and rare high-risk RHR factors were extracted from common factors and rare factors respectively through fuzzy support and conditional fuzzy support. The fuzzy support and conditional fuzzy support were normalized and then integrated to obtain the support score. The risk level was quantified through the support score.
[0097] It should be noted that:
[0098] This embodiment selects the power transmission network in a certain area of central China as an example test system, including some renewable energy power sources such as hydropower plants, wind farms, photovoltaic power stations, etc. This embodiment collects fault records in the system in the past five years. In each record, discrete features such as fault causes (icing, lightning strikes, external forces, wildfires, etc.), terrain, wind direction and tower material, etc., and continuous features include average temperature, daily precipitation, wind speed, call height and ground resistance, etc.
[0099] To identify the factors strongly associated with failure risk, the FCCPI model was used to further evaluate the rare and common factors and screen out the HR and RHR factors.
[0100] In order to further implement the above technical solution, the specific contents of constructing the input mapping space in S1 include:
[0101] The collection of m historical fault records is denoted as D = {t 1 ,…,t i ,…,t m}, where i = 1,…,m,t i is the i-th historical fault record;
[0102] The feature set is denoted as F = {f 1 ,…,f j ,…,f n}, where j = 1,…,n, n is the number of features, f j represents the jth feature, each feature is composed of a set of factors, denoted as f j ={c 1,j ,…,c i,j ,…,c m,j}∈F,c i,j Represents feature f j A factor in
[0103] will be and When X occurs, the relevant pattern that Y must occur is expressed as X→Y, where X is the conditional factor set, which is a subset of the feature set F, and Y is the target factor set;
[0104] The target factor set is denoted as Y = {y 1 ,…,yi ,…,y m}, where y i is the fault handling result, indicating a fault record t i The result of whether the current fault is successfully handled is recorded in the fault handling result, which includes: successful fault handling Y(S), barely successful Y(P) or failed Y(F). Therefore, y i =Y(r)∈{Y(S),Y(P),Y(F)};
[0105] Then the evaluation database A is:
[0106]
[0107] In order to further implement the above technical solution, S2 uses conditional correlation pattern identification CCPI to learn discrete features and distinguish common factors and rare factors that cause corresponding faults. The specific contents include:
[0108] The scores are calculated according to the association rules and compared with the preset thresholds. Then, the conditional factor set X is decomposed to obtain common factor subsets and rare factor subsets corresponding to the discrete features.
[0109] It should be noted that:
[0110] The scores calculated in association rules include support, confidence, and lift.
[0111] In order to further implement the above technical solution, S2 uses the probabilistic fuzzy inference system PFIS to learn continuous features and distinguish the common factors and rare factors that cause the corresponding faults. The specific contents include:
[0112] The continuous features are learned through the probabilistic fuzzy inference system PFIS, and the probability distribution function PDF of each continuous feature is divided into rare degrees to obtain common factors and rare factors.
[0113] It should be noted that:
[0114] Compared with discrete values, continuous values exist in the form of numerical intervals. Direct division by absolute boundaries will lead to greater subjective uncertainty; FIS can be based on overlapping boundaries and membership, which can effectively deal with it. Among them, when constructing the input membership function, the input continuous value will be divided into several fuzzy sets. In most FIS, fuzzy sets are often determined according to the values of corresponding features. For example, the average temperature is defined as "cold" when it is between -5℃ and 0℃, and the daily precipitation is defined as "drought" when it is between 0 and 5 mm. These FIS based on feature fuzzy sets have relatively clear and accurate outputs, but when faced with different features or scenarios, functions with different values must be constructed, increasing the computational burden. At the same time, rare variables in each feature will also be ignored. To this end, the present invention relies on probabilistic fuzzy sets to construct PFIS, that is, the input membership function is calculated by the frequency of occurrence of each numerical segment, which is more flexible and adjustable when processing different scenarios or features.
[0115] In the present invention, the probability distribution function (PDF) of two continuous features is plotted respectively, and the equidistant boundary values are set to divide the PDF into four regions: rare (R), uncommon (U), possible (P), and frequent (F). The probability of the continuous feature factor in the divided region can be obtained. Figure 2 As shown, in this embodiment, the rare (R) and uncommon (U) obtained by the division are used as the common factors described in S2, and the possible (P) and frequent (F) are used as the rare factors in S2.
[0116] In order to further implement the above technical solutions, the specific contents of S3 include:
[0117] S31. Expand the correlation pattern X→Y to obtain the expanded correlation pattern:
[0118] X,P→Y,Q
[0119] Where P = {p 1,1 ,p 1,2 ,…,p i,j ,…,p m,n} and Q = {q 1 ,q 2 ,…,q i ,…,q m} are fuzzy sets corresponding to X and Y, respectively, where p i,j ,q i are the factors in the fuzzy set of P and Q respectively; the expanded correlation model means that if X is related to P, then Y is considered to be related to Q;<X,P> Represents related feature-fuzzy set pairs;
[0120] S32. Input common factors and their fuzzy sets into the fuzzy support S cIn the model, the corresponding fuzzy support is calculated, where the fuzzy support is greater than the fuzzy support voting threshold Th s The common factor is the HR factor, and the fuzzy support of the HR factor After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score;
[0121] Input the rare factors and their fuzzy sets into the conditional fuzzy support S r In the model, the corresponding conditional fuzzy support is calculated, where the conditional fuzzy support is greater than the fuzzy support voting threshold Th s The rare factors are RHR factors, and the fuzzy support of RHR factors is After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score;
[0122] The support score ranges from (0 to 1), where the closer it is to 1, the higher the risk level of failure.
[0123] In order to further implement the above technical solution, the fuzzy support model S in S32 c The specific contents include:
[0124] Fuzzy support model S c for:
[0125]
[0126] Where: B() represents the cardinality of fault records after preprocessing, Th s The voting threshold indicating support, t i (c i,j ) indicates t i Medium i,j The value of Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree of for Abbreviation of .
[0127] It should be noted that:
[0128] From the above formula, we can see that when the support number of an event record exceeds zero, the event satisfies<X,P> . Using t i (x i,j ) Calculate x i,j The value of is converted into membership through membership function. The resulting membership should be greater than the support threshold, then the low membership factors will be discarded and the scores of high risk factors will be integrated into the support score.
[0129] In order to further implement the above technical solution, the conditional fuzzy support model S in S32 r The specific contents include:
[0130] Conditional fuzzy support model S r for:
[0131]
[0132] In the formula, X r is a rare factor set, is the coefficient:
[0133]
[0134] in, Represents the fuzzy set p i,j The corresponding fuzzy weight of .
[0135] It should be noted that:
[0136] The calculation methods of input membership function and output membership function are:
[0137] First, create input membership functions to fuzzify the corresponding probabilities. Construct four input membership functions for the four regions in the input PDF. The boundary between two adjacent fuzzy sets is obtained based on the probability of each region in the PDF, and the overlapping area of the boundary is determined through historical statistics. The input membership functions of the features "average wind speed" and "average humidity" are referenced. Figure 3 shown.
[0138] Secondly, a Mamdani-type hierarchical PFIS is constructed to reduce the computational complexity. Based on the fault classification regulations and historical statistical data, faults are divided into four risk levels: low (L), medium (M), high (H), and very high (E), and four fuzzy weights are assigned to them, namely 0.11, 0.22, 0.34, and 1. The fuzzy rules are shown in Table 1.
[0139] Table 1 fuzzy rule
[0140]
[0141]
[0142] After that, the output membership function is determined by four triangular functions. For the four fuzzy sets in Table 1: low (L), medium (M), high (H), and very high (E), the data proportion of each category in the four risk levels is statistically analyzed. Output membership function reference Figure 4 shown.
[0143] Finally, the aggregated risk area is defuzzified in the output membership function to obtain the final probabilistic fuzzy risk (the vertical dashed line in the figure), i.e., the fuzzy membership. Figure 5 shown. Represents the fuzzy set p i,j The corresponding fuzzy weights can be obtained from Table 1.
[0144] In order to further implement the above technical solution, S3 also includes S33:
[0145] Verify the high-risk HR factors and rare high-risk RHR factors screened out in S32:
[0146] Let the characteristic fuzzy set pair<X,P> and<Y,Q> Another characteristic fuzzy set<Z,L> A subset of And Z = X ∪ Y, And L=P∪Q, then Z={z 1 ,z 2 ,…,z i ,…,z m}, L = {l 1 ,l 2 ,…,l i ,…,l m} is a fuzzy set associated with Z, where z i ,l i are the i-th related factors in Z and L respectively;
[0147] For the HR factor, the three cases of fault handling results Y(r) are evaluated by the fuzzy certainty evaluation model Ce c and fuzzy correlation evaluation model Co c Find two characteristic fuzzy set pairs<<X,P> ,<Y,Q> >The fuzzy certainty and fuzzy relevance of the fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the fuzzy correlation is greater than the correlation evaluation threshold Th co Then the input factor is determined to be the HR factor;
[0148] At the same time, for the RHR factor, the three cases of fault handling results Y(r) are respectively analyzed by the conditional fuzzy certainty model Ce r and fuzzy correlation evaluation model Co r Find two characteristic fuzzy set pairs<<X,P> ,<Y,Q> >The conditional fuzzy certainty and conditional fuzzy relevance, the conditional fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the conditional fuzzy relevance is greater than the relevance evaluation threshold Th co Then the input factor is determined to be the RHR factor.
[0149] It should be noted that:
[0150] By evaluating the model based on certainty and association, frequent models can be discovered from all possible patterns. When both item sets and patterns are confirmed to be strongly correlated, valuable factor-risk patterns can be obtained.
[0151] Relevant theory can prove that if a large item set has been verified, then all its subsets are also large item sets.
[0152] In order to further implement the above technical solution, the specific contents of S33 include:
[0153] Fuzzy Certainty Evaluation Model Ce c for:
[0154]
[0155] Where: Th ce represents the voting threshold of the certainty measure, t i (z i ) is z i In t i The value in or Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree, M li or M li [t i (z i )] represents the fuzzified membership of the input membership function to the fuzzy set li;
[0156] Conditional Fuzzy Certainty Model Ce r for:
[0157]
[0158] Where: X r and Z r denote the rare factor set and its associated fuzzy set, respectively. and The two coefficients are:
[0159]
[0160] in, and Respectively represent the fuzzy set p i,j and fuzzy set l i The corresponding fuzzy weights; fuzzy correlation evaluation model Co c for:
[0161]
[0162] Where: Covc (X,Y) is the covariance, Var c (X) and Var c (Y) is the variance;
[0163] Cov c (X,Y)=E c [<Z,L> ]-E c [<X,P> ]×E c [<Y,Q> ]
[0164] Among them, E[…] represents the expectation:
[0165]
[0166]
[0167] and The two coefficients are:
[0168]
[0169] Var c (X) and Var c (Y) are:
[0170] Var c (X) = E c [<X,P> 2 ]-E c [<X,P> ] 2
[0171] Var c (Y) = E c [<Y,Q> 2 ]-E c [<Y,Q> ] 2
[0172]
[0173] Conditional fuzzy correlation evaluation model Co r for:
[0174]
[0175] In the formula, Cov r (X,Y) is the covariance, Var r (X) and Var r (Y) is the variance;
[0176] Cov r (X,Y)=E r [<Z,L> ]-Er [<X,P> ]×E r [<Y,Q> ]
[0177] Among them, E[…] represents the expectation:
[0178]
[0179] Var r (X) and Var r (Y) are:
[0180] Var r (X) = E r [<X,P> 2 ]-E r [<X,P> ] 2
[0181] Var r (Y) = E r [<Y,Q> 2 ]-E r [<Y,Q> ] 2
[0182]
[0183] The present invention will be further described below in conjunction with specific models and verification tools:
[0184] The model for implementing the method of the present invention is denoted as FCPIie. First, the state features and factors collected in the fault records are preprocessed as input, and the input features are divided into discrete features and continuous features; then the two types of factors are further divided into common factors and rare factors through the CCPI and PFIS models; then a qualitative analysis is performed to mine risk factors, and the common factors and rare factors are evaluated through the two types of fuzzy importance evaluations in the FCCPI model, and the HR factors and RHR factors are respectively screened out, and the action mode of the fault factors is clarified, and then the prediction of the future fault risk distribution is realized. Model comprehensive process, refer to Figure 6 shown.
[0185] Results test:
[0186] In order to fully verify the performance of the constructed model, this embodiment uses ROC, PR and KS curves to verify the three types of fault processing results. Among them, the larger the value of the area under the ROC and PR curves, or the larger the maximum value of the KS curve, the stronger the prediction ability of the model. In addition, this embodiment compares and analyzes the improvement effect of the integrated FCPIie model with the two initial models PFIS and CCPI. Finally, this embodiment uses the FCPIie model to predict and visualize the fault risk of a certain transmission line. The prediction results are as follows: Figure 7 As shown, green indicates a low risk of failure and red indicates a high risk of failure.
[0187] exist Figure 8 In the above figure, the inspection unnecessary rate (IUR) on the horizontal axis represents the proportion of samples that are incorrectly identified as positive among all samples that are actually negative, and the fault detection rate (FDR) on the vertical axis represents the proportion of samples that are correctly identified as positive among all samples that are actually positive. Figure 8 It can be seen that the AUROC score of FCPIie is 4.9% and 17.1% higher than that of PFIS and CCPI on average, respectively, indicating that the prediction performance of the FCPIie model is better for all fault handling result scenarios;
[0188] exist Fig. 9 In the figure, the horizontal axis is the Fault Detection Rate (FDR), which indicates the proportion of samples that are correctly identified as positive among all samples that are actually positive. The vertical axis is the Precise Predictive Value (PPV), which indicates the proportion of samples that are actually positive among all samples that are predicted to be positive. Fig. 9 It can be seen that by comparing the three sets of PR curves generated, the AUPR score of FCPIie is 5.2% and 14.9% higher than that of PFIS and CCPI on average, demonstrating its performance in a negative sample-dominated environment;
[0189] exist Fig.10 In the figure, the horizontal axis is the evaluation threshold, which represents the critical value of the probability of occurrence and non-occurrence, and the vertical axis is the cumulative probability, which represents the difference between FDR and IUR. The vertical line in each curve represents the maximum difference between FDR and IUR at that point, and the horizontal axis corresponding to that point is the threshold of the model, which is given by Fig.10 From the comparison of KS curves, it can be seen that the FCPIie model has better fault judgment performance than PFIS and CCPI.
[0190] Through the analysis of ROC and PR curve results, it can be seen that the fault prediction effect of successful processing results is the best, while the barely successful results are relatively the lowest. This confirms the decisive role of the scale and granularity of the data set in the prediction performance. But on the other hand, since the number of events with barely successful results in the data set is smaller, it represents a greater number of negative examples. Therefore, from the results based on AUPR evaluation (FCPIie is 5.82% better than PFIS and 18.09% better than CCPI), it can be seen that the improvement in the prediction effect of barely successful results is more significant than that of successful (FCPIie is 2.56 better than PFIS and 7.61% better than CCPI) and failed results (FCPIie is 5.24% better than PFIS and 13.42% better than CCPI), indicating that when there are more negative examples in the data set, the PR curve has a higher degree of distinction.
[0191] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors, characterized in that: The following steps are involved: S1. Obtain a set of historical line fault records, normalize the historical line fault records and corresponding characteristic factors, construct an input mapping space between the historical line fault records and the corresponding fault characteristics, and obtain an evaluation database A; S2. For data bias-heterogeneous environment, complete the parallel learning of discrete features and continuous features in fault features to obtain common factors and rare factors, including: learning discrete features through conditional correlation pattern identification CCPI and distinguishing common factors and rare factors that cause corresponding faults; The continuous features are learned through the probabilistic fuzzy inference system PFIS, and the common factors and rare factors that cause the corresponding faults are distinguished; S3. A fuzzy conditional correlation pattern recognition model FCCPI was established. The high-risk HR factors and rare high-risk RHR factors were extracted from common factors and rare factors respectively through fuzzy support and conditional fuzzy support. The fuzzy support and conditional fuzzy support were normalized and then integrated to obtain the support score. The risk level was quantified through the support score.
2. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 1 is characterized in that: The specific contents of constructing the input mapping space in S1 include: The collection of m historical fault records is recorded as D = {t1, L, t i ,L,t m }, where i = 1, L, m, t i is the i-th historical fault record; The feature set is denoted as F = {f1, L, f j ,L,f n }, where j = 1, L, n, n is the number of features, f j represents the jth feature, each feature is composed of a set of factors, denoted as f j ={c 1,j ,L,c i,j ,L,c m,j }∈F,c i,j Represents feature f j A factor in will be and When X occurs, the relevant pattern that Y must occur is expressed as X→Y, where X is the conditional factor set, which is a subset of the feature set F, and Y is the target factor set; The target factor set is denoted as Y = {y1,L,y i ,L,y m }, where y i is the fault handling result, indicating a fault record t i The result of whether the current fault is successfully handled is recorded in the fault handling result, which includes: successful fault handling Y(S), barely successful Y(P) or failed Y(F). Therefore, y i =Y(r)∈{Y(S),Y(P),Y(F)}; Then the evaluation database A is:
3. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 2 is characterized in that: In S2, conditional correlation pattern recognition (CCPI) is used to learn discrete features and distinguish between common factors and rare factors that cause corresponding failures. The specific contents include: The scores are calculated according to the association rules and compared with the preset thresholds. Then, the conditional factor set X is decomposed to obtain common factor subsets and rare factor subsets corresponding to the discrete features.
4. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 2, characterized in that: In S2, the probabilistic fuzzy inference system PFIS is used to learn continuous features and distinguish between common factors and rare factors that cause corresponding failures. The specific contents include: The continuous features are learned through the probabilistic fuzzy inference system PFIS, and the probability distribution function PDF of each continuous feature is divided into rare degrees to obtain common factors and rare factors.
5. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 2, characterized in that: The specific contents of S3 include: S31. Expand the correlation pattern X→Y to obtain the expanded correlation pattern: X,P→Y,Q Where P = {p 1,1 ,p 1,2 ,L,p i,j ,L,p m,n } and Q = {q1,q2,L,q i ,L,q m } are fuzzy sets corresponding to X and Y, among which p i,j ,q i are the factors in the fuzzy set of P and Q respectively; the expanded correlation model means that if X is related to P, then Y is considered to be related to Q;<X,P> Represents related feature-fuzzy set pairs; S32. Input common factors and their fuzzy sets into the fuzzy support S c In the model, the corresponding fuzzy support is calculated, where the fuzzy support is greater than the fuzzy support voting threshold Th s The common factor is the HR factor, and the fuzzy support of the HR factor After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score; Input rare factors and their fuzzy sets into the conditional fuzzy support S r In the model, the corresponding conditional fuzzy support is calculated, where the conditional fuzzy support is greater than the fuzzy support voting threshold Th s The rare factors are RHR factors, and the fuzzy support of RHR factors is After normalization, the support score is obtained, and the risk level is quantified according to the size of the support score; The support score ranges from (0 to 1), where the closer it is to 1, the higher the risk level of failure.
6. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 5, characterized in that: Fuzzy support model S in S32 c The specific contents include: Fuzzy support model S c for: Where: B( ) represents the cardinality of fault records after preprocessing, Th s The voting threshold indicating support, t i (c i,j ) indicates t i Medium i,j The value of Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree of for Abbreviation of .
7. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 5, characterized in that: Conditional fuzzy support model S in S32 r The specific contents include: Conditional fuzzy support model S r for: Where, X r is a rare factor set, is the coefficient: in, Represents the fuzzy set p i,j The corresponding fuzzy weight of .
8. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 6, characterized in that: S3 also includes S33: Verify the high-risk HR factors and rare high-risk RHR factors screened out in S32: Let the characteristic fuzzy set pair<X,P> and<Y,Q> Another characteristic fuzzy set<Z,L> A subset of And Z = X ∪ Y, And L=P∪Q, then Z={z1,z2,L,z i ,L,z m }, L={l1,l2,L,l i ,L,l m } is a fuzzy set associated with Z, where z i ,l i are the i-th related factors in Z and L respectively; For the HR factor, the three cases of fault handling results Y(r) are evaluated by the fuzzy certainty evaluation model Ce c and fuzzy correlation evaluation model Co c Find two characteristic fuzzy set pairs<<X,P> ,<Y,Q> >The fuzzy certainty and fuzzy relevance of the fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the fuzzy correlation is greater than the correlation evaluation threshold Th co Then the input factor is determined to be the HR factor; At the same time, for the RHR factor, the three cases of fault handling results Y(r) are respectively analyzed by the conditional fuzzy certainty model Ce r and fuzzy correlation evaluation model Co r Find two characteristic fuzzy set pairs X,P>,<Y,Q> >The conditional fuzzy certainty and conditional fuzzy relevance, the conditional fuzzy certainty is greater than the voting threshold Th of the certainty measure ce And the conditional fuzzy relevance is greater than the relevance evaluation threshold Th co Then the input factor is determined to be the RHR factor.
9. The method for predicting the spatiotemporal distribution of line fault risk under biased-heterogeneous environmental factors according to claim 6, characterized in that: The specific contents of S33 include: Fuzzy Certainty Evaluation Model Ce c for: Where: Th ce represents the voting threshold of the certainty measure, t i (z i ) is z i In t i The value in or Denotes the input membership function to the fuzzy set p i,j The fuzzy membership degree, M li or M li [t i (z i )] represents the fuzzified membership of the input membership function to the fuzzy set li; Conditional Fuzzy Certainty Model Ce r for: Where: X r and Z r denote the rare factor set and its associated fuzzy set, respectively. and The two coefficients are: in, and Respectively represent the fuzzy set p i,j and fuzzy set l i The corresponding fuzzy weight of Fuzzy correlation evaluation model Co c for: Where: Cov c (X,Y) is the covariance, Var c (X) and Var c (Y) is the variance; Number c (X,Y)=E c [<Z,L> ]-E c [<X,P> ]×E c [<Y,Q> ] Where E[L] represents the expectation: and The two coefficients are: Var c (X) and Var c (Y) are: Var c (X)=And c [<X,P> 2 ]-AND c [<X,P> ] 2 There is c (Y)=Y c [<Y,Q> 2 ]-TO c [<Y,Q> ] 2 Conditional fuzzy correlation evaluation model Co r for: In the formula, Cov r (X,Y) is the covariance, Var r (X) and Var r (Y) is the variance; Cov r (X,Y)=E r [<Z,L> ]-E r [<X,P> ]×E r [<Y,Q> ]Where E[L] represents the expectation: Var r (X) and Var r (Y) are: Var r (X)=And r [<X,P> 2 ]-AND r [<X,P> ] 2 There is r (Y)=Y r [<Y,Q> 2 ]-TO r [<Y,Q> ] 2
Citation Information
Cited By
SF6 high-voltage circuit breaker operation refusal risk prediction method, device, equipment, medium and product
CN121350705A
Method, device, equipment, medium and product for predicting risk of SF6 high-voltage circuit breaker refusal
CN121350705B