Power grid line fault knowledge base construction method and power grid line fault prediction method
By building a dynamic database and using a dual-mode dynamic tree DDT-tree for incremental mining, combined with the trained fault prediction network, the problem of insufficient construction of the power grid fault knowledge base in the existing technology is solved, and the accuracy of grid line fault prediction is significantly improved.
Patent Information
- Application Number
- CN202510333392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively build a grid fault knowledge base under complex meteorological conditions, resulting in inaccurate prediction results of grid line faults.
By collecting power grid operation data and meteorological data, building a dynamic database, designing association rule evaluation indicators, using the dual-mode dynamic tree DDT-tree for incremental mining, updating association rules, building a grid line fault knowledge base, and using the trained fault prediction network for prediction.
It significantly improves the accuracy of grid line fault prediction results, enhances the reliability and safety of grid operation, and can effectively deal with dynamically changing grid fault modes.
Smart Images

Figure CN120179652A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grid fault detection, and particularly relates to a method for constructing a knowledge base for power grid line faults and a method for predicting power grid line faults. Background Art
[0002] During the process of power transmission, the safety performance of the power grid is very important. Whether the power system can operate stably will directly affect the production and life in related areas. However, meteorological factors will significantly affect the operating environment of power grid lines or equipment. Especially under extreme weather conditions such as heavy rain and thunderstorms, power grid lines are very likely to be severely affected, resulting in power system accidents and threatening the safe operation of the power grid; it will also cause physical damage to power grid lines, trigger cascading failures, and lead to large-scale power outages. Therefore, faults in power system lines should be detected in a timely manner and eliminated as soon as possible.
[0003] In the prior art, the publication number CN 117436351 A discloses a method and system for predicting power grid equipment faults based on a knowledge graph under complex meteorology. The method includes: using Transformer and knowledge graph technologies as the basic means for prediction, and during the training process, synchronously training a Transformer prediction model and a BP neural network classifier based on the knowledge graph prediction model. With the unique data structure of the knowledge graph, related nodes and node relationships are found, and further relevant historical case data is found, and then adjusted to optimize the sample quality. However, the knowledge graph of power grid faults can be further improved, and the prediction results of power grid faults still need to be further enhanced. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a method for constructing a knowledge base for power grid line faults and a method for predicting power grid line faults.
[0005] In the first aspect, the present invention proposes a method for constructing a knowledge base for power grid line faults, and the method includes:
[0006] S1: Collect power grid operation data and meteorological data, and construct a dynamic database D for power grid line faults;
[0007] S2: Based on the dynamic database D, design an evaluation index for the association rules of power grid line faults;
[0008] S3: Construct a dual-mode dynamic tree DDT-tree of the dynamic database D, update the dynamic database D to obtain incremental data, and update the dual-mode dynamic tree DDT-tree;
[0009] S4: According to the updated dual-mode dynamic tree DDT-tree, based on the association rule evaluation metrics, update the association rules, and construct a fault knowledge base for the power grid lines according to the updated association rules.
[0010] In a second aspect, the present invention proposes a power grid line fault prediction method, which is based on the power grid line fault knowledge base and includes:
[0011] Obtain power grid operation data and meteorological prediction data, preprocess them to obtain the input data to be measured;
[0012] After inputting the input data to be measured into the trained fault prediction network, obtain the prediction result;
[0013] Combine the prediction result with the power grid line fault knowledge base for matching judgment to obtain the prediction result of the power grid line fault.
[0014] In a third aspect, the present invention proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the power grid line fault knowledge base construction method proposed in the first aspect of the present invention or the power grid line fault prediction method proposed in the second aspect of the present invention.
[0015] In a fourth aspect, the present invention proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the power grid line fault knowledge base construction method proposed in the first aspect of the present invention or the power grid line fault prediction method proposed in the second aspect of the present invention.
[0016] The beneficial effects of the present invention: The present invention analyzes the historical data of power grid faults and historical meteorological data, establishes a dynamic database of power grid faults, designs evaluation metrics for association rules, and can effectively evaluate the association rules; based on the evaluation metrics of association rules, adopts a dynamic association rule mining method (i.e., the DDT-tree incremental mining algorithm), which can not only mine the association rules of frequent variables but also mine the association rules of rare variables to adapt to the dynamic changes of the dynamic database and realize the efficient update and maintenance of the association rule set. Furthermore, construct and update the fault knowledge base of the power grid lines according to the dynamically mined association rules. Use the trained fault prediction network and combine it with the fault knowledge base to obtain the prediction result of the power grid line fault. The present invention can significantly improve the accuracy of the power grid line fault prediction result, facilitate power grid fault early warning and maintenance, and enhance the reliability and security of power grid operation. Description of the Drawings
[0017] Figure 1 It is a flowchart of the steps of an embodiment of the present invention;
[0018] Figure 2 It is the overall flowchart of the embodiment of the present invention;
[0019] Figure 3 It is the pseudo-code for constructing and updating the DDT-tree algorithm in the embodiment of the present invention;
[0020] Figure 4 It is the pseudo-code for the DDT-tree incremental mining algorithm in the embodiment of the present invention;
[0021] Figure 5 It is the schematic diagram of the association rule judgment process of the fault knowledge base in the embodiment of the present invention;
[0022] Figure 6 It is the schematic diagram of the structure of the fault knowledge base in the embodiment of the present invention;
[0023] Figure 7 It is the schematic diagram of the processing process of the power grid line fault prediction method in the embodiment of the present invention;
[0024] Figure 8 It is the schematic diagram of the network structure of the fault prediction network in the embodiment of the present invention. Detailed implementation manners
[0025] The terms "first", "second", "third", "fourth", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application.
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] Through data research and analysis, the important reasons leading to power grid line faults are natural factors, such as lightning, ice disasters, wind disasters, geological disasters, wildfires, etc.
[0028] The embodiment of the present invention proposes a method for constructing a power grid line fault knowledge base. Referring to Figure 1 、 2 shown, the method includes:
[0029] S1: Collect power grid operation data and meteorological data, and construct a dynamic database D for power grid line faults.
[0030] Analyze the impact of meteorological factors on power grid line faults from the historical operation data of the power grid and the historical meteorological data, and obtain the important meteorological data for power grid line faults, as shown in Table 1.
[0031] Table 1 Classification, attributes and elements included in power grid line faults
[0032]
[0033]
[0034] Since the impacts of different weathers on power grid equipment are different, and even for the same meteorological factor, the power grid equipment or lines affected by its different magnitudes and the degrees of influence will also vary greatly. Therefore, in the embodiments of the present invention, the meteorological data is refined according to the national classification standard and the influence degree of different meteorological factors on power grid equipment or lines, and then the weather in the fault repair report is replaced based on the relevant knowledge of the database, so as to facilitate the subsequent association rule analysis. Based on the above information, the input environmental characteristics selected in the embodiments of the present invention and the elements included are shown in Table 2.
[0035] Table 2 Classification of environmental characteristics and elements included therein
[0036]
[0037]
[0038] Combining the above information, set it as the data representing in database D, that is, the fault records that occurred.
[0039] In database D, let F = {f1, f2, f3} be the set of all fault characteristics, where f1 represents the voltage level, f2 represents the fault phase, f3 represents the fault time, and each fault characteristic f j has its corresponding elements {f j,1 , f j,2 , … f j,n}, j represents the corresponding fault characteristic, for example, j = 1 represents the voltage level, j = 2 represents the fault phase, and n represents the nth element corresponding to the corresponding fault characteristic; E = {e1, e2, e3, e4, e5} is the environmental characteristic, and each environmental characteristic e j has its corresponding elements {e j,1 , e j,2 , … e j,n}, such as the occurrence season, terrain, wind direction, cloud cover, weather, meteorological elements, etc. The meteorological elements include precipitation, snowfall, temperature, relative humidity, wind speed, etc. R is the set of fault causes when power grid circuit faults occur, {r j,1 , …, r j,n} are the elements contained in the corresponding fault cause R, such as icing, lightning strike, flashover, wind disaster, wildfire, geological disaster, etc. Y represents the processing result of the power grid circuit fault, {y j,1 ,…,y j,n} are the elements contained in the corresponding processing result Y, such as success or failure. Let I = {w1, w2, …} be a set containing all input feature variables.
[0040] According to the above content, the corresponding data space processing matrix T can be obtained:
[0041]
[0042] In the formula, each row starting from the second represents the record of a fault, w i,j The numerical value of the characteristic element corresponding to each fault record, y i,15 is the corresponding processing result, i ∈ [1, 2, …, m], m represents the number of fault record numbers, and j ∈ [1, 2, …, 15].
[0043] Combining the space processing matrix T and the number of the fault record, a dynamic database D of the power grid line fault is constructed, which is expressed as:
[0044]
[0045]
[0046] In the formula, t i represents the number of the i-th row fault record, i ∈ [1, 2, … m], m represents the number of fault record numbers, f1 represents the voltage level when the power grid circuit fault occurs, f2 represents the fault phase when the power grid circuit fault occurs, f3 represents the fault time when the power grid circuit fault occurs, e1 represents the value of the first environmental characteristic element when the power grid circuit fault occurs, such as the season of occurrence, e 10 represents the value of the tenth environmental characteristic element when the power grid circuit fault occurs, such as the wind speed, R represents the set of fault causes when the power grid circuit fault occurs, Y represents the set of processing results of the power grid circuit fault, w i,j represents the numerical value of the characteristic element corresponding to each fault record, y i,15 represents the processing result corresponding to a specific power grid circuit fault, i ∈ [1, 2, …, m], m represents the number of fault record numbers, and j ∈ [1, 2, …, 15].
[0047] When traditional association rule mining algorithms are used for fault analysis in transmission line systems, they fail to fully consider the uneven distribution of faults in different time periods. For example, icing faults are more frequent in winter and relatively rare in spring and autumn. If icing faults are relatively common in a certain system, the annual fault events of this system will mainly concentrate in winter. However, traditional algorithms still use a fixed importance threshold to equally evaluate and analyze faults in winter and spring / autumn. Since the occurrence frequency of faults in spring and autumn is relatively low, the importance scores of their corresponding environmental states are often lower than the annual set threshold, resulting in the fault patterns in these rare seasons being easily ignored or directly filtered out. Although the occurrence frequency of faults in rare time periods is low, these faults may also cause the operation interruption of the transmission line system and result in serious losses. Therefore, these rare time periods need to be considered during the analysis.
[0048] S2: Based on the dynamic database D, design the association rule evaluation index for power grid line faults.
[0049] The embodiment of the present invention designs a method for setting the threshold of the importance diagnosis standard, which can set a more reasonable threshold according to the distribution of faults in different time periods in the fault database. Specifically, the embodiment of the present invention selects a quarter as the reference time period, and the same conditional importance diagnosis standard value will be used for the power grid line faults occurring within the same quarter.
[0050] Define the general association rule of power grid line faults as:
[0051] X→Y
[0052] Wherein, X represents the cause of the power grid line fault, that is, the item set including all element values in the fault characteristics and environmental characteristics, and Y represents the processing result of the power grid line fault, that is, the reclosing processing result.
[0053] The embodiment of the present invention uses five evaluation indexes, namely support, confidence, lift, conviction and leverage, to evaluate the association rules of power grid line faults. Through these five evaluation indexes, the quality and effectiveness of the mined association rules can be evaluated from different angles. Support and confidence are the basic indexes for measuring the universality and reliability of rules, while lift, conviction and leverage further reveal the association strength and statistical significance of rules. These evaluation methods provide a scientific basis for association rule mining and help researchers screen out rules with practical value.
[0054] (1) The support of the general association rule X→Y is an important index for measuring the occurrence frequency of an item set or rule in the dataset. In association rule mining, support is often used to screen frequent item sets. By setting the minimum support threshold, those item sets with low occurrence frequency can be eliminated, and the frequent item sets are retained for subsequent association rule generation.
[0055] Support is an important indicator for measuring the universality of rules, used to calculate the proportion of transactions containing all items in the rule to the total number of transactions. The higher the support, the more representative the rule is in the dataset.
[0056] The calculation formula for Support is:
[0057]
[0058] In the formula, Support(X→Y) represents the support of the association rule X→Y, |T(X∩Y)| represents the number of transactions that contain both X and Y in the data set, and |T| represents the total number of transactions.
[0059] (2) Confidence of the general association rule X→Y
[0060] Confidence is an important indicator for measuring the reliability of association rules, used to calculate the proportion of transactions that contain the consequent Y among the transactions that contain the antecedent X. The higher the confidence, the stronger the association between X and Y.
[0061] The calculation formula for Confidence is:
[0062]
[0063] In the formula, Confidence(X→Y) represents the confidence of the association rule X→Y, and |T(X)| represents the number of transactions that contain X.
[0064] (3) The Lift of the association rule X→Y is used to measure the association strength between set X and set Y, representing the ratio of the probability that X and Y appear simultaneously to the probability that they appear independently.
[0065] The calculation formula for Lift is:
[0066]
[0067] In the formula, Confidence(X→Y) represents the lift of the association rule X→Y, Support(X→Y) represents the support of the association rule X→Y, Support(X∩Y) represents the support of the data set that contains both X and Y, Support(X) represents the support of X in the data set, and Support(Y) represents the support of Y in the data set.
[0068] When Lift > 1, it indicates a positive correlation between X and Y, that is, the occurrence of X will increase the likelihood of the occurrence of Y; when Lift = 1, it indicates that X and Y are independent, that is, the occurrence of X has no association with the occurrence of Y; when Lift < 1, it indicates a negative correlation between X and Y, that is, the occurrence of X will reduce the likelihood of the occurrence of Y.
[0069] (4) The conviction of the association rule X→Y is used to measure the strength of the association rule, indicating the degree of association between X and Y in the case where X appears but Y does not appear.
[0070] The calculation formula for conviction is:
[0071]
[0072] In the formula, Convition(X→Y) represents the conviction of the association rule X→Y, Support(Y) represents the support of Y in the data set, and Confidence(X→Y) represents the confidence of the association rule X→Y.
[0073] The greater the conviction, the stronger the association between X and Y. When Conviction = 1, it indicates that X and Y are independent.
[0074] (5) The leverage of the association rule X→Y is used to measure the difference between the probability of X and Y occurring simultaneously and the probability of them occurring independently.
[0075] The calculation formula for leverage is:
[0076] Leverage(X→Y) = Support(X∩Y) - Support(X)·Support(Y)
[0077] In the formula, Leverage(X→Y) represents the leverage of the association rule X→Y.
[0078] The greater the leverage of Leverage(X→Y), the stronger the association between X and Y. Leverage(X→Y) = 0 indicates that X and Y are independent.
[0079] Based on the dynamic database D, the first-time dynamic adjustment parameter α is set according to different seasons to adjust the support of the general association rule X→Y. Based on the dynamic database D, the second-time dynamic adjustment parameter β is set according to different seasons to adjust the lift, conviction, and leverage of the general association rule X→Y.
[0080] The calculation formula for the first-time dynamic adjustment parameter α is:
[0081]
[0082] Among them, t i represents the i-th fault record in the dynamic database D, t i ∈D y (i, 1), i = 2, 3, ..., (m + 1) represents a row in the dynamic database D, D(i, 1) represents the i-th fault record in the dynamic database D, |…| represents the number of fault records in the dynamic database D that simultaneously satisfy all the included conditions, w i,4 represents the season corresponding to the i-th fault record, w i,4 = h s represents the fault record with the season being h s h s ∈ {C, X, Q, D} represents one of the four seasons: spring, summer, autumn, and winter, represents the season with the highest fault occurrence frequency, e1 = {w 1,1 , w i,1 , … w m,1} represents the set of seasons where m fault data are located, Y = {y 1,15 , y i,15 , … y m,15} represents the set of processing results of m fault data, h y ∈ {Y, N}, Y = {y 1,15 , y i,15 , … y m,15} represents one of the results of whether the reclosing is successful, y i,15 represents the processing result corresponding to the i-th fault record, y i,15 = h y represents the fault record with the processing result being h s .
[0083] The calculation formula for the second-time dynamic adjustment parameter β is as follows:
[0084]
[0085] In the formula, t i represents the i-th fault record in the dynamic database D, t i ∈D y (i, 1), i = 2, 3, …, (m + 1) represents a row in the dynamic database D, |…| represents the number of fault records in the dynamic database D that simultaneously satisfy all the included conditions, w i,4 represents the season corresponding to the i-th fault record, w i,4 = h s represents the fault record with the season being h s h s∈{C, X, Q, D} represents one of the four seasons: spring, summer, autumn, and winter. represents the season with the highest failure occurrence frequency, y i,15 represents the processing result corresponding to the i-th failure record, y i,15 = h y represents the failure record with the processing result h y Y = {y 1,15 , y i,15 , … y m,15} represents the set of m failure data processing results, h y ∈{Y, N} represents one of the results of whether the reclosing is successful.
[0086] Based on the five evaluation indicators of support, confidence, lift, conviction, and leverage of the general association rule X→Y, more reasonable thresholds are set respectively according to the distribution of power grid line failures in each season. The thresholds of various evaluation indicators of the general association rule X→Y are adjusted by using the dynamic adjustment parameters α and β, and their expressions are as follows:
[0087]
[0088] In the formula, represents the support threshold reset according to the initial support threshold and the first dynamic adjustment parameter α. minsupp represents the initial support threshold of the general association rule X→Y, h s ∈{C, X, Q, D} represents one of the four seasons: spring, summer, autumn, and winter.
[0089]
[0090] In the formula, represents the initial confidence threshold, minconf represents the initial confidence threshold of the association rule X→Y, h s ∈{C, X, Q, D} represents one of the four seasons: spring, summer, autumn, and winter, h y ∈{Y, N} represents one of the results of whether the reclosing is successful.
[0091]
[0092] In the formula, represents the lift threshold reset according to the initial lift threshold and the second time dynamic adjustment parameter β. minlift represents the initial lift threshold of the general association rule X→Y, h s ∈{C, X, Q, D} represents one of the four seasons: spring, summer, autumn, and winter, h y ∈{Y, N} represents one of the results of whether the reclosing is successful, and β represents the second time dynamic adjustment parameter.
[0093]
[0094] wherein represents the confidence threshold reset according to the initial confidence threshold and the second-time dynamic adjustment parameter β, minconv represents the initial confidence threshold of the general association rule X→Y, β represents the second-time dynamic adjustment parameter, and h s ∈{C, X, Q, D} represents one of the seasons of spring, summer, autumn, and winter, and h y ∈{Y, N} represents one of the results of whether the reclosing is successful.
[0095]
[0096] wherein represents the leverage threshold reset according to the initial leverage threshold and the second-time dynamic adjustment parameter β, minleve represents the initial leverage threshold of the general association rule X→Y, β represents the second-time dynamic adjustment parameter, and h s ∈{C, X, Q, D} represents one of the seasons of spring, summer, autumn, and winter, and h y ∈{Y, N} represents one of the results of whether the reclosing is successful.
[0097] The transmission line system failure caused by rare fault reasons, rare fault characteristics, or rare environmental characteristics may also lead to serious losses. Therefore, it is necessary to further mine high-importance low-frequency variables from the rare variable set of the dynamic database D.
[0098] Based on the calculation formulas of support, confidence, lift, conviction, and leverage, for the rare items containing fault reasons, fault characteristics, or environmental characteristics in the dynamic database D.
[0099] (1) The rare item association rule X r →Y
[0100] The rare item association rule X r →Y's support calculation formula is:
[0101]
[0102] wherein represents the support of the dynamic database D including the rare variable h r and represents the support of the input variable with the processing result of h y when the rare variable is included, t i represents the i-th fault record in the dynamic database D, and t i ∈Dy (i, 1), where i = 2, 3, ..., (m + 1) represents a row in the dynamic database D, and D(i, 1) represents the corresponding case in the dynamic database D, that is, the data in the i-th row corresponding to the first column in the dynamic database D. |…| represents the number of fault records in the dynamic database D that simultaneously satisfy all the included conditions, w i,j ∈X r indicates that the input variable belongs to the rare variable, h r is X r is one of the rare variable sets in X, w i,j = h r is the fault record where the input variable is a rare variable, X r represents the rare variable set of the fault cause, fault feature, and environmental feature sets in the dynamic database D. r represents the rare element, specifically any one of all the rare elements in the fault cause, fault feature, and environmental feature sets, represents the empty set.
[0103] The confidence calculation formula for the rare item association rule X r →Y is:
[0104]
[0105] In the formula, represents the confidence of the rare item association rule X r →Y, represents the support degree that both the fault feature and environmental feature in the dynamic database D contain the corresponding rare element h r and the processing result h y of, represents the support degree that the dynamic database D includes the rare variable h r of, h y represents one of the reclosing situations, y i,15 = h y represents the fault record where the fault result is h y e1 = {w 1,1 , w i,1 , … w m,1} represents the set of quarters where m fault data are located.
[0106] The lift calculation formula for the rare item association rule X r →Y is:
[0107]
[0108] In the formula, represents the lift of the rare item association rule X r →Y, represents the rare item association rule X rConfidence of →Y Indicates that the rare variable h is included in the dynamic database D r Support degree
[0109] Rare item association rule X r The formula for calculating the conviction degree of →Y is as follows
[0110]
[0111] In the formula Indicates the conviction degree of the rare item association rule X r →Y Indicates that the rare variable h is included in the dynamic database D r Support degree Indicates the rare item association rule X r →Y Confidence
[0112] Rare item association rule X r The formula for calculating the leverage of →Y is as follows
[0113]
[0114] In the formula Indicates the leverage of the rare item association rule X r →Y Indicates that both the fault characteristics and environmental characteristics of the dynamic database D contain the corresponding rare element h r And the processing result h y Support degree Indicates that the rare variable h is included in the dynamic database D r Support degree Indicates that the processing result h is included in the dynamic database D y Support degree
[0115] For X f →Y, the calculation formulas for its support degree, confidence degree, lift degree, conviction degree and leverage degree, and the calculation method for the diagnostic standard of the conditional importance degree of rare variables in fault characteristics and environmental characteristics are also as shown above
[0116] (2) Association rule X that contains both rare items and frequent items r →Y
[0117] If the association rules in the dynamic database D have both rare variable sets and frequent variable sets, the association rule that contains both rare items and frequent items is defined as
[0118] X f +X r →Y
[0119] In the formula, X fThe frequent variable set representing the set of fault causes, fault characteristics, and environmental characteristics in the dynamic database D. f represents a frequent item, and X r The rare variable set representing the set of fault causes, fault characteristics, and environmental characteristics in the dynamic database D. r represents a rare item, and Y represents the set of processing results for power grid circuit faults.
[0120] The association rule X that contains both rare and frequent items r →Y, whose support The calculation formula is:
[0121]
[0122] In the formula, w i,j ∈X f Indicates that the input variable belongs to the frequent variable, and h f Represents X f One of the frequent variable sets in. w i,j =h f Is a fault record where the input variable is a frequent variable. The numerator represents the fault record where the input variable contains both the frequent variable h f , and the rare variable h r .
[0123] The association rule X that contains both rare and frequent items r →Y, whose confidence The calculation formula is:
[0124]
[0125] In the formula, Represents the support that contains both the frequent variable h f , the rare variable h r , and the reclosing result is h y . Represents the support in the fault characteristics and environmental characteristics of the dynamic database D that contains both the corresponding rare element h r and the processing result h y .
[0126] The association rule X that contains both rare and frequent items r →Y, whose lift The calculation formula is:
[0127]
[0128] In the formula, Represents the lift that contains both the frequent variable h f , the rare variable h r and its corresponding processing result h y . Indicates that the rare variable h is included in the dynamic database D r Support degree.
[0129] The association rule X that contains both rare items and frequent items r →Y, its confidence The calculation formula is as follows:
[0130]
[0131] In the formula, Indicates that the frequent variable h is included at the same time f , rare variable h r And its corresponding processing result h y Confidence degree.
[0132] The association rule X that contains both rare items and frequent items r →Y, its leverage calculation formula is as follows:
[0133]
[0134]
[0135] In the formula, Indicates that the frequent variable h is included at the same time f , rare variable h r And its corresponding processing result h y Leverage.
[0136] The dynamic database D changes over time because new records are added or previous records are deleted. In addition, when the dynamic database D is updated, a new threshold may be switched to generate the required set of association rules. For the currently generated set of association rules, a simple but ineffective solution is to re-execute the entire mining algorithm from scratch for each modified data set and updated value.
[0137] FP-tree (Frequent Pattern Tree) is a data mining technique for efficiently mining frequent item sets. The FP-tree algorithm effectively reduces the number of database scans and improves the data compression rate by converting the original transaction data set into a compact tree structure, thus significantly enhancing the efficiency of frequent item set mining. For a given dynamic database D, in the traditional FP-growth (Frequent Pattern growth) algorithm, generating frequent item sets from it requires scanning the entire database and constructing an fp-tree. However, the faulty database changes over time because new records are added or previous records are deleted. There are four problems related to frequent item set mining, specifically:
[0138] Problem 1: When several new fault instances are added to it, generating frequent item sets does not require scanning the entire database, but constructs an updated fp-tree by scanning the updated part of the database.
[0139] Problem 2: When several fault instances in it are modified, generating frequent item sets requires calculating a new minimum support threshold and constructing an updated fp-tree.
[0140] Problem 3: When several existing fault instances are deleted from it, generating frequent item sets requires calculating a new support threshold and constructing an updated fp-tree.
[0141] Problem 4: When the support threshold is modified, update the fp-tree to generate frequent item sets that meet the given threshold.
[0142] To solve the above four problems, the traditional FP-grwth algorithm needs to scan the entire database, reconstruct the FP-tree, and generate the corresponding association rules. In this way, the computational complexity is significantly increased, affecting the efficiency of generating the entire association rules.
[0143] To solve this problem, the present invention develops an efficient incremental mining technology based on DDT-tree. Based on the dynamic database D and combined with the fp-growth algorithm, it can effectively derive a new set of patterns and a set of rare association rules with less execution time and space usage when updating the database, without losing information, and generate a complete set of association rules from the dynamically changing fault database.
[0144] S3: Construct the dual-pattern dynamic tree DDT-tree of the dynamic database D, update the dynamic database D to obtain incremental data, and update the dual-pattern dynamic tree DDT-tree.
[0145] It should be noted that: The DDT-tree is the abbreviation of Dual-pattern Dynamic Tree, that is, the dual-pattern dynamic tree. The DDT-tree is based on the FP-tree and consists of a root node labeled "null" and many subtrees. Each node in the subtrees has an item (corresponding to each data in the dynamic database D) as a prefix. Each node in the tree has different attributes, including the name of the item, the list of its descendants, the count, and the pre-order value.
[0146] Transactions in the dynamic database D (each transaction contains multiple items, and each transaction corresponds to each row of data in the dynamic database D) are inserted into the DDT-tree one by one according to a predefined item order. The items in the transaction and their support counts are maintained by the header table in descending order of frequency. Inserting transactions in descending order of frequency aims to achieve the maximum compactness of the tree structure. The most frequent items are thus closer to the root of the DDT-tree because they are more likely to be shared. The header table is updated each time to achieve the current descending order of frequency items. Before inserting a transaction into the DDT-tree, if any of its paths deviate from the current descending order of frequency items in the header table, the path is dynamically rearranged by recursively swapping adjacent nodes to achieve the current order.
[0147] Figure 3 The following is the algorithm pseudocode for constructing and updating the DDT-tree in the embodiments of the present invention.
[0148] Refer to Figure 3 As shown, to construct the dual-mode dynamic tree DDT-tree of the dynamic database D, the specific process includes:
[0149] Predefinition: Each item is each row of data in the dynamic database D. An item is also called an item. Each transaction includes multiple items, and each item corresponds to each data in the dynamic database D. The root node of the dual-mode dynamic tree DDT-tree is null, where null is an empty node, and each node of each subtree in the DDT-tree is an item.
[0150] S301: Calculate the support threshold and filter out eligible items.
[0151] Specifically, according to the minimum support minsupp of the general association rule X→Y, calculate the minimum support threshold of the general association rule X→Y According to Judge the frequent variable set X in the dynamic database D f And the rare variable set X r , According to Calculate the support degree of each rare variable, and filter out the eligible rare variables in X r Qualified rare variables.
[0152] S302: Insert the items that meet the threshold in the transaction into the header table, and rearrange the header table according to the current descending order of item frequencies.
[0153] Specifically, scan the transaction, and insert the eligible common items and rare items in the dynamic database D into the header table in descending order of frequency. An item will only be inserted into the header table if it does not exist previously. If the item already exists, its count in the header table will simply be incremented by 1. Then rearrange the header table according to the current descending order of item frequencies.
[0154] S303: Insert the transactions sorted in descending order of frequency into the DDT-tree.
[0155] S304: Update the DDT-tree according to the newly arranged header table until all transactions in the dynamic database D are inserted into the DDT-tree.
[0156] When any path of the DDT-tree deviates from the current frequency-descending item sorting of the header table, the DDT-tree will be dynamically rebuilt. To achieve the new item sorting, the paths of the DDT-tree are adjusted by recursively swapping adjacent nodes. If there are intermediate nodes between the swapped nodes in any path of the DDT-tree, the upper-level node is swapped down to reach the lower-level node. A key property of the FP-Tree-like structure is that according to the order of nodes in the path, the count of any node cannot exceed the count of its parent node. Therefore, whenever the DDT-tree needs to be updated to achieve the sorting of the current header table, it will perform path adjustment to maintain this property.
[0157] Path adjustment: To maintain the properties of the DDT-tree, if a parent node needs to be swapped with its child node with a smaller count, a new node with the same name is inserted as the sibling of the parent node. The support count of the parent node is set to the support count of its child node, and the support count of the sibling node is set to the difference between the support counts of the parent node and the child node. When the support counts of the parent node and the child node are equal, only the swap operation between the nodes is performed. If there are two sibling nodes with the same name after the swap operation, they are recursively merged by adding their support counts. The split, swap, and merge operations are explained in detail through an example below.
[0158] For example, there is a path in the DDT-tree where node X is the parent of node Y, and node Y is the parent of node Z, and it is required to swap the items represented by nodes Y and Z. If the count of node Y in the DDT-tree is greater than the count of node Z, but the count of the item represented by node Y in the header table is less than the count of the item represented by node Z, then perform the following path adjustment steps 1-3; otherwise, only perform steps 2 and 3:
[0159] Step 1: Split node Y into Y`, and insert it as a child of X, such that Y`.count = Y.count – Z.count. The count of node Y is reset to the count of node Z, and all children of node Y except node Z are assigned to the newly created Y` node.
[0160] Step 2: Swap the parent-child links of nodes Y and Z. Node Z will become the new child of node X, and the new parent of node Y.
[0161] Step 3: If node X has another child node named W, and the item it represents is the same as node Z, then add the count of node W to the count of node Z and delete node W. Insert the corresponding child node path of node W as a new child node path into node Z.
[0162] Figure 4 This is the pseudocode of the incremental mining algorithm based on DDT-tree in the embodiments of the present invention.
[0163] Dynamic database update includes deleting transactions or adding transactions, which may all cause changes to the minimum threshold minsupp of the support degree of association rules. hs The change of the threshold may change the frequency of patterns, making some frequent item sets become rare and some rare item sets become frequent. Existing algorithms will execute from scratch to generate the pattern and rule sets under the new threshold value. In this case, the incremental dual-mode mining algorithm based on DDT-tree will not repeat the tree construction stage, but only update the changed part of the DDT-tree and then execute the mining of patterns, as Figure 4 shown.
[0164] S4: According to the updated dual-mode dynamic tree DDT-tree, based on the association rule evaluation index, update the association rules, and construct a fault knowledge base for the power grid lines according to the updated association rules.
[0165] After the dual-mode dynamic tree DDT-tree is updated, based on the DDT-tree incremental mining method, the association rules between the device fault types and features can be efficiently extracted. This method reveals the internal connection between the fault types and features and can dynamically update the rules as new data is added.
[0166] The main construction process of the fault knowledge base for power grid lines includes:
[0167] 1) Obtain real-time data, select relevant elements to dynamically update the dynamic database D of power grid line faults, and update the DDT-tree.
[0168] 2) Use the incremental mining based on DDT-tree to perform association rule mining.
[0169] 3) Evaluate the mined association rules according to the evaluation index of the association rules.
[0170] 4) Put the evaluated association rules into the fault knowledge base.
[0171] Figure 5 This is the schematic diagram of the association rule judgment process of the fault knowledge base in the embodiments of the present invention. Figure 5Among them, the association rule judgment process of the fault knowledge base includes: obtaining real-time fault data in the dynamic database D, matching and judging the fault data with the association rules of the fault knowledge base. If the judgment result is a match, output the fault diagnosis description and update the fault knowledge base; if the judgment result is a mismatch, use the evaluation index of the association rule to evaluate the mined association rule. If the evaluation passes, add a new association rule. If the evaluation fails, output the fault diagnosis description and update the fault knowledge base.
[0172] After constructing the fault knowledge base, online fault diagnosis is carried out through the fault knowledge base.
[0173] Figure 6 It is a schematic diagram of the fault knowledge base in the embodiment of the present invention.
[0174] Refer to Figure 6 As shown, the fault knowledge base of the power grid line specifically includes natural factors, and the natural factors include lightning, ice disaster, wind disaster, geological disaster, mountain fire and others. Extract the time distribution characteristics, space distribution characteristics and meteorological association characteristics from the natural factors. The meteorological association characteristics include the highest temperature, the lowest temperature, air pressure, wind speed, rainfall and humidity, etc.
[0175] Figure 7 It is a schematic diagram of the processing process of the power grid line fault prediction method in the embodiment of the present invention. Figure 7 Among them, the processing process of the power grid line fault prediction method includes: collecting the real-time data of the power grid line, judging whether the collected real-time data is abnormal. If the data is abnormal, after preprocessing the data, extract the fault characteristics; if the data is not abnormal, input it into the trained fault prediction network to obtain the prediction result, and judge the level of the fault risk according to the prediction result. If the fault risk is low, output the fault diagnosis result and its description. If the fault risk is high, extract the fault characteristics, match and judge the extracted fault characteristics with the fault knowledge base, and output the fault diagnosis result and its description. The training process of the fault prediction network includes: obtaining training data, where the training data includes historical power grid operation data and historical meteorological data, inputting the training data into the fault prediction network, outputting the prediction result, calculating the loss between the prediction result and the training data, and optimizing the fault prediction network through the backpropagation of this loss.
[0176] The embodiment of the present invention proposes a power grid line fault prediction method. Refer to Figure 7 As shown, this method is based on the fault knowledge base of the power grid line and includes:
[0177] Obtain power grid operation data and meteorological prediction data, preprocess them to obtain the input data to be measured;
[0178] After inputting the input data to be measured into the trained fault prediction network, obtain the prediction result;
[0179] The prediction results are combined with the power grid line fault knowledge base for matching judgment to obtain the prediction results of power grid line faults.
[0180] Figure 8 It is a schematic diagram of the network structure of the fault prediction network in the embodiment of the present invention.
[0181] Refer to Figure 8 As shown, the fault prediction network includes: an input layer, a projection layer, a Transformer-LSTM layer, a fully connected layer, and an output layer.
[0182] The input layer includes two input data processing methods for continuous variables and categorical variables. Categorical variables are initially one-hot encoded, then concatenated, and then the corresponding embedded feature space is searched to obtain an embedding matrix of dimension d. Continuous variables (after normalization) are first projected into a d-dimensional vector space so that the mapping process from predefined data features to the homogenized neural network input is fully learnable. These two types of inputs undergo similar matrix transformations and are independent of each other. The continuous variables and categorical one-hot encoded features can be directly concatenated.
[0183] Projection layer: After the continuous variables and categorical one-hot encoded features are concatenated, a projection function is applied to both of them simultaneously to obtain T×d-dimensional features, simplifying the transformation process. Therefore, the dimension of the embedding matrix becomes N×d, and the dimension d of the projection function must be adjusted according to the size of the concatenated input to produce an adjusted model dimension value. Thereafter, the effective model dimension value is referred to as d eff , which corresponds to the input size of the continuous variable projection layer, and the adjusted model dimension input size (including the dimension corresponding to the zero value of the categorical variable) is referred to as d adj .
[0184] d adj and d eff The relationship between them is given by the formula:
[0185] d adj =d eff +N
[0186] Transformer-LSTM layer: The network structure of this layer improves the model for time series tasks based on the Transformer model, mainly modifying the decoder in the traditional Transformer model. Considering that LSTM can capture the features of time series, and the output of the Transformer model encoder is essentially also a sequence. Based on this, the embodiment of the present invention combines the different characteristics of Transformer and LSTM, adds an LSTM layer to the decoder of the Transformer model, and constructs a new Transformer-LSTM fusion model (i.e., the fusion network layer).
[0187] The specific working process of the Transformer-LSTM fusion model is as follows:
[0188] First, relative position encoding is adopted to inject the processed sequence data X m into the multi-head attention mechanism of the encoding layer of the Transformer-LSTM model, and its local feature A m is calculated by the formula:
[0189] A m = multiheadAttention(X m )
[0190] Subsequently, the feature weight S m and the local attention mechanism M (m) are calculated, and their formulas are as follows:
[0191] S m = Softmax(Relu((W q A m ))) T W k A m ))
[0192]
[0193] In the formula, W k , W q , W v are the weight matrices of the multi-head attention mechanism respectively, used to capture multi-level change features.
[0194] Then, the multi-level data feature vectors output by the encoding layer are concatenated to obtain Z, and input into the decoder LSTM model to continue the time series features of the data, and finally complete the prediction. The formula for extracting time series features is as follows:
[0195]
[0196] In the formula, is the hidden state at time t, is the hidden state at time t-1, W m is the parameter matrix of the decoder LSTM layer, and θ m is the hyperparameter of the decoder LSTM layer.
[0197] The fully connected layer is a structure in the neural network where each neuron is connected to all neurons in the previous layer. It performs a non-linear transformation on the features through the weight matrix to achieve high-level feature integration and abstraction.
[0198] The output layer is the final layer of the neural network. Depending on the type of task, different activation functions are used to map the features learned by the network into the final prediction results.
[0199] An embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the power grid line fault knowledge base construction method as described above or the power grid line fault knowledge base construction method as described above.
[0200] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the power grid line fault knowledge base construction method as described above or the power grid line fault prediction method as described above.
[0201] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: ROM, RAM, magnetic disk, optical disk, etc.
[0202] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a knowledge base of power grid line faults, characterized in that: include: Collect power grid operation data and meteorological data to build a dynamic database of power grid line faults; Based on the dynamic database, design an association rule evaluation index for power grid line faults; Constructing a dual-mode dynamic tree of the dynamic database, updating the dynamic database, mining incremental data, and updating the dual-mode dynamic tree; According to the updated dual-mode dynamic tree, based on the association rule evaluation index, the association rules are updated, and according to the updated association rules, a fault knowledge base of the power grid line is constructed.
2. The method for constructing a knowledge base of power grid line faults according to claim 1, characterized in that: The dynamic database (D) of the power grid line fault is expressed as: Where, t i represents the number of the fault record in the i-th row, i∈[1,2,…m], m represents the number of fault record numbers, f1 represents the voltage level when the power grid circuit fault occurs, f2 represents the fault phase when the power grid circuit fault occurs, f3 represents the fault time when the power grid circuit fault occurs, e1 represents the value of the first environmental characteristic element when the power grid circuit fault occurs, e 10 represents the value of the tenth environmental characteristic element when a power grid circuit fault occurs, R represents the set of fault causes when a power grid circuit fault occurs, Y represents the set of processing results for the power grid circuit fault, and w i,j Indicates the characteristic element value corresponding to each fault record, y i,15 Indicates the processing result corresponding to a specific power grid circuit fault, i∈[1,2,…,m], m represents the number of fault record numbers, j∈[1,2,…,15].
3. The method for constructing a power grid line fault knowledge base according to claim 1, characterized in that: The association rule evaluation indicators of the power grid line fault include: support, confidence, promotion, conviction and leverage. Thresholds are set according to the five association rule evaluation indicators, the association rules mined based on the dynamic database (D) are evaluated, and the association rules are updated.
4. The method for constructing a knowledge base of power grid line faults according to claim 1 or 3, characterized in that: For the dynamic database (D), the association rules mined include: general association rule X→Y, rare item association rule X r →Y and association rule X containing both rare items and frequent items r →Y; Calculate the evaluation indicators of these three association rules respectively.
5. The method for constructing a knowledge base of power grid line faults according to claim 1, characterized in that: The dual-mode dynamic tree DDT-tree of the dynamic database (D) is constructed, and the construction process includes: Predefined: Each item is a row of data in the dynamic database (D), each transaction includes multiple items, and each item corresponds to each data in the dynamic database; the root node of the dual-mode dynamic tree (DDT-tree) is null, null is an empty node, and each node of each subtree in the dual-mode dynamic tree is an item; Calculate the support threshold and filter out items that meet the conditions; Insert the items in the transaction that meet the threshold into the header table, and re-arrange the header table in descending order according to the current item frequency; Insert the transactions sorted in descending frequency into a dual-mode dynamic tree (DDT-tree); The dual-mode dynamic tree (DDT-tree) is updated according to the newly arranged header table until all transactions in the dynamic database (D) are inserted into the dual-mode dynamic tree (DDT-tree).
6. A method for predicting power line faults, the method being based on the power line fault knowledge base as claimed in claim 1, characterized in that: The method includes: Obtain power grid operation data and meteorological forecast data, pre-process them, and obtain input data to be tested; After inputting the input data to be tested into the trained fault prediction network, the prediction result is obtained; The prediction results are combined with the power grid line fault knowledge base for matching and judgment to obtain the prediction results of the power grid line fault.
7. The method for predicting power grid line faults according to claim 6, characterized in that: The fault prediction network includes: an input layer, a projection layer, a Transformer-LSTM layer, a fully connected layer and an output layer.
8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for constructing a knowledge base of power grid line faults as claimed in any one of claims 1 to 6 or the method for predicting power grid line faults as claimed in claim 6 or 7 is implemented.
9. A computer storable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a power grid line fault knowledge base as described in any one of claims 1 to 6 or the method for predicting power grid line faults as described in claim 6 or 7 is implemented.
Citation Information
Patent Citations
Power grid equipment fault prediction method and system based on knowledge graph under complex weather
CN117436351A
Cited By
Entertainment culture content transmission mode recognition method based on improved FP-GROWTH algorithm
CN121188730A
Network communication health assessment method based on FP-Growth algorithm
CN121691107A