Intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism
Through the intelligent substation fault diagnosis method with multi-source data fusion and self-learning mechanism, the problems of inconsistent alarm information and redundancy, inaccurate feature extraction, inaccurate fault positioning and poor system adaptability are solved, efficient and accurate fault detection and positioning are achieved, and the safety and operation and maintenance efficiency of the power system are improved.
Patent Information
- Application Number
- CN202510399080.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional fault diagnosis methods have problems such as inconsistent alarm information and redundancy, inaccurate feature extraction, inaccurate fault positioning and poor system adaptability in smart substations, resulting in low safety and operation and maintenance efficiency of the power system.
Multi-source data fusion and self-learning mechanism are adopted to collect and standardize alarm information, establish a dynamic weight fault feature extraction model, combine feature collaboration enhancement and topological structure for rapid positioning, and introduce an adaptive optimization mechanism of the knowledge graph to form a closed-loop learning mechanism.
It realizes efficient and accurate fault detection and positioning, improves the safety and operation and maintenance efficiency of the power system, reduces false alarms and missed alarms, and adapts to equipment aging and environmental changes.
Smart Images

Figure CN120334624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent substation technology, and particularly relates to an intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism. Background Art
[0002] In modern power systems, especially in the operation and management of intelligent substations, fault diagnosis and rapid response are important links to ensure the stability and reliability of the power grid. Traditional fault diagnosis methods mainly rely on alarm information provided by a single data source (such as SCADA system or IED device), and perform fault identification and location through preset rules or static models. However, these traditional methods have many limitations and are difficult to meet the requirements of increasingly complex power systems.
[0003] First of all, due to the time synchronization problem between different devices, the alarm information collected from multiple data sources may have time deviation, which affects the accuracy of the fault time sequence relationship. For example, when judging the sequence of "protection blocking" and "GOOSE link break", a time difference of milliseconds may lead to misjudgment, thus delaying the best time for fault handling. In addition, the same fault may trigger multiple system repeated alarms, increasing the complexity of data analysis and calculation burden, and reducing the diagnosis efficiency.
[0004] Secondly, traditional feature extraction methods usually rely on manually setting weights or using fixed algorithms. This method is prone to ignoring some key features or overemphasizing unimportant features, resulting in inaccurate fault feature extraction. Especially in the face of multi-type and multi-level fault scenarios, single-dimensional feature analysis often cannot comprehensively reflect the essential features of faults, thus affecting the accuracy of fault diagnosis.
[0005] Furthermore, most of the existing fault location technologies adopt simple rule matching or single-index judgment, lacking comprehensive consideration of complex fault scenarios. For example, when dealing with cross-device link faults or multi-feature coupling faults, it is difficult to accurately locate the fault point only by relying on basic feature matching, resulting in an extended fault troubleshooting time, increasing the system downtime and maintenance cost.
[0006] Finally, once the traditional fault diagnosis system is deployed, it is difficult to self-adjust according to the actual operation situation. Over time, equipment aging or environmental changes may lead to a decline in its diagnosis ability, resulting in false alarms or missed alarms. Such a system lacking an adaptive optimization mechanism is difficult to maintain an efficient and stable fault detection ability for a long time, affecting the overall operation and maintenance level of the power system.
[0007] In view of the above background, there is an urgent need for a new fault diagnosis method that can collect and standardize multi-source alarm information in real time, dynamically adjust feature weights, quickly and accurately locate faults, and have the ability of self-learning and optimization, so as to improve the safety and operation and maintenance efficiency of the power system. The present invention precisely proposes an intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism, aiming to comprehensively improve the accuracy and timeliness of fault diagnosis. Summary of the Invention
[0008] The purpose of the present invention is to provide an intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism, which solves the problems of inconsistent and redundant alarm information, inaccurate feature extraction, inaccurate fault location, and poor system adaptability, realizes efficient and accurate fault detection and location, and improves the safety and operation and maintenance efficiency of the power system.
[0009] The technical solution of the present invention to solve the above technical problems is as follows:
[0010] The present invention provides an intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism, including the following steps:
[0011] S1: Real-time collection and standardized processing of multi-source alarm information;
[0012] S2: Extract key fault features from the standardized alarm data, and establish a dynamic weight fault feature extraction model;
[0013] S3: Fast fault location based on feature collaboration enhancement and dynamic weights;
[0014] S4: Fault correlation adaptive optimization based on knowledge graph.
[0015] In the step S1, alarm information is collected from a multi-source system to eliminate time deviation and redundant alarms.
[0016] Furthermore, in the step S1, by aligning timestamps, the clock difference between SCADA and IED is eliminated.
[0017] Timestamp alignment formula:
[0018]
[0019] Wherein, Δt is the maximum time deviation between the timestamp t SCADA recorded by the SCADA system and the timestamp t IED recorded by the intelligent electronic device; t SCADA is the alarm timestamp recorded by the SCADA system; t IEDThe alarm timestamp recorded for the intelligent electronic device; τ is the upper limit of the clock error between systems; ε is the maximum allowable time deviation threshold.
[0020] Further, in the step S1, based on the redundant alarm elimination method of hybrid similarity and dynamic threshold, calculate the semantically repeated alarms that can be merged, reducing the computational complexity of subsequent feature extraction;
[0021]
[0022] Among them, Merge(A i ,A j ) is a logical judgment expression used to determine whether to merge alarm A i and A j ; A i ,A j represent two different alarms in the system respectively; V i ,V j are the enhanced vectors of the alarm text; cos(V i ,V j ) is the cosine similarity, measuring the distribution relationship between two vectors in the high-dimensional space; J(A i ,A j ) is the Jaccard similarity, measuring the proportion of common words in two pieces of text; λ is the weight parameter used to adjust the relative importance of the cosine similarity and the Jaccard similarity; θ merge is the basic threshold; η, σ, Δt are the time decay coefficient, time constant, and alarm time difference.
[0023] In the step S2, extract the key fault features from the alarm data, convert the original alarm data into a weighted feature vector, and quantify its contribution degree to different fault types;
[0024]
[0025] Among them, ω j (t) is the weight of the j-th feature at time point t, reflecting the importance of this feature for fault diagnosis at the current moment; E j is the entropy value of the j-th feature, used to measure the dispersion degree of the distribution of this feature; n is the total number of features, that is, the number of all possible features considered in the analysis process; α is the proportional coefficient of the dynamic correction term, used to adjust the influence degree of the historical fault frequency on the weight; N fail (j) is the number of times the j-th feature has occurred in past faults, reflecting the association strength between a specific feature and known faults; N total is the total number of faults, used as the denominator to help standardize N fail (j).
[0026] In step S3, based on the feature vectors and the alarm-fault association strength, combined with the topological structure and the real-time status, the fault point is located.
[0027]
[0028] Among them, F k represents the k-th candidate fault point or hypothesis. During the fault diagnosis process, multiple possible fault points are generated as candidates, and F k is one of these candidate faults; Score(F k ) is the confidence score for the candidate fault F k , which is used to evaluate the possibility of this candidate fault becoming a real fault; w i is the weight of the i-th feature, reflecting the importance of a specific feature in fault diagnosis; f i (F k ) is the basic feature matching item, indicating the matching degree between the candidate fault F k and the i-th feature; δ is the proportionality coefficient of the feature collaboration item, which is used to adjust the influence of the key features that appear simultaneously on the final score; β(t) is the time-dependent function of the dynamic topology weight; TopoRank(F k ) is the ranking of the candidate fault F k in the power system topological structure; γ is the proportionality coefficient of the real-time status penalty item, which determines the influence degree of the real-time device status data on the final score; S real (F k ) is the health score calculated based on the real-time status of the device.
[0029] In step S4, the association strength between the alarm and the fault is dynamically adjusted according to the actual operation data, and a self-learning mechanism is introduced to dynamically optimize the fault location effect, forming a closed-loop learning mechanism.
[0030] Furthermore, in step S4, when some alarm data are proven to be frequently misreported or fail to identify specific faults, the association strength between the alarm and the fault is allowed to be automatically adjusted.
[0031] Furthermore, in step S4, when the performance of the device deteriorates over time, which may cause the alarm signals sent by it to become unreliable, by continuously updating R(A→F), the system gradually reduces the trust in the alarms of this device and instead relies more on other more reliable metrics;
[0032]
[0033] Among them, R new(A→F) is the updated relationship strength from alarm A to fault F, which reflects the likelihood of a connection between alarm A and fault F in the historical records; μ is the forgetting factor used to control the impact of historical data on the current relationship strength; R old is the previous, old relationship strength value, which is the degree of association between alarm A and fault F calculated based on previous data. In each update process, R old is the relationship strength value obtained from the previous calculation; N correct (A,F) is the number of times that alarm A correctly diagnoses fault F; N Total (A) is the total number of times that alarm A appears.
[0034] The present invention has the following beneficial effects:
[0035] The present invention proposes an intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism, which has significant advantages compared with traditional technologies. The following are the main beneficial effects of the present invention and its advantages over traditional technologies:
[0036] (1) In traditional fault detection systems, usually relying on single or limited data sources for fault detection, which often leads to incomplete or biased alarm information. In contrast, through real-time collection and standardized processing of multi-source alarm information (S1) in the present invention, it can not only collect alarm information in real time from multiple sources (such as SCADA systems, IED devices, etc.), but also ensure the consistency and accuracy of alarm information through timestamp alignment and redundant alarm elimination technologies. This method effectively solves the time deviation problem between different systems and reduces the interference caused by repeated alarms, providing a high-quality data basis for subsequent analysis.
[0037] (2) Traditional methods often rely on manually setting weights or using fixed algorithms in the feature extraction process, which may lead to some key features being ignored or overemphasized, affecting the accuracy of fault diagnosis. In contrast, the dynamic weight fault feature extraction model (S2) in the present invention uses the entropy weight method to automatically evaluate the effectiveness of features and dynamically adjusts feature weights in combination with historical fault frequencies to ensure the selection of key features that can effectively distinguish different types of faults. In addition, this model also avoids the subjective bias in the traditional expert scoring method, making the fault feature extraction more objective and accurate.
[0038] (3) Traditional fault location methods usually adopt simple rule matching or single-index judgment, making it difficult to meet the fault diagnosis requirements in complex scenarios. The fault rapid location method (S3) based on feature collaborative enhancement and dynamic weights proposed by the present invention comprehensively considers multiple factors such as basic feature matching, feature collaborative effect, dynamic topology weight, and real-time status penalty, and realizes the rapid location of faults by calculating the confidence score of each candidate fault point. This method can not only identify the most likely fault point, but also improve the location accuracy for complex fault scenarios, thus significantly shortening the fault handling time.
[0039] (4) Once a traditional fault diagnosis system is deployed, it is difficult to self-adjust according to the actual operation situation. Over time, equipment aging or environmental changes may lead to a decline in its diagnostic ability. The present invention introduces a fault association adaptive optimization mechanism based on a knowledge graph (S4). By continuously updating the relationship strength between alarms and faults, the system can adapt to long-term influencing factors such as equipment aging and environmental changes, reducing false alarms and missed alarms. This self-learning mechanism enables the system to continuously optimize its fault diagnosis ability during operation, maintaining high accuracy and reliability.
[0040] In summary, through a series of innovative technologies and methods, the present invention shows significant advantages in improving the accuracy of alarm information, enhancing the accuracy of fault feature extraction, achieving rapid and accurate fault location, and having the ability of adaptive optimization. Compared with traditional technologies, it greatly improves the safety and stability of the power system, and has important application value and development prospects. Brief Description of the Drawings
[0041] Figure 1 It is a flow schematic block diagram of the fault diagnosis method of the present invention. Detailed Embodiments
[0042] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0043] Embodiment 1
[0044] The present invention provides a fault diagnosis method in the field of electricity. Referring to Figure 1 as shown, the fault diagnosis method in the field of electricity includes:
[0045] S1: Real-time collection and standardized processing of multi-source alarm information;
[0046] Collect alarm information from multi-source systems, eliminate time deviation and redundant alarms, and provide high-quality input for subsequent analysis.
[0047] ①There may be a millisecond-level deviation between the clocks of the substation SCADA system and IED devices. Therefore, by aligning the timestamps, the clock differences between SCADA and IED are eliminated to ensure the accuracy of the alarm timing relationship in subsequent analyses, such as determining the order of "protection blocking" and "GOOSE link break".
[0048] Timestamp alignment formula:
[0049]
[0050] Among them, Δt is the maximum time deviation between the timestamp t SCADA recorded by the SCADA system and the timestamp t IED recorded by the intelligent electronic device (IED); t SCADA is the alarm timestamp recorded by the SCADA system; t IED is the alarm timestamp recorded by the intelligent electronic device (IED); τ is the upper limit of the clock error between systems; ε is the maximum allowable time deviation threshold
[0051] ②The same fault may trigger multiple system repeated alarms (such as the network recorder and IED reporting "GOOSE link break" at the same time). The present invention proposes a redundant alarm elimination method based on hybrid similarity and dynamic threshold to calculate and merge semantically repeated alarms, reducing the computational complexity of subsequent feature extraction.
[0052]
[0053] Among them, Merge(A i , A j ) is a logical judgment expression used to determine whether to merge alarm A i and A j ; A i , A j respectively represent two different alarms in the system; V i , V j are the enhanced vectors of the alarm text. Based on the traditional TF-IDF, core keywords such as "GOOSE link break" are weighted and device entity information is embedded to enhance the sensitivity to the core semantics of the fault, reduce the interference of common words, and make the text representation closer to the actual business requirements; cos(V i , V j ) is the cosine similarity, which measures the distribution relationship (semantic similarity) of two vectors in the high-dimensional space; J(A i , A j ) is the Jaccard similarity, which measures the proportion of common words in two texts (lexical overlap); λ is a weight parameter used to adjust the relative importance of the cosine similarity and the Jaccard similarity; θ mergeis the basic threshold; η, σ, and Δt are the time decay coefficient, time constant, and alarm time difference, respectively.
[0054] S2: Establish a dynamic weight fault feature extraction model;
[0055] Extract key fault features from the standardized alarm data in S1, and assign dynamic weights to each feature to form a feature vector.
[0056] Extract key fault features from a large number of alarms, convert the original alarms into weighted feature vectors, and quantify their contribution degrees to different fault types.
[0057]
[0058] Among them, ω j (t) is the weight of the j-th feature at time point t, reflecting the importance of this feature for fault diagnosis at the current moment; E j is the entropy value of the j-th feature, used to measure the discreteness of the distribution of this feature. The smaller the entropy value, the higher the discrimination degree of this feature in distinguishing different fault types. For example, the concentrated occurrence of the feature "switch port DOWN" in physical link faults means that it has a lower entropy value and a higher discrimination degree; n is the total number of features, that is, the number of all possible features considered in the analysis process; α is the proportional coefficient of the dynamic correction term, used to adjust the influence degree of historical fault frequency on the weight, and this parameter can be adjusted according to the requirements of the actual application scenario; N fail (j) is the number of times the j-th feature has occurred in past faults, reflecting the correlation strength between a specific feature and known faults; N total is the total number of faults, used as the denominator to help standardize N fail (j) to make the comparison between different features more fair and reasonable.
[0059] The present invention evaluates the effectiveness of features through the entropy value E j to ensure the selection of features that can effectively distinguish different types of faults. This is crucial for reducing false alarms and improving diagnostic accuracy. At the same time, introducing historical fault frequency allows the model to dynamically adjust the feature weights according to the actual operating status of the substation. For example, if a certain optical port has frequent faults recently, its weight is correspondingly increased to detect related problems earlier. Finally, the entropy weight method does not require manual intervention to set weights, reducing the subjective bias existing in the traditional expert scoring method, and meeting the strict requirements of the power system for the objectivity of the decision-making process.
[0060] S3: Fast fault location based on feature collaborative enhancement and dynamic weights;
[0061] Use the eigenvector of S2 and the alarm-fault correlation strength, combined with the topological structure and real-time status, to quickly locate the fault point.
[0062]
[0063] Among them, F k represents the k-th candidate fault point or hypothesis. During the fault diagnosis process, multiple possible fault points are generated as candidates, and F k is one of these candidate faults; Score(F k ) is the confidence score for the candidate fault F k , which is used to evaluate the possibility of this candidate fault becoming a real fault. The higher the score, the more likely this candidate fault is the actual fault that occurred; w i is the weight of the i-th feature, which reflects the importance of a specific feature in fault diagnosis. Different features have different contribution degrees to locating different types of faults. Therefore, an appropriate weight needs to be assigned to each feature; f i (F k ) is the basic feature matching item, which represents the matching degree between the candidate fault F k and the i-th feature. For example, "multiple IED disconnections under the same switch" can be used to match physical link faults. The larger the value of f i (F k ), the better the match between the candidate fault F k and feature i; δ is the proportionality coefficient of the feature collaboration item, which is used to adjust the influence of key feature pairs that appear simultaneously (such as "GOOSE disconnection + SV packet loss") on the final score. By introducing δ, the interaction between features can be considered during the calculation to improve the location accuracy in complex fault scenarios; β(t) is the time-dependent function of the dynamic topology weight. According to different states of the power grid (transient or steady state), the value of β(t) will change to reflect the change of the criticality of equipment in the topological structure over time. For example, under transient conditions, the failure impact range of the core switch is larger, and the value of β(t) may be higher at this time; TopoRank(F k ) is the ranking of the candidate fault F k in the power system topological structure. Based on the SCD file (substation configuration description file), the importance of each device in its topological position can be calculated, so as to guide the priority investigation of those devices located in key positions; γ is the proportionality coefficient of the real-time status penalty item, which determines the influence degree of real-time device status data (such as load, temperature, etc.) on the final score. By adjusting γ, the role of online monitoring data in fault location can be controlled; S real (F k ) is the health score calculated based on the real-time status of the device. For example, if a certain device is currently in an overloaded state, then its Sreal (F k ) may be a negative value, indicating that the device has a high risk, which helps to suppress the misjudgment of nodes with high risks but normal operation.
[0064] S4: Adaptive Optimization of Fault Association Based on Knowledge Graph;
[0065] Dynamically adjust the association strength between alarms and faults according to the actual operation data, optimize the positioning effect of S3, and form a closed-loop learning mechanism.
[0066] S3 relies on information such as pre-set or learned feature weights and topological structures for fault location. However, over time, the device may age and the environmental conditions may change, resulting in the original model no longer being fully applicable. Therefore, in S4, the present invention introduces a self-learning mechanism to dynamically correct these relationships to adapt to the following new situations:
[0067] (1) Dynamically correct the impact of false alarms / missing alarms: As new data is added, especially when certain alarms are proven to be frequently false or fail to identify specific faults, the method allows for automatic adjustment of the association strength between alarms and faults, thereby reducing the future false alarm rate and missing alarm rate.
[0068] (2) Adapt to scenarios such as device aging: For example, the performance degradation of a certain sensor over time may cause the alarm signals it emits to become less reliable. By continuously updating R(A→F), the system can gradually reduce the trust in the alarms of this sensor and instead rely more on other more reliable metrics.
[0069]
[0070] Among them, R new (A→F) is the updated relationship strength from alarm A to fault F, and this value reflects the likelihood of a connection between alarm A and fault F in the historical record; μ is the forgetting factor, which is used to control the influence degree of historical data on the current relationship strength. A higher μ value means more emphasis on past data, while a lower value allows for a faster response to the latest information; R old is the previous, old relationship strength value, which is the association degree between alarm A and fault F calculated based on previous data. In each update process, R old is the relationship strength value obtained from the previous calculation; N correct (A,F) is the number of times that alarm A correctly diagnoses fault F; N Total (A) is the total number of times that alarm A appears.
[0071] Embodiment 2
[0072] An intelligent substation fault diagnosis system based on multi-source data fusion and self-learning mechanism, including the following units:
[0073] A data acquisition and processing unit, used for real-time acquisition and standardized processing of multi-source alarm information;
[0074] A model establishment unit, used to extract key fault features from the standardized alarm data and establish a dynamic weight fault feature extraction model;
[0075] A fault location unit, used for fast fault location based on feature collaborative enhancement and dynamic weight;
[0076] An optimization unit, used for adaptive optimization of fault association based on a knowledge graph.
[0077] In the data acquisition and processing unit, alarm information is collected from multi-source systems to eliminate time deviation and redundant alarms.
[0078] By aligning timestamps, the clock difference between SCADA and IED is eliminated.
[0079] Timestamp alignment formula:
[0080]
[0081] Where, Δt is the maximum time deviation between the timestamp t SCADA recorded by the SCADA system and the timestamp t IED recorded by the intelligent electronic device; t SCADA is the alarm timestamp recorded by the SCADA system; t IED is the alarm timestamp recorded by the intelligent electronic device; τ is the upper limit of the clock error between systems; ε is the maximum allowable time deviation threshold.
[0082] A redundant alarm elimination method based on hybrid similarity and dynamic threshold is used to calculate and merge semantically repeated alarms, reducing the computational complexity of subsequent feature extraction;
[0083]
[0084] Where, Merge(A i , A j ) is a logical judgment expression used to determine whether to merge alarm A i and A j ; A i , A j represent two different alarms in the system respectively; V i , V j are the enhancement vectors of the alarm text; cos(V i , V j) is the cosine similarity, which measures the distribution relationship between two vectors in a high-dimensional space; J(A i , A j ) is the Jaccard similarity, which measures the proportion of common words in two texts; λ is the weight parameter used to adjust the relative importance of the cosine similarity and the Jaccard similarity; θ merge is the basic threshold; η, σ, and Δt are the time decay coefficient, time constant, and alarm time difference.
[0085] In the model establishment unit, key fault features are extracted from the alarm data, the original alarm data is converted into a weighted feature vector, and its contribution degree to different fault types is quantified;
[0086]
[0087] Among them, ω j (t) is the weight of the j-th feature at time point t, reflecting the importance of this feature for fault diagnosis at the current moment; E j is the entropy value of the j-th feature, used to measure the dispersion degree of the distribution of this feature; n is the total number of features, that is, the number of all possible features considered in the analysis process; α is the proportional coefficient of the dynamic correction term, used to adjust the influence degree of the historical fault frequency on the weight; N fail (j) is the number of times the j-th feature has occurred in past faults, reflecting the association strength between a specific feature and known faults; N total is the total number of faults, used as the denominator to help standardize N fail (j).
[0088] In the fault location unit, based on the described feature vector and the alarm-fault association strength, combined with the topological structure and real-time status, the fault point is located,
[0089]
[0090] Among them, F k represents the k-th candidate fault point or hypothesis. During the fault diagnosis process, multiple possible fault points are generated as candidates, and F k is one of these candidate faults; Score(F k ) is the confidence score for the candidate fault F k , used to evaluate the possibility of this candidate fault becoming a real fault; w i is the weight of the i-th feature, reflecting the importance of a specific feature in fault diagnosis; f i (F k ) is the basic feature matching item, indicating the candidate fault F kThe degree of matching with the i-th feature; δ is the proportionality coefficient of the feature collaboration term, used to adjust the impact of the key features that appear simultaneously on the final score; β(t) is the time-dependent function of the dynamic topology weight; TopoRank(F k ) is the ranking of the candidate fault F k in the power system topology; γ is the proportionality coefficient of the real-time state penalty term, which determines the influence degree of the real-time device state data on the final score; S real (F k ) is the health score calculated based on the real-time state of the device.
[0091] In the optimization unit, the association strength between alarms and faults is dynamically adjusted according to the actual operation data, and a self-learning mechanism is introduced to dynamically optimize the fault location effect, forming a closed-loop learning mechanism.
[0092] When some alarm data are proven to be frequently misreported or fail to identify specific faults, the association strength between alarms and faults is allowed to be automatically adjusted.
[0093] When the performance of a device degrades over time, which may cause the alarm signals it sends to become unreliable, by continuously updating R(A→F), the system gradually reduces the trust in the alarms of this device and instead relies more on other more reliable metrics;
[0094]
[0095] Among them, R new (A→F) is the updated relationship strength from alarm A to fault F, and this value reflects the likelihood of a connection between alarm A and fault F in the historical record; μ is the forgetting factor, used to control the influence degree of historical data on the current relationship strength; R old is the previous, old relationship strength value, which is the degree of association between alarm A and fault F calculated based on previous data. In each update process, R old is the relationship strength value obtained in the previous calculation; N correct (A,F) is the number of times that alarm A correctly diagnoses fault F; N Total (A) is the total number of times that alarm A appears.
[0096] A computer device, comprising:
[0097] A processor; and a memory configured to store machine-readable instructions that, when executed by the processor, perform the above-mentioned intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism.
[0098] A storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism.
[0099] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a reference structure" does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0100] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism, characterized in that, Including the following steps: S1: Real-time collection and standardization processing of multi-source alarm information; S2: Extract key fault features from the alarm data after the above-mentioned standardization processing, and establish a dynamic weight fault feature extraction model; S3: Fast fault location based on feature collaborative enhancement and dynamic weight; S4: Adaptive optimization of fault association based on knowledge graph.
2. The intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 1, wherein In the step S1, alarm information is collected from multi-source systems to eliminate time deviation and redundant alarms.
3. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 2, characterized in that, In the step S1, by aligning timestamps, the clock difference between SCADA and IED is eliminated. Timestamp alignment formula: where, Δt is the maximum time deviation between the time stamp t recorded by the SCADA system SCADA and the time stamp t recorded by the intelligent electronic device IED ; t SCADA is the alarm time stamp recorded by the SCADA system; t IED is the alarm time stamp recorded by the intelligent electronic device; τ is the upper limit of the clock error between systems; ε is the threshold of the maximum allowable time deviation.
4. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 2, characterized in that, In the step S1, based on the redundant alarm elimination method of hybrid similarity and dynamic threshold, semantically repeated alarms that can be merged are calculated to reduce the computational complexity of subsequent feature extraction; Among them, Merge(A i , A j ) is a logical judgment expression used to determine whether to merge alarm A i and A j ; A i , A j respectively represent two different alarms in the system; V i , V j are enhancement vectors of the alarm text; cos(V i , V j ) is the cosine similarity, which measures the distribution relationship between two vectors in the high-dimensional space; J(A i , A j ) is the Jaccard similarity, which measures the proportion of common words in two texts; λ is a weight parameter used to adjust the relative importance of the cosine similarity and the Jaccard similarity; θ merge is the basic threshold; η, σ, and Δt are the time decay coefficient, the time constant, and the alarm time difference.
5. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 1, characterized in that, In the step S2, key fault features are extracted from the alarm data, and the original alarm data is transformed into a weighted feature vector to quantify its contribution degree to different fault types; where, ω j (t) is the weight of the j-th feature at time point t, reflecting the importance of this feature for fault diagnosis at the current moment; E j is the entropy value of the j-th feature, used to measure the degree of dispersion of the distribution of this feature; n is the total number of features, that is, the number of all possible features considered in the analysis process; α is the proportionality coefficient of the dynamic correction term, used to adjust the influence degree of the historical fault frequency on the weight; N fail (j) is the number of times the j-th feature has failed in the past, reflecting the correlation strength between a specific feature and known faults; N total is the total number of faults, used as the denominator to help standardize N fail (j).
6. The intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 1, characterized in that, In the step S3, based on the above-mentioned feature vector and the alarm-fault association strength, combined with the topological structure and real-time status, the fault point is located. Among them, F k represents the k-th candidate fault point or hypothesis. During the fault diagnosis process, multiple possible fault points will be generated as candidates, and F k is one of these candidate faults; Score(F k ) is the confidence score for the candidate fault F k , which is used to evaluate the possibility of this candidate fault becoming the real fault; w i is the weight of the i-th feature, reflecting the importance of a specific feature in fault diagnosis; f i (F k ) is the basic feature matching item, indicating the matching degree between the candidate fault F k and the i-th feature; δ is the proportionality coefficient of the feature collaboration item, which is used to adjust the influence of the key features that appear simultaneously on the final score; β(t) is the time-dependent function of the dynamic topology weight; TopoRank(F k ) is the ranking of the candidate fault F k in the power system topology; γ is the proportionality coefficient of the real-time state penalty item, which determines the influence degree of the real-time device state data on the final score; S real (F k ) is the health score calculated based on the real-time state of the device.
7. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 1, characterized in that, In the step S4, the association strength between alarms and faults is dynamically adjusted according to the actual operation data, and a self-learning mechanism is introduced to dynamically optimize the fault location effect, forming a closed-loop learning mechanism.
8. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 7, characterized in that, In the step S4, when some alarm data is proved to be frequently misreported or fails to identify specific faults, the association strength between alarms and faults is allowed to be automatically adjusted.
9. An intelligent substation fault diagnosis method based on multi-source data fusion and self-learning mechanism according to claim 7, characterized in that, In the step S4, when the performance of the device degrades over time, which may cause the alarm signals sent by it to become unreliable, by continuously updating R(A→F), the system gradually reduces the trust in the alarms of this device and instead relies more on other more reliable indicators. Among them, R new (A→F) is the updated relationship strength from alarm A to fault F, and this value reflects the likelihood of a connection between alarm A and fault F in the historical records; μ is the forgetting factor, which is used to control the influence degree of historical data on the current relationship strength; R old is the previous, old relationship strength value, which is the degree of association between alarm A and fault F calculated based on previous data. In each update process, R old is the relationship strength value obtained from the previous calculation; N correct (A,F) is the number of times that alarm A correctly diagnoses fault F; N Total (A) is the total number of times that alarm A appears.
Citation Information
Cited By
Secondary equipment fault diagnosis and positioning method and system based on correlation analysis
CN121145032A