Fault source positioning method, fault source positioning device, electronic equipment and product
By acquiring the current alarm data of the distributed energy system, matching fault association rules, and constructing a fault propagation graph, the problem of being unable to quickly locate the source of the fault in existing technologies is solved, and rapid and accurate fault source location is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SIGENERGY TECHNOLOGY (JIANGSU) CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing equipment management systems are unable to quickly locate the source of faults in distributed energy systems, relying mainly on periodic inspections and passive alarms, resulting in delayed responses.
By acquiring the current alarm data of the distributed energy system, matching fault association rules, constructing a fault propagation graph, and calculating the target probability of candidate fault sources based on time sequence and association rules, the fault source can be quickly located.
It enables rapid location of the source of the fault, significantly shortens the fault diagnosis time, changes the traditional delayed response mode, and improves the fault response speed.
Smart Images

Figure CN121935869A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a fault source location method, a fault source location device, electronic equipment and products. Background Technology
[0002] With the rapid development of distributed energy systems, the types and quantities of energy equipment have increased dramatically, leading to a corresponding increase in the complexity of equipment management. Existing equipment management systems mainly rely on periodic inspections and passive alarms for equipment fault location, which cannot quickly pinpoint the source of the fault. Summary of the Invention
[0003] This application provides a fault source location method, a fault source location device, an electronic device, and a product that can quickly locate the fault source.
[0004] In a first aspect, embodiments of this application provide a method for locating the source of a fault, including:
[0005] Obtain current alarm data from the distributed energy system;
[0006] Match the current alarm data with each association rule in the fault association rules;
[0007] Based on the matched association rules, a first fault propagation graph is constructed for each device involved in the current alarm data;
[0008] Based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time sequence association rules in the fault association rules, each candidate fault source is determined, and the target probability of each candidate fault source as a fault source is calculated; the target probability of each candidate fault source as a fault source represents the possibility that each candidate fault source is a fault source.
[0009] The candidate fault source with the highest target probability is determined as the fault source of the current alarm data.
[0010] In this embodiment, current alarm data of the distributed energy system is acquired and matched with various association rules in the fault association rules. Based on this, a first fault propagation graph of each device involved in the current alarm data can be constructed. Based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, each candidate fault source can be determined, and the target probability of each candidate fault source as the fault source is calculated. Thus, the candidate fault source with the highest target probability is determined as the fault source of the current alarm data. This solution intelligently matches the current alarm data with fault association rules and dynamically constructs a fault propagation graph focusing on the current fault scenario. It can integrate the time sequence of the current alarm data, the matched association rules, the constructed fault propagation graph, and the time series association rules for multi-dimensional probabilistic reasoning. Based on the target probability of each candidate fault source, the fault source can be located, changing the traditional delayed response mode that relies on manual inspection and passive alarms, and realizing rapid fault source location.
[0011] Secondly, embodiments of this application provide a fault source location device, comprising:
[0012] The data acquisition module is used to acquire the current alarm data of the distributed energy system;
[0013] The rule matching module is used to match the current alarm data with each association rule in the fault association rules;
[0014] The propagation graph construction module is used to construct a first fault propagation graph of each device involved in the current alarm data based on the matched association rules;
[0015] The candidate determination module is used to determine each candidate fault source based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time sequence association rules in the fault association rules, and to calculate the target probability of each candidate fault source as a fault source; the target probability of each candidate fault source as a fault source represents the possibility that each candidate fault source is a fault source.
[0016] The source determination module is used to determine the candidate fault source with the highest target probability as the fault source of the current alarm data.
[0017] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device enables the fault source localization method as described in the first aspect above.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a computer, implements the fault source localization method as described in the first aspect above.
[0019] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when run, causes the fault source localization method as described in the first aspect above to be executed.
[0020] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating the fault source localization method provided in an embodiment of this application;
[0023] Figure 2 This is another flowchart illustrating the fault source localization method provided in the embodiments of this application;
[0024] Figure 3 This is an example diagram illustrating the continuous learning and optimization of the system provided in the embodiments of this application;
[0025] Figure 4 This is a flowchart of the rapid fault source location and processing provided in the embodiments of this application;
[0026] Figure 5 This is a schematic diagram of the fault source location device provided in the embodiments of this application;
[0027] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0029] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0030] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0031] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0033] The fault source localization method provided in this application can be applied to electronic devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, desktop computers, servers, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.
[0034] To illustrate the technical solution of this application, specific embodiments are described below.
[0035] Please see Figure 1 , Figure 1 The flowchart illustrating the fault source localization method provided in this application is shown as an example and not a limitation. The method includes the following steps:
[0036] Step 101: Obtain the current alarm data of the distributed energy system.
[0037] In some embodiments, the current alarm data of the distributed energy system can be obtained in real time from the monitoring platform of the distributed energy system.
[0038] The current alarm data of the aforementioned distributed energy system may include a set of alarm events (which can be referred to as current alarm events) that have been monitored and reported by one or more devices in the distributed energy system in the recent period or within the current time window. Optionally, the current time window can be set according to actual needs or empirical values.
[0039] For example, the current time window can be a short-term window (e.g., 5-15 minutes). Based on the current alarm data collected within this short-term window, the immediate propagation of faults can be analyzed, such as the inverter's protection mechanism operating within 5-15 minutes after a battery pack failure. The current time window can also be a medium-term window (e.g., 1-6 hours). Based on the current alarm data collected within this medium-term window, the delayed impact of faults can be analyzed, such as the gradual decline in equipment performance after an abnormal temperature. Finally, the current time window can be a long-term window (e.g., 24 hours). Based on the current alarm data collected within this long-term window, periodic fault patterns can be analyzed, such as an increase in fault frequency during peak electricity consumption periods.
[0040] In this embodiment, the data used to locate the source of the fault is the alarm data acquired in real time (such as alarm data in the last hour), which does not require waiting for the batch collection to be completed, thus greatly improving the fault response speed.
[0041] In a distributed energy system, the devices can be of the same type or different types, and this application does not limit this. For example, the devices in a distributed energy system can be 30 inverters of the same type, or they can be multiple different types of devices such as inverters, solar panels, energy storage systems, charging piles, smart meters, gateways, and temperature sensors.
[0042] A device's current alarm event indicates that a fault has been detected in the device recently or within the current time window. Examples include battery pack failure, gateway communication anomalies, and temperature sensor malfunctions. In other words, the generation of an alarm event by a device indicates a fault in that device.
[0043] A device's current alarm event includes, but is not limited to, device identifier, fault type (e.g., over-temperature, communication interruption, over-voltage protection, etc.), alarm time, device physical location coordinates, device type, alarm level, and other information.
[0044] Step 102: Match the current alarm data with each association rule in the fault association rules.
[0045] The aforementioned fault association rules can include a number of predefined association rules. These fault association rules can be pre-stored in a rule base.
[0046] Each association rule may include fields such as rule identifier, rule type, association pattern description (e.g., "Device A fails → Device B fails, time window 5-15 minutes", meaning Device B fails within 5-15 minutes after Device A fails), confidence level, support level, detailed description, and applicable conditions. Confidence level represents the reliability of the corresponding association rule, and support level represents the generality of the corresponding association rule.
[0047] In this embodiment, by matching the current alarm data of the distributed energy system with each association rule in the fault association rules, it is possible to detect whether the overall pattern of the current alarm data of the distributed energy system conforms to the association pattern described by a certain association rule.
[0048] In some embodiments, before performing step 102, the current alarm data may be preliminarily filtered, cleaned and standardized.
[0049] In some embodiments, the fault association rules described above may include four types of association rules: time series association rules, spatial topology association rules, functional dependency association rules, and environmental factor association rules. Each of these four types of association rules may include multiple association rules. The association rules in the fault association rules described above are all the association rules included in these four types of association rules.
[0050] Time series association rules describe the statistical dependencies between device failures in terms of occurrence time. These rules typically include time-related information such as sliding time windows, specific time windows, and synchronization times. They can be used to analyze the temporal patterns of failure propagation (i.e., the chronological order of alarm events) and discover causal relationships and temporal dependencies between devices. Association patterns in this type of rule can include lag association patterns, synchronous association patterns, and periodic association patterns. A lag association pattern refers to the pattern where, after one device fails, another device fails within a specific time window. For example, when device A fails, device B has an 87% probability of experiencing a related failure within 5 to 15 minutes. Common scenarios include: after a battery pack failure, the inverter may activate its protection mechanism within 5 to 15 minutes; after a gateway communication anomaly, downstream devices may go offline within 5 to 15 minutes; after a temperature sensor malfunction, the cooling system may activate within 5 to 15 minutes.
[0051] Synchronous association mode refers to a mode in which multiple devices fail simultaneously within the same synchronization time. Common causes may include external power grid anomalies (such as voltage fluctuations, power outages, etc.), environmental factors (such as lightning strikes, high temperatures, floods, etc.), system-level configuration errors, malicious attacks, or network security incidents.
[0052] Optionally, the synchronization time can be set according to actual needs or experience. For example, the synchronization time can be set to 30 seconds.
[0053] Periodic correlation patterns refer to patterns in which faults recur at specific intervals. This is typically manifested as a significantly increased probability of equipment failure during specific time periods (e.g., peak electricity consumption periods). Common causes may include protection system activation due to changes in equipment load and power grid fluctuations.
[0054] Spatial topology association rules can describe fault propagation patterns based on the physical proximity of devices or logical network connections. These rules typically include spatially relevant information such as physical distance and proximity relationships, and can be used to analyze fault propagation among spatially adjacent devices. This is crucial for understanding the physical propagation mechanisms of faults and network cascading effects. Association patterns for these rules can include proximity device association patterns, network topology association patterns, and critical node association patterns.
[0055] The proximity device association mode refers to the fault propagation mode between devices whose physical distance is less than a preset distance threshold. Examples include: heat dissipation issues: high-temperature equipment affects neighboring equipment; vibration propagation: mechanical failures cause malfunctions in neighboring equipment; environmental sharing: sharing environmental threats (dust, humidity); and electromagnetic interference: high-voltage equipment affects low-voltage equipment. Optionally, a preset distance threshold can be set based on actual needs or empirical values. For example, a preset distance threshold of 50 meters.
[0056] Network topology association patterns can refer to the cascading propagation patterns of faults along network connections. For example, an upstream device failure causes downstream devices to lose their data source; a master control device malfunction causes subordinate devices to enter protection mode; a communication link interruption causes all devices on the link to go offline; a power supply failure causes devices on the power supply chain to shut down sequentially.
[0057] The critical node correlation pattern can refer to the ripple effect of a core device failure (i.e., a core device failure leading to a chain reaction of failures in surrounding devices). For example, a central switch failure causes communication interruption for all devices in the subnet; a main controller malfunction causes all controlled devices to malfunction; a power bus failure causes all connected devices to lose power.
[0058] Functional dependency association rules describe the fault dependencies between devices caused by logical functions, data flows, or control relationships. They focus on the logical functional relationships between devices, analyzing how data interaction, control signals, and functional cooperation form fault propagation paths. These association rules typically include function-related information such as master-slave relationships and functional dependencies, and can be used to analyze the fault dependencies between master and slave devices. Association patterns for this type of rule can include master-slave device association patterns, complementary device association patterns, and data dependency association patterns.
[0059] The master-slave device association mode can refer to a mode in which a failure of the master control device leads to a degradation of the function of the slave device. For example, a failure of the Battery Control Unit (BCU) causes the Battery Management Unit (BMU) to enter local protection mode; a gateway malfunction causes smart devices to lose remote control; an inverter failure causes the battery pack to stop charging and discharging; and a smart meter malfunction causes the energy management system to lose data.
[0060] Complementary equipment association mode can refer to fault modes where equipment depends on each other or fault response modes where equipment compensates for each other. For example, a fault in one inverter causes other inverters to increase their output; an abnormality in some battery packs causes normal battery packs to bear more load; an interruption in the main communication link causes the backup link to be automatically activated; a cooling system fault causes derating protection devices to operate.
[0061] Data dependency patterns can be fault propagation patterns caused by data flow or patterns where data source failures lead to malfunctions in data-consuming devices. For example, sensor failures can cause control systems to make incorrect decisions; database service anomalies can cause applications that depend on the data to be interrupted; and time synchronization service failures can cause timing discrepancies in distributed systems.
[0062] Environmental factor association rules describe the correlation between changes in performance indicators and the occurrence of equipment failures. They analyze the relationship between various environmental parameters and equipment failures, providing a scientific basis for preventative maintenance. These association rules typically include environmentally relevant information such as performance indicators and triggering conditions, and can be used to analyze the correlation between abnormal performance indicators and failures. Association patterns for these rules can include temperature association patterns, humidity association patterns, power grid quality association patterns, and air quality association patterns.
[0063] Temperature-related modes refer to equipment failure modes triggered by abnormal ambient temperatures. For example, an ambient temperature above 40°C causes the inverter to operate at a reduced rating; a battery temperature above 45°C leads to charging power limitations; a cabinet temperature above 50°C triggers forced cooling activation; and sustained high temperatures accelerate equipment lifespan degradation. Corresponding preventative measures may include improving ventilation and heat dissipation design, installing temperature monitoring systems, and setting temperature warning thresholds.
[0064] Humidity-related modes refer to equipment failure modes caused by high humidity environments. For example, humidity greater than 85% leads to decreased insulation performance of equipment; prolonged high humidity causes corrosion of metal components; humidity fluctuations cause condensation on circuit boards; and extreme humidity poses electrical safety hazards. Corresponding protective measures may include sealed protection design, dehumidification equipment configuration, and regular maintenance and inspection.
[0065] Power grid quality-related modes can refer to equipment failure modes caused by power grid anomalies. For example, voltage fluctuations greater than 10% may trigger equipment protection; frequency anomalies (e.g., 47–53 Hz) may cause synchronization equipment failure; excessive harmonic content may cause equipment overheating; and voltage sags may cause sensitive equipment to restart. Corresponding protection recommendations may include installing voltage stabilizers, configuring uninterruptible power supply (UPS) systems, and strengthening harmonic mitigation.
[0066] Air quality correlation patterns can refer to the patterns of equipment performance degradation caused by air pollutants. For example, dust accumulation reduces the heat dissipation efficiency of equipment; high particulate matter concentration causes filter clogging; corrosive gases damage electronic components; and poor air circulation leads to the formation of localized hotspots. Corresponding mitigation measures may include strengthening environmental control in the data center of distributed energy systems, regular cleaning and maintenance, and installing air purification equipment.
[0067] In some embodiments, the four types of association rules mentioned above can be traversed in the rule base to match the current alarm data with each type of association rule.
[0068] The matching of current alarm data with time-series association rules involves: traversing the time-series association rules in the database; for each association rule (e.g., "Device A failure → Device B failure, time window 5-15 minutes"), checking if there are alarm events for both Device A and Device B in the current alarm data, provided that Device A's alarm time precedes Device B's alarm time, and the time difference is within the time window specified by the rule (5-15 minutes). If a match is successful, Device A is likely the source of the failure, and the matching rule identifier, confidence level, support level, and other information are recorded. For example: in the current alarm data, Device X fails at 10:00, and Device Y fails at 10:08 (time difference 8 minutes), and the rule "Device X failure → Device Y failure, time window 5-15 minutes, confidence level 0.87" exists, then Device X is likely the source.
[0069] The matching of current alarm data with spatial topology association rules includes: traversing the spatial topology association rules in the rule base; for each association rule (e.g., "fault propagation of spatially adjacent devices (physical distance < 50 meters): device C fault → device D fault"), checking whether there are alarm events for devices C and D in the current alarms, and calculating whether the physical distance between the two is less than the preset distance threshold (e.g., 50 meters) specified by the association rule using the physical location coordinates of the devices. If the match is successful, device C is likely the source of the fault. For example: if devices M and N are both faulty in the current alarm data, and the distance between them is 30 meters, and there is a rule "fault propagation of spatially adjacent devices: device M fault → device N fault, confidence level 0.82", then device M is likely the source.
[0070] The matching of current alarm data with functional dependency association rules includes: traversing the functional dependency association rules in the rule base; for each association rule (e.g., "master device failure → slave device failure"), checking if there are any slave device alarm events in the current alarm data, and confirming the existence of a master-slave relationship through device topology information. If the match is successful, the master device is likely the source of the failure. For example, if slave device P fails in the current alarm data, and the rule "master device O failure → slave device P failure, confidence 0.90" exists, then master device O is likely the source of the failure.
[0071] The matching of current alarm data with environmental factor association rules involves: traversing the environmental factor association rules in the rule base; for each rule in the environmental factor association rules (e.g., "When performance indicators are abnormal (voltage fluctuation > 10%), related equipment will fail"), checking if there are any abnormal performance indicators (obtained through device monitoring data) and related device alarm events in the current alarm data. If the match is successful, the abnormal performance indicator may be the source of the fault. For example: if a 12% voltage fluctuation is detected in device Q, and device Q fails, and there is an association rule "Voltage fluctuation > 10% → Device Q fails, confidence level 0.85", then the abnormal voltage fluctuation may be the source of the fault.
[0072] Step 103: Based on the matched association rules, construct the first fault propagation graph of each device involved in the current alarm data.
[0073] In the first fault propagation graph, the nodes represent the devices involved in the current alarm data, the edges represent the fault propagation paths between devices, and the edge weights (i.e., path weights) represent the confidence of the matched association rules.
[0074] As an example rather than a limitation, if the association rules “Device A→Device B, confidence 0.87” and “Device B→Device C, confidence 0.82” are matched, then the fault propagation path “Device A→Device B→Device C” can be constructed, with a path weight of 0.87×0.82=0.7134.
[0075] It should be noted that the current alarm data of a distributed energy system includes not only the alarm devices listed in the current alarm data, but also potentially key devices that have not yet generated an alarm, inferred based on matching association rules. The aforementioned alarm devices can refer to those that have explicitly generated an alarm in the current alarm data of the distributed energy system. The aforementioned key devices that have not generated an alarm can refer to devices that are dependent on the alarm devices and whose failure may trigger an alarm from the alarm devices, but which themselves may not have generated an alarm at present.
[0076] In this embodiment, a first fault propagation graph of each device involved in the current alarm data is constructed based on the matched association rules. This avoids calculating the complete topology of the entire distributed energy system, reduces the size of the analysis model, and improves the efficiency of fault location. Furthermore, by including key devices that have not generated the current alarm in the first fault propagation graph from the matched association rules, the hidden fault source can be revealed, improving the coverage and accuracy of fault source location.
[0077] In some embodiments, a first fault propagation graph can be constructed based on the matched association rules and a multi-level fault propagation model. This model comprehensively considers device importance, historical fault data, and network topology, and can accurately predict the propagation path and impact range of faults, providing a key basis for the construction of the first fault propagation graph.
[0078] The aforementioned multi-level fault propagation model forms the knowledge foundation for fault source localization, providing the patterns and probabilistic information of fault propagation for streaming fault source tracing. It can be constructed based on multi-dimensional correlation features extracted from historical alarm data and mined correlation rules, describing the multi-level relationships of fault propagation between devices (time-series propagation, spatial topology propagation, functional dependency propagation, and environmental factor propagation).
[0079] Step 104: Based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, determine each candidate fault source and calculate the target probability of each candidate fault source as the fault source.
[0080] Among them, the target probability of each candidate fault source as a fault source represents the likelihood that each candidate fault source is a fault source.
[0081] The time order of the current alarm data can refer to the logical sequence determined by sorting all the current alarm events in the current alarm data according to the alarm time recorded in the current alarm event.
[0082] This embodiment achieves intelligent fusion of multi-source information by integrating the time sequence of current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules for multi-dimensional probabilistic reasoning, thereby improving the accuracy and robustness of fault source inference in complex distributed energy scenarios.
[0083] In this embodiment, by quantifying the probability of each candidate fault source as a fault source through target probability quantification, the fault source can be located quickly.
[0084] Step 105: Identify the candidate fault source with the highest target probability as the fault source of the current alarm data.
[0085] Since the target probability directly represents the likelihood that each candidate fault source is the fault source, the candidate fault source with the highest target probability can be determined as the fault source of the current alarm data.
[0086] In this embodiment, when a device in a distributed energy system fails, the current alarm data is acquired and matched with each association rule in the fault association rules to find possible fault propagation paths. Then, an artificial intelligence assistant is called to perform in-depth analysis to locate the fault source, which can achieve streaming location of the fault source.
[0087] In this embodiment, by intelligently matching the current alarm data with fault association rules and dynamically constructing a fault propagation graph focused on the current fault scenario, the time sequence of the current alarm data, the matched association rules, the constructed fault propagation graph, and the time series association rules can be integrated for multi-dimensional probabilistic reasoning. Based on the target probability of each candidate fault source, the fault source can be located, achieving millisecond-level fault source location capability, significantly shortening the fault investigation time, changing the traditional delayed response mode that relies on manual inspection and passive alarms, and realizing rapid location of the fault source.
[0088] In some embodiments of this application, before matching the current alarm data with the various association rules in the fault association rules, fault association rules can be generated based on historical alarm data. The generation method may include, for example... Figure 2 Steps 201 to 205 are shown.
[0089] Step 201: Obtain historical alarm data of the distributed energy system.
[0090] It should be noted that the devices in the historical alarm data can be of the same type or different types, and this application does not limit this.
[0091] For example, if historical alarm data includes 30 inverters of the same type, the fault correlation patterns among these 30 inverters can be analyzed based on the historical alarm data. For instance, it can be analyzed how, after one inverter fails, the probability of other inverters failing within a specific time window can be determined. This embodiment, through fault correlation analysis of similar equipment, helps to identify batch problems, environmental factors, or design flaws.
[0092] It should be noted that the historical alarm data mentioned above is typically alarm data from distributed energy systems within a certain time frame. For example, this time frame is the most recent 3 to 6 months, to ensure that there are sufficient historical alarm events in the data.
[0093] In some embodiments, after obtaining historical alarm data, the historical alarm data can be initially filtered, cleaned, and standardized.
[0094] Step 202: Extract multi-dimensional correlation features from historical alarm data to obtain the time series features, spatial correlation features, network topology features, and performance index features of each device pair in the historical alarm data.
[0095] In some embodiments, after extracting multidimensional correlation features from historical alarm data, these features can be passed to an AI assistant for deep correlation rule mining to generate fault correlation rules. The entire process adopts streaming processing, and the features are immediately passed to the AI assistant for analysis after feature extraction, thus realizing real-time fault correlation analysis.
[0096] Time series features can analyze the temporal patterns of fault occurrence, identifying lagging, synchronous, and periodic correlation patterns. Spatial correlation features analyze fault propagation patterns based on the physical location of equipment. Network topology features consider the logical connections between devices. Performance index features can correlate changes in equipment performance parameters with fault occurrence.
[0097] It should be noted that time-series features, spatial correlation features, and network topology features are characteristics between any two devices in historical alarm data, used to describe the correlation between device pairs. For example, 30 devices will generate 435 device pairs, and each device pair has time-series features, spatial correlation features, and network topology features. Performance metric features are characteristics of each device, describing the performance status of an individual device.
[0098] Step 203: Convert the numerical features in the time series features into rules, add preset fields, and generate time series association rules.
[0099] In some embodiments, an AI assistant can be used to transform the numerical features in the time series into understandable rules and add preset fields to generate time series association rules.
[0100] Step 204: Convert the physical distance in the spatial association features and the network connection relationship in the network topology features into rules, and add preset fields to generate spatial topology association rules and functional dependency association rules.
[0101] In some embodiments, an AI assistant can be used to convert physical distance in spatial association features and network connection relationships in network topology features into rules, and add preset fields to generate spatial topology association rules and functional dependency association rules.
[0102] Step 205: Convert the correlation between performance index anomalies and faults in the performance index features into rules, add preset fields, and generate environmental factor association rules.
[0103] In some embodiments, an AI assistant can be used to convert the correlation between performance index anomalies and faults in performance index features into rules, and add preset fields to generate environmental factor association rules.
[0104] The preset fields include rule identifier, rule type, confidence level, support level, detailed description, and applicable conditions. Confidence level represents the reliability of the corresponding association rule; the higher the confidence level, the more reliable the association rule. Support level represents the universality of the corresponding association rule; the higher the support level, the more common the association rule.
[0105] The support of a association rule represents the frequency with which the rule appears in historical alarm data. The formula is as follows: Support = (Number of fault events in historical alarm data that simultaneously meet the applicable conditions of the association rule) / (Total number of fault events). A fault event typically consists of one or more historical alarm events, describing the overall scenario of a system anomaly or fault.
[0106] For example, if there are 1,000 fault events in the historical alarm data, and 730 of these 1,000 fault events meet the applicable conditions of a certain association rule, then the support of that association rule is 0.73 (73%).
[0107] The confidence score of an association rule represents the conditional probability that the conclusion condition also occurs when the premise condition of the association rule occurs. The formula is as follows: Confidence score = (Number of times the conclusion condition also occurs when the premise condition of the association rule occurs) / (Total number of times the premise condition of the association rule occurs).
[0108] For example, for a time series association rule, if device A fails 1000 times in historical alarm data (i.e., the precondition), and device B fails 870 times within 5 to 15 minutes (i.e., the conclusion condition), then the confidence level of the time series association rule is 0.87 (87%).
[0109] In some embodiments, minimum confidence threshold and minimum support threshold can be set according to actual needs or empirical values. The minimum confidence threshold can be used to filter some association rules with low confidence, and the minimum support threshold can be used to filter some association rules with low support.
[0110] It should be noted that the aforementioned time series characteristics, spatial correlation characteristics, network topology characteristics, and performance indicator characteristics can be understood as data-level correlations. Time series correlation rules, spatial topology correlation rules, functional dependency correlation rules, and environmental factor correlation rules can be understood as knowledge-level summaries of patterns.
[0111] In some embodiments, when transforming time-series features, spatial correlation features, network topology features, and performance indicator features into corresponding types of association rules using an artificial intelligence model, these extracted features can first be formatted into natural language descriptions to construct prompt words containing the following requirements: a) The AI assistant is required to identify four types of association rules: time-series association rules (analyzing the temporal order of failures between devices), spatial topology association rules (analyzing the propagation of failures in spatially adjacent devices), functional dependency association rules (analyzing the failure dependency relationship between master and slave devices), and environmental factor association rules (analyzing the association between performance indicator anomalies and failures); b) Each rule must include: rule identifier, rule type, association pattern description (e.g., "Device A failure → Device B failure, time window 5-15 minutes"), confidence level, support level, detailed description, and applicable conditions; c) An example format is provided to guide the AI assistant to output structured JSON-formatted association rules. The formatted prompt words are then sent to the AI assistant, which performs in-depth analysis based on these features to identify failure association rules. Next, fault association rules are parsed from the streaming response of the AI assistant, including: a) extracting association patterns discovered during the thinking process; b) parsing the final structured rules; c) verifying that the confidence and support of the rules are within a reasonable range (0-1) and that the rule patterns meet the format requirements; d) deduplicating and merging the rules. Then, the rules are classified: the identified rules are classified into four categories according to type: a) Time series association rules: rules containing time-related information such as time windows and lag times; b) Spatial topology association rules: rules containing spatial-related information such as spatial distance and proximity relationships; c) Functional dependency association rules: rules containing function-related information such as master-slave relationships and functional dependencies; d) Environmental factor association rules: rules containing environmental-related information such as performance indicators and triggering conditions. The entire identification process uses streaming processing, real-time analysis, and rule generation to ensure the timeliness and accuracy of the rules.
[0112] In this embodiment, carefully designed prompts enable the AI assistant to focus on key dimensions, provide a standardized rule output format, and combine intermediate results and final outputs during the AI assistant's thinking process to obtain a more comprehensive set of association rules.
[0113] In some embodiments of this application, the extraction method of time series features for each device includes:
[0114] The historical alarm data is sorted by alarm time, and the sorted data is traversed using multiple sliding time windows of different scales.
[0115] For each historical alarm event during the traversal process, find other historical alarm events that occurred after its alarm time and before the current sliding time window ends, and obtain the search results;
[0116] Based on the search results, the frequency of historical alarm events of the corresponding device pair appearing in pairs within the current sliding time window is statistically analyzed, and the correlation strength of the historical alarm events of the device pair is calculated based on the frequency.
[0117] Based on the time difference of alarm times of the device pair, determine the association mode type of the device pair;
[0118] Based on the device identifier, association strength, time difference, association pattern type, and current sliding time window of the device pair, the time series features of the device pair are generated.
[0119] The above sorting of historical alarm data by alarm time can refer to sorting all historical alarm events in the historical alarm data in order from earliest to latest alarm time.
[0120] Optionally, the scale of the sliding time window can be set according to actual needs or empirical values. For example, multiple sliding time windows such as 60 seconds, 120 seconds, 240 seconds, ..., 3600 seconds can be used to traverse the sorted data to capture correlation patterns at different time scales.
[0121] For each historical alarm event, the search results for that historical alarm event in the current sliding window can include a list of all other historical alarm events found in the sorted data after the alarm time of that historical alarm event and before the end of the current sliding window. The device that generated the historical alarm event and the device that generated each other historical alarm event in the search results form a device pair. The frequency of each device pair appearing in pairs within the current sliding window is determined by dividing the number of times the device pair appears in pairs within the current sliding window by the total number of times the device that generated the historical alarm event appears.
[0122] In some embodiments, the frequency with which historical alarm events occur in pairs within the current sliding time window can be determined as the correlation strength of the device with historical alarm events within the current sliding window.
[0123] In some embodiments, for each historical alarm event and a corresponding other historical alarm event, if the alarm time in the historical alarm event is earlier than the alarm time of the other historical alarm event, and the time difference between these two alarm times is within a specific time window, then the association mode type between the device that generated the historical alarm event and the device that generated the other historical alarm event can be determined to be a lag association mode. If the time difference between these two alarm times is within a synchronization time, then the association mode type between the device that generated the historical alarm event and the device that generated the other historical alarm event can be determined to be a synchronous association mode.
[0124] It should be noted that the correlation strength, time difference, and current sliding time window of the above-mentioned device pairs are numerical features of the time series characteristics.
[0125] In some embodiments of this application, the method for extracting spatial association features for each device includes:
[0126] Calculate the physical distance between the equipment pairs based on the physical location coordinates of each equipment in the equipment pair;
[0127] If the physical distance between the devices is less than a preset distance threshold, the frequency of fault co-occurrence between the devices is counted, and the spatial association features of the devices are generated based on the physical distance between the devices and the frequency of fault co-occurrence.
[0128] If the physical distance between the device pairs is greater than or equal to a preset distance threshold, then the spatial association features of the device pairs are generated based on the physical distance between them.
[0129] Optionally, the physical location coordinates of the device can be represented by its Global Positioning System (GPS) coordinates or the location coordinates of the computer room where the device is located.
[0130] In this embodiment, if the physical distance between the device pairs is less than a preset distance threshold, the device pairs can be determined to be neighboring device pairs. The co-occurrence frequency of device pair failures can refer to the frequency at which two devices in the device pair fail simultaneously or sequentially within a specified time window.
[0131] In some embodiments of this application, the method for extracting network topology features for each device pair includes:
[0132] Obtain the tree topology of each device in the historical alarm data; the tree topology represents the physical connection, logical dependency and data flow relationship between devices;
[0133] Based on a tree topology, a device relationship graph is constructed; nodes in the device relationship graph represent devices, and edges represent the network connection relationships between devices.
[0134] In the device relationship diagram, for each device pair, analyze the fault propagation path along the network connection;
[0135] If there is a core device in the equipment pair, the scope of the failure impact of the core device is identified based on the tree topology structure.
[0136] Based on the device identifier, network connection relationship, fault propagation path and fault impact range of the device pair, the network topology characteristics of the device pair are generated.
[0137] The network connection relationships between the aforementioned devices can include gateway-device relationships, master-slave relationships, data flow relationships, functional dependencies, etc.
[0138] The scope of impact of the aforementioned core equipment failure includes the equipment affected when the core equipment fails.
[0139] In some embodiments of this application, the extraction methods for the performance metric features of each device include:
[0140] Extract performance metrics from real-time monitoring data of the equipment;
[0141] Analyze the trends in the equipment's performance indicators before and after the failure;
[0142] Determine whether the equipment's performance indicators are abnormal based on the changing trends;
[0143] If the performance indicators are abnormal, the correlation between the abnormal performance indicators and the fault is calculated, and the correlation between the abnormal performance indicators and the fault is determined as the performance indicator characteristics of the equipment.
[0144] The aforementioned performance indicators include, but are not limited to, voltage, current, temperature, power, and communication quality.
[0145] In some embodiments, the total number of performance metric anomalies within a target time window and the number of failures that occurred after the performance metric anomalies can be counted. The value obtained by dividing the number of failures that occurred after the performance metric anomalies by the total number of performance metric anomalies is determined as the correlation between performance metric anomalies and failures. Of course, it is understood that the correlation between performance metric anomalies and failures can also be calculated in other ways, and this application does not limit this.
[0146] Optionally, a target time window can be set based on actual needs or experience.
[0147] In some embodiments of this application, obtaining the current alarm data of the distributed energy system includes:
[0148] If a critical device failure is detected in the distributed energy system, or if multiple devices fail within a preset time, the current alarm data is obtained.
[0149] Optionally, the aforementioned key equipment can be core equipment or other equipment (such as equipment with alarm levels higher than the preset level).
[0150] Optionally, a preset time can be set based on actual needs or experience. For example, the preset time is 5 minutes.
[0151] In this embodiment, the current alarm data is only acquired when critical equipment fails or multiple equipment fails within a short period of time. This avoids the need for real-time processing of massive alarm data under normal conditions, thereby reducing real-time processing pressure, saving storage costs, reducing network transmission, and conserving computing resources.
[0152] In some embodiments of this application, three analysis methods—Bayesian inference, AI assistant analysis, and time series analysis—are used to calculate the target probability. Specifically, based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, each candidate fault source is determined, and the target probability of each candidate fault source as a fault source is calculated, including:
[0153] Candidate fault sources are inferred based on the current alarm data and the first fault propagation graph. Based on the first fault propagation graph, the likelihood probability of the inferred candidate fault source being the fault source is calculated. The likelihood probability is the probability of observing the current alarm data when the inferred candidate fault source is the fault source.
[0154] Based on the likelihood probability of the inferred candidate fault source as the fault source and the preset prior probability, the posterior probability of the inferred candidate fault source as the fault source is calculated.
[0155] Format the current alarm data, the matched association rules, and the first fault propagation graph into prompt words;
[0156] Based on the prompt words, the AI assistant is guided to explore different candidate fault sources and determine the AI probability of each candidate fault source as the fault source; the AI probability is the probability of the corresponding candidate fault source determined by the AI assistant as the fault source.
[0157] Based on the time sequence of the current alarm data, time series association rules are used to analyze the time pattern of fault propagation in order to determine the association patterns that match the time patterns in the time series association rules; the time pattern represents the temporal regularity of fault propagation.
[0158] Candidate fault sources are identified based on the association pattern, and the time series probability of the identified candidate fault sources is determined based on the confidence of the association rule corresponding to the association pattern.
[0159] For the same candidate fault source, the posterior probability, artificial intelligence probability, and time series probability of the candidate fault source as a fault source are weighted and fused to obtain the target probability of the candidate fault source as a fault source.
[0160] In some embodiments, in the absence of observational evidence, the probability (i.e., prior probability) of each device potentially being a source of failure can be pre-defined by considering factors such as device importance (critical devices, such as the main controller, have a higher prior probability), historical failure rate (devices that frequently fail have a higher prior probability), and device type (certain device types, such as battery packs, are more likely to be sources of failure). For example, a prior probability of 0.30 for device A means that, without observing any alarms, device A has a 30% probability of being a source of failure. Prior probabilities reflect historical experience and device characteristics, providing an initial estimate for Bayesian inference.
[0161] The likelihood probability of a candidate fault source being the fault source can be defined as the probability that, given a candidate fault source is indeed the fault source (hypothesis), the current alarm data (observational evidence) is observed. For example, assuming device A is the fault source, a likelihood probability of 0.85 means that if device A is indeed the fault source, then the probability of observing the current alarm data is 85%. Likelihood probability reflects the degree of matching between the observed evidence and the hypothesis. If the observed alarm data has a strong propagation relationship with the hypothetical fault source, the likelihood probability is high.
[0162] In some embodiments, for each node in the first fault propagation graph, a downstream influence set for that node can be calculated. If all or most of the alarm devices in the current alarm data are within the downstream influence set of this node, then the device represented by this node can be identified as a candidate fault source. The downstream influence set of this node can include devices represented by all nodes reachable from this node in the first fault propagation graph.
[0163] In some embodiments, Bayesian inference can be used to multiply the likelihood probability of the inferred candidate fault source as a fault source by a preset prior probability to obtain the posterior probability of the inferred candidate fault source as a fault source.
[0164] The aforementioned association pattern that matches the current alarm data can refer to the association pattern that matches the current alarm data among the various association patterns of the time series association rules.
[0165] In some embodiments, the source device corresponding to the time pattern (i.e., the alarm device with the earliest alarm time) can be identified as a candidate fault source, and the confidence level of the association rule corresponding to the association pattern matching the time pattern can be determined as the time series probability of the candidate fault source. It should be noted that an association pattern can correspond to multiple association rules, so the average confidence level of these multiple association rules can be determined as the time series probability of the candidate fault source.
[0166] For example, if the current alarm data is in the time sequence of device A (10:00), device B (10:08), and device C (10:15), and the analysis shows that the fault propagation time pattern is device A (10:00) → device B (10:08) → device C (10:15), and the time difference meets the time difference requirement of the lag pattern in the time series association rule, then device A can be identified as a candidate fault source.
[0167] In some embodiments, for the same candidate fault source, the posterior probability, artificial intelligence probability, and time series probability of the candidate fault source as a fault source can be weighted and summed based on the weights of Bayesian inference, the weights of AI assistant analysis, and the weights of time series analysis to obtain the target probability of the candidate fault source as a fault source.
[0168] Optionally, the weights of Bayesian inference, AI assistant analysis, and time series analysis can be set according to actual needs or experience.
[0169] Since the posterior probability is derived from Bayesian inference and the AI probability from an AI assistant, both of which consider the first fault propagation graph and association patterns, their reliability is relatively high. Therefore, the weights of the posterior probability and the AI probability can be set relatively large. The time series probability is derived from time series association rules, which primarily focus on time patterns and have relatively simple information. Therefore, the weights of the time series probability can be set relatively small. For example, the weights of the posterior probability, the AI probability, and the time series probability can be set to 0.4, 0.4, and 0.2, respectively.
[0170] In some embodiments, after obtaining the target probability of all candidate fault sources as fault sources, the sum of the target probabilities of all candidate fault sources as fault sources can be calculated. Then, the target probability of each candidate fault source as a fault source is divided by the sum to ensure that the sum of all target probabilities is 1.
[0171] For example, if the target probabilities of equipment A, equipment B, and equipment C as failure probabilities are 0.46, 0.36, and 0.20 respectively, with a total of 1.02, then after normalization, they are 0.46 / 1.02 = 0.451, 0.36 / 1.02 = 0.353, and 0.20 / 1.02 = 0.196 respectively. The source of the failure is located based on the normalized probabilities.
[0172] It should be noted that if the probability of a candidate fault source being a fault source includes only one or two of the posterior probability, artificial intelligence probability, and time series probability, then the probability not included can be set to 0, and then weighted fusion can be performed on this basis.
[0173] It should be noted that the number of candidate fault sources inferred can be one or more. When there are multiple candidate fault sources inferred, the posterior probabilities of multiple candidate fault sources as fault sources can be normalized to ensure that the sum of the posterior probabilities of multiple candidate fault sources as fault sources is 1.
[0174] In this embodiment, by weighted fusion of the posterior probability, artificial intelligence probability, and time series probability of the candidate fault source as the fault source, the results of multiple information sources can be integrated, and the positioning results of different analysis methods can be comprehensively considered, thereby improving the accuracy and reliability of fault source positioning.
[0175] In some embodiments of this application, based on the first fault propagation graph, the likelihood probability of the inferred candidate fault source being the fault source is calculated, including:
[0176] From the first fault propagation diagram, find the fault propagation path from the inferred candidate fault source to each alarm device in the current alarm data;
[0177] The path probability of a fault propagation path is obtained by multiplying the weights of all edges in any fault propagation path; the weights of the edges represent the confidence of the corresponding association rule.
[0178] Obtain the total delay time of the fault propagation path;
[0179] Based on the total delay time of the fault propagation path and the alarm time of the corresponding alarm device, the time factor of the fault propagation path is calculated; the time factor reflects the degree of matching between the alarm time of the current alarm data and the total delay of the fault propagation path.
[0180] The likelihood probability of the corresponding alarm device is obtained by multiplying the path probability of the fault propagation path by the time factor.
[0181] Calculate the product of the likelihood probabilities of each alarm device to obtain the likelihood probability of the inferred candidate fault source as the fault source.
[0182] The total delay time of the fault propagation path from the aforementioned candidate fault source to a certain alarm device in the current alarm data can refer to the total time required to reach the alarm device along the fault propagation path from the candidate fault source.
[0183] In some embodiments, for any fault propagation path, the actual delay time of the fault propagation path can be calculated based on the alarm time of the candidate fault source and the alarm device corresponding to the fault propagation path (i.e. the alarm device that the candidate fault source can reach along the fault propagation path). Based on the actual delay time and total delay time of the fault propagation path, the time factor corresponding to the fault propagation path is calculated in combination with the Gaussian kernel function.
[0184] In this embodiment, the likelihood probability corresponding to the fault propagation path is calculated by using a time factor, taking into account the matching degree between alarm time and propagation delay, which can improve the accuracy of fault source location.
[0185] It should be noted that if no fault propagation path from the candidate fault source to a certain alarm device is found in the first fault propagation graph, the likelihood probability of the alarm device can be determined to a small preset value (e.g., 0.01).
[0186] For example, devices A and E are inferred candidate fault sources. Other devices besides A and E are also inferred candidate fault sources, but they are not listed here. Devices B, C, and D are alarm devices. a) Calculation process of the posterior probability of device A as a fault source: Prior probability: The prior probability of device A as a fault source is 0.30; Likelihood probability calculation: Find the fault propagation path from device A to device B: Find the fault propagation path "device A → device B", path probability = 0.87 (edge weight), time factor = 0.95, likelihood probability of device B = 0.87 × 0.95 = 0.8265; Find the fault propagation path from device A to device C: Find the fault propagation path "device A → device B → device B". "Backup C", path probability = 0.87 × 0.82 = 0.7134, time factor = 0.90, likelihood probability of device C = 0.7134 × 0.90 = 0.6421; Search for propagation path from device A to device D: no propagation path found, likelihood probability of device D = 0.01; likelihood probability of device A as the source of failure = 0.8265 × 0.6421 × 0.01 = 0.0053; posterior probability of device A as the source of failure = 0.30 × 0.0053 = 0.00159. b) Calculation of the posterior probability of device E as the source of the fault: The prior probability of device E as the source of the fault is 0.20; Likelihood probability calculation: After searching for propagation paths from device E to devices B, C, and D respectively, none were found. The likelihood probability of device E as the source of the fault = 0.01 × 0.01 × 0.01 = 0.000001; Posterior probability: The posterior probability of device E as the source of the fault = 0.20 × 0.000001 = 0.0000002. c) Normalization: Assuming that the sum of the posterior probabilities of all candidate fault sources as the source of the fault is 0.5, then the normalized posterior probability of device A as the source of the fault = 0.00159 / 0.5 = 0.00318, and the normalized posterior probability of device E as the source of the fault = 0.0000002 / 0.5 = 0.0000004.
[0187] In some embodiments of this application, before guiding the AI assistant to explore different candidate sources of failure based on prompt words, the method further includes:
[0188] Various types of evidence are embedded in the prompts, which guide the AI assistant to adjust the evidence weights, filter fault scenarios, and explore different candidate fault sources. The evidence weight is the degree of influence of the corresponding evidence on the fault source location during the fault source analysis process. Various types of evidence include current alarm data, performance indicators of each device involved in the current alarm data, log information, and network topology.
[0189] Based on prompts, the AI assistant is guided to explore different candidate fault sources and determine the AI probability of each candidate fault source being the fault source, including:
[0190] Based on the prompts, the AI assistant is guided to adjust the evidence weights of various types of evidence and determine the matching fault scenarios.
[0191] Based on the matched fault scenarios and various types of evidence, explore different candidate fault sources;
[0192] Based on the adjusted evidence weights, the artificial intelligence probability of different candidate fault sources as fault sources is determined.
[0193] In this embodiment, guiding the AI assistant to use evidence weighting and fault hypothesis through prompts can improve the accuracy and interpretability of the analysis.
[0194] The aforementioned matched fault scenarios can refer to the fault types or system states that can explain these pieces of evidence after comprehensive inference. For example, when temperature sensor alarms, fan malfunction logs, and related equipment overheating performance indicators occur simultaneously, the inferred fault scenario might be a heat dissipation failure scenario.
[0195] In this embodiment, by adjusting the evidence weights, the AI assistant can place greater emphasis on reliable and relevant evidence. Current alarm events are the most direct evidence of faults, and therefore have a high evidence weight. Performance metrics, including voltage, current, temperature, and power, reflect the device's operating status and have a medium evidence weight. Log information, including device operation logs and error logs, provides contextual information about the fault's occurrence and has a medium evidence weight. Network topology, including device connection relationships and master-slave relationships, reflects the dependencies between devices and has a high evidence weight.
[0196] For example, suppose there are three alarm events: device A (alarm level: emergency), device B (alarm level: normal), and device C (alarm level: warning); suppose the performance indicator shows that the voltage of device A fluctuates by 12% (abnormal), then the evidence weight of the performance indicator increases to 0.9; when analyzing, the artificial intelligence assistant will pay more attention to the alarm and performance indicator of device A, and consider device A to be more likely to be the source of the fault.
[0197] In some embodiments, when a certain type of evidence is more reliable or more relevant, its evidence weight can be increased, making the AI assistant pay more attention to that type of evidence; when a certain type of evidence is unreliable or irrelevant, its evidence weight can be reduced, reducing its impact on the analysis results; for example, if the alarm level is "urgent", the evidence weight of the current alarm event is higher; if the performance indicator is abnormal (such as voltage fluctuation >10%), the evidence weight of the performance indicator is higher.
[0198] In some embodiments, the above evidence weights can be adjusted as follows: the importance of various types of evidence can be clearly indicated in the prompt words, such as "Please pay special attention to alarm events with an emergency alarm level" or "Prioritize devices with abnormal performance indicators"; the evidence weights can be adjusted in multi-source information fusion by adjusting the weights of Bayesian inference, AI assistant analysis, and time series analysis (e.g., 0.4, 0.4, 0.2).
[0199] In some embodiments, users can embed various types of evidence into prompts through an interactive fault analysis interface. This interactive interface is a high-level interaction layer of the system, allowing users to directly participate in the fault analysis process. By adjusting evidence weights, filtering fault scenarios, and exploring different fault hypotheses, users can obtain analysis results that better reflect the actual situation. This human-computer collaborative analysis mode significantly improves the flexibility and accuracy of fault diagnosis.
[0200] The aforementioned fault hypotheses refer to multiple possible assumptions about the source of the fault proposed by the AI assistant during the fault source analysis process. Each hypothesis corresponds to a possible fault source device and fault propagation path. By exploring and evaluating different hypotheses, the most likely fault source can be found.
[0201] The generation of hypotheses includes: based on the current alarm and the first fault propagation graph, the AI assistant generates multiple fault hypotheses; for example: Hypothesis 1: Device A is the source of the fault, and the fault propagates to Device B and Device C through time series; Hypothesis 2: Device B is the source of the fault, and the fault propagates to Device C through spatial topology; Hypothesis 3: Device C is the source of the fault, but the faults of Device A and Device B occur independently.
[0202] Hypothesis exploration: The AI assistant evaluates each hypothesis and calculates its probability (confidence level); by comparing the confidence levels of different hypotheses, the most likely hypothesis is selected as the final result; for example: if the confidence level of hypothesis 1 is 0.85, the confidence level of hypothesis 2 is 0.60, and the confidence level of hypothesis 3 is 0.30, then hypothesis 1 is selected as the final result.
[0203] Hypothesis exploration is implemented by prompting the AI assistant to "explore different fault hypotheses" in the prompt, such as "Please consider multiple possible fault sources and evaluate the probability of each hypothesis"; the AI's thinking process is displayed in real time in the streaming response, including "Exploring hypothesis 1...", "Evaluating hypothesis 2...", etc.; different hypotheses explored by the AI are displayed through the candidate fault source and excluded fault source fields.
[0204] Examples of fault hypotheses: Hypothesis 1: Device A is the source of the fault (confidence: 0.85): Evidence: Device A has the earliest alarm time (10:00), while Device B and Device C have alarm times of 10:08 and 10:15 respectively; Fault propagation path: Device A → Device B → Device C; Supporting rule: Time series association rule "Device A fault → Device B fault, time window 5-15 minutes, confidence 0.87"; Hypothesis 2: Device B is the source of the fault (confidence: 0.60): Evidence: Device B Device B is spatially adjacent to Device C (30 meters away); Fault propagation path: Device B → Device C; Supporting rule: Spatial topology association rule "fault propagation of spatially adjacent devices, confidence level 0.82"; Hypothesis 3: Device C is the source of the fault (confidence level: 0.30); Evidence: Device C has the highest alarm level (urgent); Fault propagation path: None (the faults of Device A and Device B occurred independently); Supporting rule: No explicit association rule to support it; After evaluation by the AI assistant, Hypothesis 1 is selected as the final result.
[0205] It's important to note that evidence weighting influences hypothesis generation: different types of evidence have different weights, and evidence with higher weighting is more likely to generate hypotheses. For example, if the evidence weight of an alarm event is high, a hypothesis based on the alarm time sequence (time series propagation) is more likely to be considered. Evidence weighting also affects hypothesis evaluation: when evaluating the confidence level of each hypothesis, evidence with higher weighting contributes more to the confidence level. For example, if hypothesis 1 is supported by high-weighted alarm events, its confidence level is high. Hypothesis exploration helps adjust evidence weighting: by exploring different hypotheses, AI assistants can discover which evidence is more important and thus adjust the evidence weighting accordingly. For example, if a hypothesis based on performance metrics has high confidence, the evidence weighting of performance metrics is increased.
[0206] In some embodiments of this application, after determining the candidate fault source with the highest target probability as the fault source of the current alarm data, the method further includes:
[0207] The number of accurate predictions for each association rule and the total number of predictions are calculated in the statistical fault association rules.
[0208] Divide the number of accurate predictions by the total number of predictions to obtain the actual accuracy of the association rule.
[0209] Divide the total number of predictions by the first preset value to obtain the target value;
[0210] The minimum value between the target value and the second preset value is determined as the adjustment weight; the second preset value is less than the first preset value.
[0211] Based on the adjusted weights, the confidence of the association rule, and the actual accuracy, the adjusted confidence of the association rule is calculated as follows: Adjusted confidence = Confidence of association rule × (1 - Adjusted weights) + Actual accuracy × Adjusted weights.
[0212] Update the confidence level of the association rule to the adjusted confidence level.
[0213] In this embodiment, after acquiring each current alarm data and determining the fault source of the acquired current alarm data, the accuracy of the fault source predicted based on the matched association rules can be determined based on the actual fault source. If the fault source predicted based on the matched association rules is the same as the actual fault source, the number of accurate predictions by the matched association rules is incremented by 1, and the total number of predictions by the matched association rules is incremented by 1. If the fault source predicted based on the matched association rules is different from the actual fault source, the number of incorrect predictions by the matched association rules is incremented by 1, and the total number of predictions by the matched association rules is incremented by 1. Based on this, the number of accurate predictions for each association rule in the fault association rules and the total number of predictions can be statistically obtained.
[0214] Optionally, a first preset value and a second preset value can be set according to actual needs or experience. The first preset value is an integer greater than 1, and the second preset value is a positive number less than 1. For example, the first preset value is 100, and the second preset value is 0.8.
[0215] The formula for calculating the above adjustment weights is as follows:
[0216] Adjust weight = min(total number of predictions / first preset value, second preset value)
[0217] As the total number of predictions increases, the adjustment weight gradually increases, but the maximum value does not exceed a second preset value. This ensures that a certain proportion of the initial confidence level (i.e., the confidence level before adjustment) is always maintained, preventing drastic changes in the association rules due to the influence of a small number of outliers. The adjustment weight determines the proportion of the initial confidence level and the actual accuracy when adjusting the confidence level. The larger the adjustment weight, the greater the impact on the actual accuracy; the smaller the adjustment weight, the greater the impact on the initial confidence level. In this way, the association rules can be ensured to have a sufficient adaptation period, while preventing outliers from excessively affecting the rule adjustment. Each current alarm data obtained serves as a validation sample.
[0218] In this embodiment, as the number of validation samples increases, the adjustment weight gradually increases, and the system gradually shifts from relying on the initial confidence level to relying on the actual accuracy. This ensures that the system continuously learns and optimizes during operation, achieving incremental learning.
[0219] In some embodiments, an AI assistant can calculate the initial confidence level of a correlation rule based on historical alarm data. This initial confidence level reflects the reliability of the correlation rule within the historical alarm data. The AI assistant can dynamically adjust the confidence level based on the actual performance of the correlation rule. The adjusted confidence level reflects the reliability of the correlation rule in practical applications. Using the adjusted confidence level for subsequent fault source localization can make the localization results more accurate and reliable.
[0220] In this embodiment, by dynamically learning and adjusting the adjustment weights and confidence levels of fault association rules, the system can continuously adapt to changing operating environments and equipment states.
[0221] In this embodiment, adjusting the weights and confidence levels together constitutes the adaptive learning mechanism of the association rules, which can continuously optimize the fault association rules, improve the quality of the rule base, adapt to various dynamic scenarios such as equipment aging, environmental changes and network topology adjustments, and significantly improve the accuracy and reliability of fault location.
[0222] like Figure 3 The diagram shown is an example of continuous learning and optimization of the system provided in this application. Figure 3 As shown, the rule base can be initialized based on historical alarm data. This includes mining association rules through an AI assistant to obtain four types of association rules, and setting the initial confidence and support of each association rule; constructing a multi-level fault propagation model; locating the fault source through multi-source information fusion based on the rule base and the multi-level fault propagation model to realize the application of fault source location; generating verification samples based on the location results; calculating the actual accuracy of each association rule based on the verification samples, and updating the confidence based on the actual accuracy, adjusted weights, and confidence; updating the rule base and the multi-level fault propagation model based on the updated confidence (e.g., updating the edge weights), and applying the updated rule base and the multi-level fault propagation model to the next fault source location, realizing continuous learning and optimization of the system.
[0223] In some embodiments of this application, before matching the current alarm data with each association rule in the fault association rules, the method further includes:
[0224] From the fault association rules, candidate rules are selected; each candidate rule is an association rule whose adjusted confidence is greater than or equal to the confidence threshold and whose support is greater than or equal to the support threshold.
[0225] The candidate rules are sorted in descending order of their adjusted confidence levels.
[0226] Match the current alarm data with each association rule in the fault association rules, including:
[0227] Match the current alarm data with the sorted candidate rules.
[0228] Optionally, confidence and support thresholds can be set based on actual needs or empirical values. For example, the confidence threshold can be 0.5 and the support threshold can be 0.1.
[0229] In this embodiment, retaining association rules with support greater than or equal to the support threshold can filter noise and avoid misleading rules. Retaining association rules with confidence greater than or equal to the confidence threshold can filter out rules with low reliability and avoid misleading rules.
[0230] In some embodiments of this application, after calculating the target probability of each candidate fault source as a fault source, the method further includes:
[0231] The candidate fault sources are sorted in descending order of target probability, and the top N candidate fault sources (excluding the first one) are determined as alternative sources, where N is an integer greater than 1.
[0232] To visualize fault location, complex fault analysis results are presented to users in an intuitive and interactive manner. In some embodiments of this application, after determining the candidate fault source with the highest target probability as the fault source of the current alarm data, the method further includes:
[0233] Based on the multi-level fault propagation model, we can find the affected devices from the fault source and the fault propagation paths to each affected device. The multi-level fault propagation model is used to describe the laws and probabilities of fault propagation among devices in a distributed energy system.
[0234] Based on the fault source, the affected devices from the fault source, and the fault propagation path to each affected device, a second fault propagation diagram is drawn, and the fault propagation process is displayed step by step using animation effects.
[0235] The aforementioned fault propagation process can refer to the real-time display of the complete path and time sequence of a fault propagating from its source to the affected device through visualization methods such as animation and charts on the user interface (UI).
[0236] In some embodiments, a real-time fault propagation visualization component can be used to display the fault propagation process. This component can dynamically display the fault propagation path by extending the message renderer. It supports interactive operation, allowing users to explore different fault scenarios and propagation paths, thus improving the efficiency and accuracy of fault analysis.
[0237] In some embodiments, at least one of the following information is displayed on the fault propagation path of the second fault propagation graph: the delay time of the fault propagation path, the propagation type, and the confidence level of the corresponding association rule.
[0238] In some embodiments, the dynamically displayed content may include the AI assistant's analysis process, fault propagation path, propagation time information, propagation type, and confidence level, etc., and its specific implementation is as follows:
[0239] The analysis process of the AI assistant is displayed in real time: the analysis results of the AI assistant are acquired in real time, including the analysis ideas and the final location results; the analysis steps of the AI assistant are displayed on the UI in real time, such as "Analyzing device A...", "Discovering the propagation path from device A to device B...", "Calculating the propagation probability...", etc.
[0240] Dynamically display the fault propagation path: Taking device A→device B→device C as an example, draw a fault propagation diagram on the UI, use arrows to indicate the propagation direction, and use different colors to indicate the propagation probability (red for high probability, yellow for low probability); use animation effects to gradually display the fault propagation process in chronological order, such as first showing the fault of device A (flashing red), then showing the arrow from A to B (animation effect), then showing the fault of device B (flashing red), and so on.
[0241] Display propagation time information: Mark the propagation delay time on each side of the fault propagation path, such as "5-15 minutes", to help users understand the time pattern of fault propagation; mark the fault occurrence time (i.e., alarm event) of each device on the time axis, such as "Device A:10:00", "Device B:10:08", "Device C:10:15".
[0242] Display propagation type and confidence level: Mark the propagation type on the fault propagation path, such as "time series propagation", "spatial topology propagation", etc.; mark the propagation probability (i.e. the confidence level of the association rule) on each edge, such as "87%", to help users understand the reliability of the propagation relationship.
[0243] In some embodiments, the design of the real-time fault propagation visualization component includes fault propagation graph design, timeline design, analysis process panel design, and results panel design. Its specific implementation is as follows:
[0244] Fault propagation graph: Draw device nodes and propagation edges using a graphics library (such as Flutter's CustomPaint or a third-party charting library); nodes represent devices, using circles or rectangles, and labeled with the device name and status (normal / fault); edges represent propagation relationships, using arrows to indicate direction, and using color and thickness to indicate propagation probability; supports interactive functions such as zooming, dragging, and clicking on nodes to view details.
[0245] Time axis: The horizontal axis represents time, and the vertical axis represents equipment; the time point of failure of each equipment is marked on the time axis; the connection is used to represent the fault propagation relationship, and the slope of the connection represents the propagation delay time; the time axis can be dragged to view the fault propagation situation in different time periods.
[0246] Analysis Process Panel: Displays the AI assistant's analysis steps and thought processes in real time; shows analysis progress in list or card format, such as "Step 1: Get Current Alarms", "Step 2: Match Association Rules", "Step 3: Calculate Probability", etc.; supports expanding / collapsed details.
[0247] Results panel: Displays the final fault location results; presented in card or table format, and supports functions such as copying and exporting.
[0248] In some embodiments, the above-described dynamic display process is as follows:
[0249] Step 1: The user clicks the "Fault Source Analysis" button on the alarm details page;
[0250] Step 2: Invoke streaming fault source tracing to begin analysis;
[0251] Step 3: The UI displays the AI assistant's analysis process in real time, including steps such as "Analyzing..." and "Detecting fault propagation path...";
[0252] Step 4: After the AI assistant completes the analysis, it obtains the source of the fault, the affected devices reached by the fault source, and the fault propagation path to each affected device;
[0253] Step 5: Draw a fault propagation diagram based on the fault source, each affected device, and the fault propagation path to each affected device;
[0254] Step 6: Users can click on nodes to view device details, click to view propagation rule details, and drag the timeline to view the propagation situation in different time periods.
[0255] In this embodiment, through the aforementioned dynamic display, users can intuitively understand the fault propagation path and time sequence via graphical representation, which is easier to understand than a plain text description. Through streaming processing, users can see the AI assistant's analysis process in real time, rather than waiting for all analyses to complete before seeing the results, thus improving the user experience. Users can gain a deeper understanding of the details of fault propagation through interactive operations such as clicking and dragging, such as viewing device details and propagation rule details. By displaying information such as propagation type, confidence level, and supporting evidence, users can understand why the system considers a certain device to be the source of the fault, making the results explainable.
[0256] like Figure 4The diagram shown is a flowchart of the rapid fault source location process provided in an embodiment of this application. Figure 4 As shown, the electronic device can acquire the current alarm data and match it with four types of association rules. The matching results are then summarized, and a first fault propagation graph is constructed based on the matching results. Subsequently, three analysis methods—AI assistant analysis, Bayesian inference, and time series analysis—are used in parallel to search for candidate fault sources. The target probabilities of the found candidate fault sources are used to fuse multi-source information to generate a location result (i.e., the fault source of the current alarm data is found). The second fault propagation graph, timeline, analysis process panel, and result panel are then displayed in the UI. Validation samples are collected, and the location results of the validation samples are compared with the actual fault sources. Based on the comparison results, the actual accuracy of the association rules is calculated, the confidence of the association rules is adjusted, and the rule base and multi-level fault propagation model are updated to achieve continuous learning and optimization of the system.
[0257] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0258] Corresponding to the fault source location method described in the above embodiments, Figure 5 A schematic diagram of the fault source location device provided in the embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0259] Reference Figure 5 The device includes:
[0260] Data acquisition module 501 is used to acquire the current alarm data of the distributed energy system;
[0261] The rule matching module 502 is used to match the current alarm data with each association rule in the fault association rules;
[0262] The propagation graph construction module 503 is used to construct a first fault propagation graph of each device involved in the current alarm data based on the matched association rules;
[0263] The candidate determination module 504 is used to determine each candidate fault source based on the time order of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, and to calculate the target probability of each candidate fault source as a fault source; the target probability of each candidate fault source as a fault source represents the possibility that each candidate fault source is a fault source.
[0264] The source determination module 505 is used to determine the candidate fault source with the highest target probability as the fault source of the current alarm data.
[0265] In some embodiments, the candidate determination module 504 includes:
[0266] The first calculation unit is used to infer candidate fault sources based on the current alarm data and the first fault propagation graph, and to calculate the likelihood probability of the inferred candidate fault source being the fault source based on the first fault propagation graph; the likelihood probability is the probability of observing the current alarm data when the inferred candidate fault source is the fault source.
[0267] The second calculation unit is used to calculate the posterior probability of the inferred candidate fault source as the fault source based on the likelihood probability of the inferred candidate fault source as the fault source and the preset prior probability.
[0268] The format conversion unit is used to format the current alarm data, the matched association rules, and the first fault propagation graph into prompt words;
[0269] The probability determination unit is used to guide the AI assistant to explore different candidate fault sources based on the prompt words, and to determine the AI probability of the different candidate fault sources as fault sources; the AI probability is the probability of the corresponding candidate fault source as a fault source determined by the AI assistant.
[0270] The pattern determination unit is used to analyze the time pattern of fault propagation based on the time sequence of the current alarm data using the time series association rules, so as to determine the association pattern that matches the time pattern in the time series association rules; the time pattern represents the temporal pattern of fault propagation.
[0271] The time determination unit is used to identify candidate fault sources based on the association pattern, and to determine the time series probability of the identified candidate fault sources based on the confidence of the association rule corresponding to the association pattern.
[0272] The weighted fusion unit is used to perform weighted fusion of the posterior probability, artificial intelligence probability and time series probability of the candidate fault source as a fault source for the same candidate fault source, so as to obtain the target probability of the candidate fault source as a fault source.
[0273] In some embodiments, the first computing unit described above is specifically used for:
[0274] From the first fault propagation graph, find the fault propagation path from the inferred candidate fault source to each alarm device in the current alarm data;
[0275] The path probability of a fault propagation path is obtained by multiplying the weights of all edges in any of the fault propagation paths; the weights of the edges represent the confidence of the corresponding association rules.
[0276] Obtain the total delay time of the fault propagation path;
[0277] Based on the total delay time of the fault propagation path and the alarm time of the corresponding alarm device, the time factor of the fault propagation path is calculated; the time factor reflects the degree of matching between the alarm time of the current alarm data and the total delay of the fault propagation path.
[0278] The likelihood probability of the corresponding alarm device is obtained by multiplying the path probability of the fault propagation path by the time factor.
[0279] The product of the likelihood probabilities of each alarm device is calculated to obtain the likelihood probability of the inferred candidate fault source as the fault source.
[0280] In some embodiments, the candidate determination module 504 further includes:
[0281] An evidence embedding unit is used to embed various types of evidence into the prompt words and guide the AI assistant to adjust the evidence weights, filter fault scenarios, and explore different candidate fault sources in the prompt words; the evidence weight is the degree of influence of the corresponding evidence on the fault source location during the fault source analysis process; the various types of evidence include the current alarm data, the performance indicators of each device involved in the current alarm data, log information, and network topology;
[0282] The aforementioned probability determination unit is specifically used for:
[0283] Based on the prompt words, the AI assistant is guided to adjust the evidence weights of various types of evidence and determine the matching fault scenarios;
[0284] Based on the matched fault scenarios and the various types of evidence, explore the different candidate fault sources;
[0285] Based on the adjusted evidence weights, the artificial intelligence probability of the different candidate fault sources as fault sources is determined.
[0286] In some embodiments, the above-described apparatus further includes:
[0287] The data acquisition module is used to acquire historical alarm data of the distributed energy system;
[0288] The feature extraction module is used to extract multi-dimensional correlation features from the historical alarm data to obtain the time series features, spatial correlation features, network topology features, and performance index features of each device pair in the historical alarm data.
[0289] The first generation module is used to convert the numerical features in the time series features into rules, add preset fields, and generate the time series association rules.
[0290] The second generation module is used to convert the physical distance in the spatial association features and the network connection relationship in the network topology features into rules, and add the preset fields to generate spatial topology association rules and functional dependency association rules.
[0291] The third generation module is used to convert the correlation between performance index anomalies and faults in the performance index features into rules, and add the preset fields to generate environmental factor association rules.
[0292] The preset fields include rule identifier, rule type, association pattern description, confidence level, support level, detailed description, and applicable conditions. The confidence level represents the reliability of the corresponding association rule, and the support level represents the universality of the corresponding association rule.
[0293] In some embodiments, the above-described apparatus further includes:
[0294] The frequency statistics module is used to count the number of times each association rule in the fault association rules is accurately predicted and the total number of predictions.
[0295] The accuracy calculation module is used to divide the number of accurate predictions by the total number of predictions to obtain the actual accuracy of the association rule;
[0296] The target determination module is used to divide the total number of predictions by a first preset value to obtain the target value;
[0297] The weight determination module is used to determine the minimum value between the target value and a second preset value as the adjustment weight; the second preset value is less than the first preset value.
[0298] The confidence calculation module is used to calculate the adjusted confidence of the association rule based on the adjusted weight, the confidence of the association rule, and the actual accuracy. The adjusted confidence = the confidence of the association rule × (1 - adjusted weight) + actual accuracy × adjusted weight.
[0299] The confidence update module is used to update the confidence of the association rule to the adjusted confidence.
[0300] In some embodiments, the above-described apparatus further includes:
[0301] The path finding module is used to find the affected devices from the fault source and the fault propagation paths to the affected devices based on a multi-level fault propagation model; the multi-level fault propagation model is used to describe the rules and probabilities of fault propagation among devices in the distributed energy system.
[0302] The propagation display module is used to draw a second fault propagation diagram based on the fault source, each affected device from the fault source, and the fault propagation path to each affected device, and to use animation effects to gradually display the fault propagation process.
[0303] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0304] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 6 (Only one is shown in the diagram), memory 61, and computer program 62 stored in said memory 61 and executable on said at least one processor 60, which, when executed, implements the steps in any of the above method embodiments.
[0305] The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that... Figure 6 This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0306] The processor 60 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0307] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may be an external storage device of the electronic device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 6. Furthermore, the memory 61 may include both internal and external storage units of the electronic device 6. The memory 61 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0308] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0309] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0310] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0311] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0312] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0313] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0314] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for locating the source of a fault, characterized in that, include: Obtain current alarm data from the distributed energy system; Match the current alarm data with each association rule in the fault association rules; Based on the matched association rules, a first fault propagation graph is constructed for each device involved in the current alarm data; Based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time sequence association rules in the fault association rules, each candidate fault source is determined, and the target probability of each candidate fault source as a fault source is calculated. The target probability of each candidate fault source as a fault source represents the likelihood that each candidate fault source is a fault source. The candidate fault source with the highest target probability is determined as the fault source of the current alarm data.
2. The fault source localization method according to claim 1, characterized in that, The process of determining each candidate fault source based on the time sequence of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, and calculating the target probability of each candidate fault source as a fault source, includes: Based on the current alarm data and the first fault propagation graph, candidate fault sources are inferred, and based on the first fault propagation graph, the likelihood probability of the inferred candidate fault source being the fault source is calculated; the likelihood probability is the probability of observing the current alarm data when the inferred candidate fault source is the fault source. Based on the likelihood probability of the inferred candidate fault source as the fault source and the preset prior probability, the posterior probability of the inferred candidate fault source as the fault source is calculated. The current alarm data, the matched association rules, and the first fault propagation graph are formatted into prompt words; Based on the prompt words, the AI assistant is guided to explore different candidate fault sources and determine the AI probability of each different candidate fault source as a fault source; the AI probability is the probability of the corresponding candidate fault source as a fault source determined by the AI assistant. Based on the time sequence of the current alarm data, the time series association rules are used to analyze the time pattern of fault propagation in order to determine the association pattern that matches the time pattern in the time series association rules; the time pattern represents the temporal regularity of fault propagation. Candidate fault sources are identified based on the association patterns, and the time series probability of the identified candidate fault sources is determined based on the confidence of the association rules corresponding to the association patterns. For the same candidate fault source, the posterior probability, artificial intelligence probability, and time series probability of the candidate fault source as a fault source are weighted and fused to obtain the target probability of the candidate fault source as a fault source.
3. The fault source localization method according to claim 2, characterized in that, The step of calculating the likelihood probability of the inferred candidate fault source as the fault source based on the first fault propagation graph includes: From the first fault propagation graph, find the fault propagation path from the inferred candidate fault source to each alarm device in the current alarm data; The path probability of a fault propagation path is obtained by multiplying the weights of all edges in any of the fault propagation paths; the weights of the edges represent the confidence of the corresponding association rules. Obtain the total delay time of the fault propagation path; Based on the total delay time of the fault propagation path and the alarm time of the corresponding alarm device, a time factor of the fault propagation path is calculated; the time factor reflects the degree of matching between the alarm time of the current alarm data and the total delay of the fault propagation path. The likelihood probability of the corresponding alarm device is obtained by multiplying the path probability of the fault propagation path by the time factor. The product of the likelihood probabilities of each alarm device is calculated to obtain the likelihood probability of the inferred candidate fault source as the fault source.
4. The fault source location method according to claim 2, characterized in that, Before guiding the AI assistant to explore different candidate sources of failure based on the prompt words, the process also includes: Various types of evidence are embedded in the prompt words, and the prompt words guide the AI assistant to adjust the evidence weights, filter fault scenarios, and explore different candidate fault sources. The evidence weights are the degree of influence of corresponding evidence on fault source location during fault source analysis. The various types of evidence include the current alarm data, the performance indicators of each device involved in the current alarm data, log information, and network topology. The process of guiding the AI assistant to explore different candidate fault sources based on the prompt words, and determining the AI probability of each different candidate fault source as a fault source, includes: Based on the prompt words, the AI assistant is guided to adjust the evidence weights of various types of evidence and determine the matching fault scenarios; Based on the matched fault scenarios and the various types of evidence, explore the different candidate fault sources; Based on the adjusted evidence weights, the artificial intelligence probability of the different candidate fault sources as fault sources is determined.
5. The fault source localization method according to any one of claims 1 to 4, characterized in that, Before matching the current alarm data with the association rules in the fault association rules, the process also includes: Obtain historical alarm data of the distributed energy system; Multi-dimensional correlation feature extraction is performed on the historical alarm data to obtain the time series features, spatial correlation features, network topology features, and performance index features of each device pair in the historical alarm data. The numerical features in the time series features are transformed into rules, and preset fields are added to generate the time series association rules; The physical distance in the spatial association features and the network connection relationship in the network topology features are transformed into rules, and the preset fields are added to generate spatial topology association rules and functional dependency association rules. The correlation between performance index anomalies and faults in the performance index features is transformed into rules, and the preset field is added to generate environmental factor association rules. The preset fields include rule identifier, rule type, association pattern description, confidence level, support level, detailed description, and applicable conditions. The confidence level represents the reliability of the corresponding association rule, and the support level represents the universality of the corresponding association rule.
6. The fault source localization method according to claim 5, characterized in that, After identifying the candidate fault source with the highest target probability as the fault source of the current alarm data, the method further includes: Calculate the number of accurate predictions for each association rule in the fault association rules and the total number of predictions; Divide the number of accurate predictions by the total number of predictions to obtain the actual accuracy of the association rule; Divide the total number of predictions by the first preset value to obtain the target value; The minimum value between the target value and the second preset value is determined as the adjustment weight; the second preset value is less than the first preset value. Based on the adjusted weights, the confidence of the association rule, and the actual accuracy, the adjusted confidence of the association rule is calculated, where the adjusted confidence = confidence of the association rule × (1 - adjusted weights) + actual accuracy × adjusted weights. The confidence level of the association rule is updated to the adjusted confidence level.
7. The fault source localization method according to any one of claims 1 to 4, characterized in that, After identifying the candidate fault source with the highest target probability as the fault source of the current alarm data, the method further includes: Based on a multi-level fault propagation model, the affected devices from the fault source and the fault propagation paths to the affected devices are identified; the multi-level fault propagation model is used to describe the rules and probabilities of fault propagation among devices in the distributed energy system. Based on the fault source, the affected devices from the fault source, and the fault propagation path to each affected device, a second fault propagation diagram is drawn, and the fault propagation process is displayed step by step using animation effects.
8. A fault source location device, characterized in that, include: The data acquisition module is used to acquire the current alarm data of the distributed energy system; The rule matching module is used to match the current alarm data with each association rule in the fault association rules; The propagation graph construction module is used to construct a first fault propagation graph of each device involved in the current alarm data based on the matched association rules; The candidate determination module is used to determine each candidate fault source based on the time order of the current alarm data, the matched association rules, the first fault propagation graph, and the time series association rules in the fault association rules, and to calculate the target probability of each candidate fault source as a fault source. The target probability of each candidate fault source as a fault source represents the likelihood that each candidate fault source is a fault source. The source determination module is used to determine the candidate fault source with the highest target probability as the fault source of the current alarm data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the fault source location method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, It includes a computer program, which, when run, causes the fault source localization method as described in any one of claims 1 to 7 to be performed.