Fault positioning method and device for power master station service, medium and equipment
By acquiring and classifying multi-source data in the power master station system, utilizing a distributed message bus and correlation analysis model, and combining it with a dedicated network topology model for the power master station, the accuracy and efficiency issues of fault location in power master station services were resolved. This enabled full-dimensional data collection and fault impact assessment, improving operational efficiency and business continuity.
Patent Information
- Application Number
- CN202610229728.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot accurately and efficiently locate faults in power station operations. Especially in distributed architectures, monitoring systems become performance bottlenecks, data silos make fault location difficult, and the lack of cross-data source correlation analysis capabilities fails to meet the business observability requirements of power systems.
By acquiring multi-source data categorized by business topology hierarchy in a distributed environment, asynchronous decoupling is achieved using a distributed message bus, data cleaning and unified identification are performed in conjunction with power master station business rules and tagging systems, cross-dimensional mining is conducted by inputting the data into a correlation analysis model, and fault propagation path tracing is performed by combining the power master station's dedicated network topology model to generate fault location results.
It enables full-dimensional data collection from infrastructure to business logic, improves the real-time performance and stability of the system, accurately locates the root cause of faults and quantifies the impact of faults on business, provides a comprehensive basis for operation and maintenance decisions, and significantly improves the accuracy and efficiency of fault location.
Smart Images

Figure CN122053340A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault location, and in particular to a fault location method, apparatus, medium and equipment for power station operations. Background Technology
[0002] As power systems transform towards intelligence and distributed architecture, the power master station system, as the core of power grid dispatching and operation management, is also gradually evolving its architecture from traditional centralized to distributed. Traditional monitoring systems are built on a centralized architecture, using a centralized monitoring server as the data processing core. Data acquisition agents are deployed at various nodes to collect infrastructure-level performance indicators such as CPU utilization and memory usage from physical machines, virtual machines, and containers. Meanwhile, application logs, database logs, and other operational data are collected by independent log collection systems, resulting in a technical landscape where indicator data and log data are collected and managed separately. The collected monitoring indicators are typically stored in time-series databases, while log data is imported into a dedicated log management system, forming isolated data storage systems that lack an effective data fusion mechanism.
[0003] At the monitoring, alarming, and fault handling level, existing technologies mostly focus on the usage of infrastructure resources and the survival status of service processes. Alarm mechanisms are mainly based on threshold judgments from a single data source, lacking the ability to perform cross-data source correlation analysis. This technical solution exposes significant limitations when dealing with the massive amounts of renewable energy access data and high concurrency access pressure brought by distributed architectures: centralized data processing architecture makes the monitoring system itself a performance bottleneck and a single point of failure, affecting the reliability and real-time performance of monitoring; when data silos cause system failures, maintenance personnel need to manually switch between multiple independent monitoring systems to troubleshoot, making it difficult to quickly locate the root cause of the problem and prolonging the mean time to repair (MTBL); at the same time, existing monitoring systems are too biased towards technical indicators, lacking measurement and alarm capabilities from the perspective of power business services, and cannot correlate system anomalies with specific power grid business functions, failing to meet the urgent needs of modern power systems for business observability.
[0004] Furthermore, there is a serious disconnect between power system monitoring and business impact analysis in existing technologies. After discovering equipment failures on the technical monitoring platform, maintenance personnel must manually consult power grid topology diagrams and business system files, relying on experience to infer the business impact. This is not only inefficient but also prone to improper handling due to human error. These shortcomings prevent existing technologies from accurately and efficiently locating faults in power substation operations. Summary of the Invention
[0005] This invention provides a fault location method, apparatus, medium, and equipment for power station services, to solve the problem that existing technologies cannot accurately and efficiently locate faults in power station services.
[0006] Firstly, this application provides a fault location method for power station operations, including: Acquire multi-source data categorized by business topology hierarchy in a distributed environment; The multi-source data is asynchronously decoupled through a pre-defined distributed message bus; Based on the preset power master station business rules and combined with the preset tag system based on the power master station business scenario, the asynchronously decoupled multi-source data is sequentially cleaned and uniformly identified to obtain the processed multi-source data. The processed multi-source data is input into an association analysis model trained based on the business dependency relationship of the power master station and historical fault data, so that the association analysis model can perform cross-dimensional association mining on the processed multi-source data and identify the root cause of the fault. Based on the identified root cause of the fault, combined with the pre-set dedicated network topology model of the power master station, the fault propagation path is traced through graph computing technology and the impact range of the fault on related power services is deduced, resulting in fault location results including the root cause of the fault, propagation path, fault impact score, business impact range, and handling priority.
[0007] This application achieves full-dimensional data collection from infrastructure to business logic by acquiring multi-source data categorized by business topology in a distributed environment, including infrastructure operation indicators, business process logs, and business operation status data. Asynchronous decoupling of multi-source data using a distributed message bus effectively solves the synchronization bottleneck problem in data transmission, improving the system's real-time performance and stability. Data cleaning and unified labeling are performed using power master station business rules and business scenario tagging systems, achieving standardization and enhanced correlation of heterogeneous data, providing a high-quality data foundation for subsequent analysis. By inputting the processed multi-source data into a correlation analysis model trained on power master station business dependencies and historical fault data, cross-dimensional fault root cause identification is achieved, overcoming the limitations of traditional monitoring systems that rely on a single data source. Furthermore, based on a dedicated network topology model of the power master station, graph computing technology is used to trace the fault propagation path and infer the impact range of the fault on related power businesses, generating fault location results including fault root cause, propagation path, fault impact score, business impact range, and handling priority. This process not only accurately locates the fault root cause but also quantitatively assesses the actual impact of the fault on the business, providing a comprehensive and quantitative basis for operation and maintenance decisions. This application effectively solves the problem that existing technologies cannot accurately and efficiently locate faults in power station operations.
[0008] Furthermore, the asynchronous decoupling of the multi-source data through a preset distributed message bus specifically involves: The collected multi-source data is categorized by data type and business importance level and then connected to the distributed message bus through a preset topic partitioning mechanism. After storing multi-source data through a multi-replica mechanism of a distributed message bus, the multi-source data is asynchronously decoupled so that the asynchronously decoupled multi-source data can run independently at the acquisition end and the processing end.
[0009] This application achieves efficient separation and collaborative operation of data acquisition and processing by asynchronously decoupling multi-source data through a pre-defined distributed message bus. First, the acquired multi-source data is accessed through a pre-defined topic partitioning mechanism based on data type and business importance. This partitioning mechanism ensures that critical business data receives priority processing, improving the system's response speed and processing efficiency for important data. Second, the distributed message bus's multi-replica mechanism stores multi-source data, effectively guaranteeing data reliability and availability. Even in the event of partial node failures, data is not lost, ensuring high system availability. Based on this, asynchronous decoupling of multi-source data allows the data acquisition and processing ends to operate independently. The acquisition end does not need to wait for feedback from the processing end, significantly improving the overall system throughput and response speed. This asynchronous decoupling mechanism not only solves the synchronization bottleneck problem in traditional monitoring systems but also enhances the system's scalability and stability, providing an efficient and reliable data transmission and processing foundation for distributed system monitoring, significantly improving the performance and reliability of the monitoring system.
[0010] Furthermore, this application processes the asynchronously decoupled multi-source data through a pre-defined power master station business rules and business scenario tagging system, achieving deep data integration and standardization. First, a stream processing engine loads the power master station business rules to clean and standardize the asynchronously decoupled multi-source data, removing invalid or abnormal data and unifying the data format, thus obtaining initially standardized first-source data. Next, the first-source data is input into the rule engine, which matches the corresponding data center and service information based on device identifiers and IP addresses, further enriching the data's contextual information to obtain second-source data. Finally, based on the pre-defined power master station business scenario tagging system, a multi-dimensional unified identifier (cluster identifier, service name, and business unit) is added to each piece of data in the second-source data, and this data is stored in a distributed time-series database, forming processed multi-source data. This process not only improves data readability and usability but also establishes relationships between data through multi-dimensional unified identifiers, providing a high-quality, structured data foundation for subsequent intelligent analysis and fault location, significantly enhancing the efficiency and accuracy of distributed system monitoring.
[0011] Furthermore, the process of inputting the processed multi-source data into a correlation analysis model trained based on the business dependencies of the power station and historical fault data, so that the correlation analysis model can perform cross-dimensional correlation mining on the processed multi-source data to identify the root causes of faults, specifically involves: The processed multi-source data is monitored in real time. When an indicator that violates a preset rule is detected in the processed multi-source data, the processed multi-source data is input into the correlation analysis model so that the correlation analysis model can retrieve the corresponding historical indicator data, error logs and call chain data in parallel based on the unified identifier and time window in the processed multi-source data. Extract various multi-source features from the retrieved multi-source correlation data; the multi-source features include anomaly type, numerical deviation degree, duration, topological dependency relationship, and time series fluctuation pattern; Based on the evidence-weighted inference rules trained on the business dependencies of the power master station and historical fault data, the spatiotemporal correlation between various multi-source features is analyzed, and the confidence probability of each fault point is calculated by combining the preset evidence weights in the rule template library. Sort the confidence probabilities from high to low and output the sorted root causes of the fault.
[0012] This application achieves efficient and accurate root cause identification of faults by inputting processed multi-source data into a correlation analysis model trained on the business dependencies of power substations and historical fault data. First, the processed multi-source data is monitored in real time. When an indicator violating preset rules is detected, the correlation analysis model is triggered. The model utilizes unified identifiers and time windows in the data to retrieve corresponding historical indicator data, error logs, and call chain data in parallel, quickly obtaining multi-dimensional information related to the anomaly. Next, multi-source features are extracted from these correlated data, including anomaly type, numerical deviation, duration, topological dependencies, and time series fluctuation patterns, providing a comprehensive feature foundation for fault analysis. Based on evidence-weighted inference rules trained on the business dependencies of power substations and historical fault data, the model analyzes the spatiotemporal correlation between multi-source features and, combined with preset evidence weights in the rule template library, calculates the confidence probability of each fault point. Finally, the confidence probabilities are sorted from high to low, and the sorted root causes of the faults are output.
[0013] This application achieves efficient and accurate root cause identification of faults by inputting processed multi-source data into a correlation analysis model trained on the business dependencies of power substations and historical fault data. First, the processed multi-source data is monitored in real time. When an indicator violating preset rules is detected, the correlation analysis model is triggered. The model utilizes unified identifiers and time windows in the data to retrieve corresponding historical indicator data, error logs, and call chain data in parallel, quickly obtaining multi-dimensional information related to the anomaly. Next, multi-source features are extracted from these correlated data, including anomaly type, numerical deviation, duration, topological dependencies, and time series fluctuation patterns, providing a comprehensive feature foundation for fault analysis. Based on evidence-weighted inference rules trained on the business dependencies of power substations and historical fault data, the model analyzes the spatiotemporal correlation between multi-source features and, combined with preset evidence weights in the rule template library, calculates the confidence probability of each fault point. Finally, the confidence probabilities are sorted from high to low, and the sorted root causes of the faults are output. This process breaks through the limitations of traditional monitoring systems that rely on a single data source and simple threshold judgment. Through cross-dimensional correlation mining and intelligent reasoning, it significantly improves the accuracy and efficiency of fault location, providing fast and accurate technical support for the operation and maintenance of power station systems.
[0014] Furthermore, based on the identified root cause of the fault, combined with a preset dedicated network topology model for the power master station, graph computing technology is used to trace the fault propagation path and deduce the impact range of the fault on related power services, resulting in a fault location result that includes the root cause of the fault, propagation path, fault impact score, service impact range, and handling priority. Specifically: An extensible plug-in framework is integrated into a preset power grid business model, so that the power grid business model, based on a preset business service mapping table, matches the fault entity corresponding to the identified fault root cause with the technical service identifier in the mapping table, and outputs the corresponding power grid equipment identifier and the business function code to which it belongs. The dedicated network topology model of the power master station is loaded into the memory graph calculation engine, and the vertex corresponding to the power grid equipment identifier is determined as the starting vertex for fault propagation analysis; Based on graph computing technology, traversal analysis is performed. Starting from the initial vertex, all reachable downstream vertices are traversed sequentially according to the connection direction of the edges in the topology model to form a fault propagation path. The set of downstream business nodes affected by the fault is selected based on the fault propagation path. The fault impact score is calculated by summing the importance weights of each business node based on the preset importance weights of each node. Based on the fault impact score, the affected nodes are sorted, and combined with the business function scope associated with each node and the constraint rules in the power grid business model, the impact range of the fault on the associated power business is deduced, generating fault location results that include fault root cause, propagation path, fault impact score, business impact range and handling priority.
[0015] This application combines fault root cause identification with a dedicated network topology model of a power substation, utilizing graph computing technology to achieve accurate fault propagation path tracing and business impact range estimation, thereby generating comprehensive fault location results. First, an extensible plug-in framework is used to access the power grid business model, matching the fault entity corresponding to the fault root cause with the technical service identifier in the business service mapping table, outputting the power grid equipment identifier and its associated business function code. Next, the dedicated network topology model of the power substation is loaded into an in-memory graph computing engine. Starting from the vertex corresponding to the fault equipment identifier, all reachable downstream vertices are traversed using graph computing technology to form a fault propagation path. The set of affected downstream business nodes is selected based on the fault propagation path, and a fault impact score is calculated based on preset importance weights. Finally, combining the business function scope and the constraints of the power grid business model, the impact range of the fault on related power businesses is estimated, generating fault location results that include the fault root cause, propagation path, fault impact score, business impact range, and handling priority. This process not only accurately pinpointed the root cause of the fault, but also quantitatively assessed the actual impact of the fault on the business, providing a comprehensive and quantitative basis for operation and maintenance decisions, and significantly improving the operation and maintenance efficiency and business continuity assurance capabilities of the power station system.
[0016] Furthermore, the graph computation technique performs traversal analysis, starting from the initial vertex and sequentially traversing all reachable downstream vertices according to the connection direction of the edges in the topological model, to form a fault propagation path, specifically as follows: Initialize the set of affected nodes, the queue of nodes to be visited, and the set of visited nodes to an empty set, and add the starting vertex to the queue of nodes to be visited and the set of visited nodes. When the queue of nodes to be visited is not empty, retrieve the current vertex from the queue of nodes to be visited and add it to the set of affected nodes; Iterate through all downstream vertices corresponding to the current vertex, and add downstream vertices that have not been added to the visited node set to the visited node set and the unvisited node queue. The fault propagation path is obtained when the queue of nodes to be visited is empty.
[0017] This application utilizes graph computation-based traversal analysis to efficiently construct fault propagation paths, thereby achieving accurate tracking of the fault's impact range. Specifically, the system first initializes the affected node set, the queue of nodes to be visited, and the set of visited nodes to an empty set, and adds the starting vertex to both the queue and the set of visited nodes. Then, while the queue of nodes to be visited is not empty, the current vertex is sequentially retrieved from the queue and added to the affected node set. Simultaneously, all downstream vertices of the current vertex are traversed, and unvisited downstream vertices are added to both the set of visited nodes and the queue of nodes to be visited. This process continues until the queue of nodes to be visited is empty, ultimately yielding the complete fault propagation path. This graph computation-based traversal method not only ensures the accuracy and completeness of the fault propagation path but also significantly improves computational efficiency through an efficient queue management mechanism. This provides a solid technical foundation for rapidly assessing the impact range of faults on power services, thereby enhancing the power system's operation and maintenance response capabilities and business continuity assurance level.
[0018] Secondly, this application provides a fault location device for power station operations. The fault location device for power station operations includes: The acquisition module is used to acquire multi-source data categorized by business topology hierarchy in a distributed environment; A decoupling module is used to asynchronously decouple the multi-source data through a preset distributed message bus; The processing module is used to perform data cleaning and unified identification processing on the asynchronously decoupled multi-source data in sequence, based on the preset power master station business rules and the preset tag system based on the power master station business scenario, to obtain the processed multi-source data. The identification module is used to input the processed multi-source data into an association analysis model trained based on the business dependency relationship of the power master station and historical fault data, so that the association analysis model can perform cross-dimensional association mining on the processed multi-source data and identify the root cause of the fault. The location module is used to trace the fault propagation path and deduce the impact range of the fault on related power services based on the identified root cause of the fault and the preset dedicated network topology model of the power master station through graph computing technology. The result is a fault location result that includes the root cause of the fault, the propagation path, the fault impact score, the scope of the business impact, and the handling priority.
[0019] This application achieves end-to-end optimization of distributed system monitoring and fault location through modular design, significantly improving the operational efficiency and business continuity assurance capabilities of the power master station system. First, the acquisition module collects multi-source data categorized by business topology level, including infrastructure operation indicators, business process logs, and business operation status data, providing a comprehensive data foundation for subsequent analysis. The decoupling module uses a distributed message bus to asynchronously decouple multi-source data, solving the synchronization bottleneck problem in data processing and improving the system's real-time performance and stability. The processing module cleanses and uniformly identifies the data based on the power master station's business rules and business scenario tagging system, enhancing data usability and relevance. The identification module uses a correlation analysis model to perform cross-dimensional correlation mining on the processed data, accurately identifying the root cause of the fault and overcoming the limitations of traditional monitoring relying on a single data source. Finally, the location module combines the power master station's dedicated network topology model and graph computing technology to trace the fault propagation path and deduce the impact range of the fault on related power services, generating fault location results that include the root cause, propagation path, fault impact score, business impact range, and handling priority. This end-to-end optimization not only achieves accurate mapping from technical failures to business impacts, but also provides comprehensive and quantitative basis for operation and maintenance decisions, significantly improving the operation and maintenance efficiency and business continuity assurance capabilities of the power station system.
[0020] Thirdly, this application provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a fault location method for power station services as described above. Its beneficial effects are the same as those of the fault location method for power station services provided in the first aspect of this application.
[0021] Fourthly, this application provides a terminal device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement any of the fault location methods for power station services as described in the first aspect. Attached Figure Description
[0022] Figure 1 A schematic flowchart of an embodiment of the fault location method for power station services provided in this application; Figure 2 A schematic diagram of an embodiment of the data acquisition framework provided in this application; Figure 3 A schematic diagram of an embodiment of the logical architecture provided in this application; Figure 4 A schematic diagram of one embodiment of the overall framework diagram provided in this application; Figure 5 This is a schematic diagram of one embodiment of the fault location device for power station services provided in this application. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example 1 Please refer to Figure 1 In order to solve the problem that existing technologies cannot accurately and efficiently locate faults in power station services, this invention provides a fault location method for power station services, including steps S01-S05.
[0025] S01: Obtain multi-source data categorized by business topology level in a distributed environment.
[0026] As a preferred embodiment of this invention, the step of acquiring multi-source data classified by business topology hierarchy in a distributed environment specifically includes: In a distributed power station system environment, the data acquisition process must first adapt to the system's distributed architecture, covering all nodes such as physical servers, virtual machines, and containers. Then, the acquired data is categorized and organized according to a preset power grid business topology hierarchy (such as substation level, feeder level, business function unit level, etc.) to ensure a clear correspondence between data and business topology. This embodiment constructs a comprehensive data acquisition system covering infrastructure, application logs, and business metrics, such as... Figure 2 As shown, through the collaborative work of three components—distributed data acquisition agent, integrated log collection, and business plugins—the core problems of scattered monitoring data sources, inconsistent formats, and missing business indicators are systematically solved, providing complete and consistent data raw materials for subsequent correlation analysis. This step was specifically designed to meet the stringent stability and real-time requirements of the power station system, ensuring that the data acquisition process itself does not impact the performance of the business system.
[0027] For infrastructure operation metrics, comprehensive data collection is achieved through distributed data collection agents deployed on various nodes. During system initialization, the agent runs as a system daemon, completing initialization by reading a locally pre-defined YAML configuration file. This file explicitly sets key parameters such as the control center's service address and the cluster's unique identifier. The agent then initiates a registration request to the control center via HTTP, establishing a stable service connection. To effectively monitor containerized environments, the agent integrates a Kubernetes client library. By listening for changes in Pod events on the API Server, it detects newly created container instances in real time. Once a new instance is detected, it is immediately added to the monitoring scope, ensuring comprehensive coverage of all running workloads. During data collection, the agent obtains data through various standardized interfaces: by accessing the Docker daemon's REST API interface, it obtains basic metrics such as container-level CPU utilization, memory usage, network traffic, and storage usage; by calling the API interface of the Metrics-Server component in the Kubernetes cluster, it obtains aggregated resource usage data (including CPU and memory usage and limits) at the Pod and node levels. This multi-layered data collection architecture constructs a three-dimensional and hierarchical view of resource usage, from the micro-state of a single container to the meso-state of the entire Pod, and then to the macro-state of cluster nodes. It truly achieves end-to-end observability from infrastructure to application platform, forming a complete infrastructure operation indicator system.
[0028] For business process logs, a unified log collection mechanism enables full collection. During the initialization phase, the collection agent locates various log sources through a multi-path discovery mechanism. Based on preset configuration files or dynamic discovery rules, it automatically identifies the log file paths to be collected (supporting wildcard matching of application log directories), while also being compatible with network log sources such as system log services and Syslog protocol receivers, ensuring effective coverage of various log sources in complex distributed environments. Once the log sources are configured, the system starts a real-time monitoring mechanism, proactively monitoring file system change events (including file creation, content updates, and rotation operations) using the event notification mechanism provided by the operating system kernel, achieving true real-time collection and completely avoiding the detection delays of traditional polling methods. When an update to the log file is detected, the agent immediately reads the new content and sends it to the subsequent processing pipeline: First, it processes the merging of multiple log lines. For multiple log lines such as exception stack traces, the timestamp pattern is used as the starting identifier for the new event, and subsequent consecutive lines with mismatched timestamp formats are automatically merged into the previous event, ensuring the integrity of each exception event. Next, a pattern-matching parsing engine splits the unstructured log text into structured key-value pairs containing timestamps, log levels, thread information, logger names, and core messages. Then, a tag enhancement step is performed to attach a unified resource identifier tag (including the tag of the business service to which it belongs and the specific running instance tag) to each log record, establishing a link between logs and metric data. Finally, a reliable buffered transmission mechanism is used to first write the processed log data to the local disk buffer (ensuring that data is not lost in the event of network interruption), and then transmit it to subsequent stages in an asynchronous non-blocking mode, ensuring the stability and throughput of the entire system and forming standardized business process log data.
[0029] Business operation status data is collected through lightweight plugins integrated into the business applications. The system achieves deep integration with business applications by providing a lightweight software development kit (SDK). Developers first introduce the SSD during the application initialization phase, selecting the appropriate package type for dependency configuration based on the specific technology stack, and specifying key parameters such as the target address for data reporting and the sampling frequency to ensure that the collection process meets the performance requirements of the business system. After completing the basic configuration, the system collects business metrics through two complementary data tracking methods: First, the SSD has built-in intelligent interceptors for common development frameworks, which can automatically capture key performance data without modifying the business code. This non-intrusive tracking can comprehensively collect basic performance metrics such as response time, throughput, and error rate of service interfaces. Second, for the core business logic unique to the power industry, developers can insert precise measurement points at key business nodes through the concise application programming interface provided by the SSD. For example, at the start and end positions of complex business transactions such as power grid model calculations, time consumption statistics and performance data recording can be achieved through simple code calls.
[0030] In addition to performance metrics, the system also supports proactive reporting of business status. Applications can call the metering and counting interfaces provided by the toolkit to report key indicators reflecting business operation status in real time, such as the number of currently online users and the length of the queue of pending computational tasks, achieving comprehensive monitoring coverage from technical performance to business status. In the data processing stage, the toolkit implements an intelligent data optimization mechanism: to avoid frequent reporting impacting system performance, the toolkit aggregates multiple measurements within a short period in memory, significantly reducing the reporting frequency while maintaining data representativeness through statistical methods such as calculating quantile values. Simultaneously, the system automatically attaches semantically relevant tags to each piece of business data, clearly identifying the business unit and transaction type to which the data belongs, laying the foundation for subsequent in-depth correlation analysis with the power grid business model.
[0031] During the data reporting phase, the system employs a reliable transmission mechanism to ensure data integrity: through independent asynchronous processing threads and local cache queues, processed business indicator data is continuously sent to a unified message bus, achieving logical isolation between data collection and business processing, ensuring that monitoring data reporting does not significantly impact the performance of the main business thread. The entire data transmission process utilizes an efficient serialization protocol and a retransmission mechanism to guarantee both transmission efficiency and reliable data delivery. Through this complete business indicator collection and reporting process, the system successfully establishes a data channel between technical indicators and business performance, forming standardized and complete business operation status data, providing a solid data foundation for building a business-aware intelligent monitoring system.
[0032] S02: Asynchronously decouple the multi-source data through a preset distributed message bus.
[0033] In a preferred embodiment of this invention, the asynchronous decoupling of the multi-source data via a preset distributed message bus specifically involves: The system employs a distributed message middleware to build a unified data access layer, forming a standardized data access platform. All multi-source data from acquisition agents and business plugins are first sent to this distributed message bus. Data access adopts a topic partitioning mechanism, logically dividing data according to data type (infrastructure operation indicators, business process logs, business operation status data) and business importance. For example, core business-related data such as power grid dispatching and electricity metering are allocated to dedicated partitions to ensure that critical business data receives priority processing resources. The message bus ensures data reliability through a multi-replica mechanism, synchronously storing each piece of data on multiple nodes. Even if a single node fails, data can be recovered from other replicas, avoiding data loss. Simultaneously, the bus supports configurable data retention policies, allowing operations and maintenance personnel to set data retention durations according to actual needs. This meets the immediate needs of short-term data processing while also supporting subsequent historical data backtracking analysis and fault recovery.
[0034] Through this mechanism, data producers (each acquisition component) and consumers (subsequent data processing stages) achieve complete asynchronous decoupling: the acquisition end does not need to pay attention to the operating status, processing progress, and load of downstream processing stages; it only needs to continuously send the acquired data to the message bus according to the specifications. The processing end can retrieve data from the bus for subsequent processing through pull or push modes according to its own processing capabilities, and can also flexibly adjust the number of processing nodes according to business needs to achieve elastic expansion. This decoupling design effectively improves the stability and data throughput of the entire system, avoiding the impact of pressure fluctuations on one end (such as a sudden surge in data at the acquisition end or a temporary failure at the processing end) on the normal operation of the other end, ensuring the continuity of data transmission and processing.
[0035] S03: Based on the preset power master station business rules and combined with the preset tag system based on the power master station business scenario, the asynchronously decoupled multi-source data is sequentially cleaned and uniformly identified to obtain the processed multi-source data.
[0036] In a preferred embodiment of this invention, based on preset power station business rules and a preset tagging system based on power station business scenarios, the asynchronously decoupled multi-source data is sequentially cleaned and uniformly identified to obtain processed multi-source data. Specifically: The stream processing engine first loads the preset power master station business rules and performs full-process standardized processing on the multi-source data after asynchronous decoupling via the distributed message bus. In the initial stage of processing, the format verification module verifies whether the data structure conforms to the preset standard pattern of the power master station, automatically identifies and removes abnormal data with missing key fields or values exceeding reasonable ranges, and filters redundant information generated by repeated collections, completing the data cleaning stage. Subsequently, a standardized conversion operation is performed to uniformly convert the timestamps of different collection sources into nanosecond-level UTC time format, standardize the measurement units of numerical indicators such as voltage and power into power industry standard units, and uniformly standardize character data with different encoding formats to form a unified first multi-source data.
[0037] The first multi-source data is then input into the rule engine, which performs pattern matching and logical reasoning based on a pre-set power station business rule library. Based on key information such as device identifiers and IP addresses carried in the data, it automatically queries a pre-set resource mapping table to match corresponding contextual information such as the physical location of the data center, the service cluster to which it belongs, and the equipment operation and maintenance responsibility unit. This completes the semantic meaning of the data, giving the originally isolated raw data business-related attributes, thus obtaining the second multi-source data.
[0038] Based on a pre-defined tagging system for power station business scenarios, the system performs multi-dimensional unified identification processing on each piece of data in the second multi-source data. The tagging system covers three core dimensions: cluster identifier and node number at the infrastructure level; service name and instance ID at the application service level; and business unit and functional module code at the business level. This ensures that each piece of data has a unique and complete identity, establishing a crucial link for subsequent cross-data source correlation analysis. After tagging is completed, the system transmits the standardized and tagged multi-source data to a distributed time-series database for persistent storage. This database optimizes the storage structure and query performance for the time-series characteristics of power station monitoring data, ultimately forming well-structured, semantically clear, and directly usable multi-source data for subsequent correlation analysis.
[0039] S04: Input the processed multi-source data into the correlation analysis model trained based on the power station business dependency relationship and historical fault data, so that the correlation analysis model can perform cross-dimensional correlation mining on the processed multi-source data and identify the root cause of the fault.
[0040] In a preferred embodiment of this invention, the process of inputting the processed multi-source data into a correlation analysis model trained based on the business dependencies of the power station and historical fault data, so that the correlation analysis model can perform cross-dimensional correlation mining on the processed multi-source data and identify the root causes of faults, specifically includes: The core of the correlation analysis in this embodiment lies in the construction of a complete closed-loop mechanism of "rule visualization configuration - automatic triggering analysis - in-depth root cause mining". Through the collaboration of a configurable rule engine and intelligent algorithms, the accuracy and flexibility of fault analysis are ensured.
[0041] First, the system provides a graphical association rule editor, allowing operations and maintenance personnel to build complex monitoring rules without requiring professional programming skills. The editor includes four functional collaboration areas, forming a complete rule definition workflow: The data source selection area uses a fine-grained tag filtering mechanism to accurately identify monitoring objects from massive amounts of data. For example, selecting the "service indicator" type and setting the tag "service = power metering" limits the rule's scope to the power metering service cluster; The rule condition configuration area supports hierarchical judgments, from basic conditions with single indicator thresholds, intermediate conditions with logical combinations of multiple indicators, to advanced conditions with time series pattern recognition, and then to composite conditions combining business and technical indicators. It also allows setting durations (e.g., conditions lasting 3 minutes) to filter out instantaneous jitter false alarms; The action definition area can configure multiple parallel query actions, such as querying concurrent error logs, associating database performance indicators, and retrieving call chain data, forming a multi-dimensional analysis chain; The rule testing area supports verifying the rule triggering effect and data retrieval completeness based on historical data, helping to optimize configurations. Meanwhile, the system has a built-in rule template library specifically for the power industry, covering typical scenarios such as database connection pool exhaustion, microservice call chain avalanche, and power transaction timeout. Users can apply these rules directly or make personalized adjustments, significantly reducing the configuration threshold.
[0042] Taking the monitoring of resource bottlenecks in electricity metering services as an example, the rule configuration process is as follows: In the data source selection area, select "Service Indicators" and lock the electricity metering service cluster by label; in the condition configuration area, set the composite condition "CPU utilization > 90% AND memory utilization > 85% AND lasting for 3 minutes", and add the advanced condition "more than 3 times of the same peak pattern within 10 minutes"; in the action definition area, configure parallel actions such as querying the same-period error log, associating database performance indicators, and retrieving call chain data; after testing and verifying through historical fault data, apply the "resource bottleneck causing business anomalies" template to improve the analysis dimensions and complete the rule deployment.
[0043] When the system monitors processed multi-source data in the distributed time-series database in real time, it immediately activates an automated correlation analysis process when any indicator violates preset rules. First, it accurately captures alarm events and generates standardized records containing timestamps, indicator types, trigger values, and associated tags. Then, it quickly matches the alarm events with a preset rule base, selecting applicable analysis processes based on rule priority. Next, it initiates parallel multi-source data retrieval, simultaneously retrieving historical indicator trends from the time-series database, extracting error records from the log storage system, collecting dependency data from the call chain system, and obtaining topology information from the configuration management system, based on a unified identifier and time window, significantly improving data collection efficiency.
[0044] After entering the deep analysis phase, the analysis engine first uses time correlation analysis to clarify the causal sequence of the anomaly, and then performs topological impact analysis based on the service dependency graph to identify the fault propagation path. Subsequently, it activates the intelligent analysis algorithm, based on a multi-source fault root cause inference algorithm using weighted evidence matching. This algorithm intelligently infers the most likely root cause by integrating multi-dimensional monitoring data. Its detailed logic is as follows: The algorithm is formalized as an evidence-weighted scoring model. Its core is to calculate the confidence score of candidate root causes by quantifying the degree of matching between features and root cause rules. Its mathematical model is defined as follows: Let F be the feature vector corresponding to the current anomalous event, and let F be a set of candidate root causes. (Where m is the total number of candidate root causes). For each candidate root cause... Define a set of evidence rules (where k is the total number of predefined evidence rules for the candidate root cause), each rule Described when When true, it represents the expected pattern of observed data features, associated with a confidence weight. , Used for normalization.
[0045] The core of the algorithm is to calculate each candidate root cause. Relative to current features Total matching score: in It is a matching function that returns With rules The degree of fit of the described pattern takes a value in the range [0, 1].
[0046] The algorithm input includes three key types of data: First, the feature vector F, which is a set of structured features extracted from the multi-source data (indicator M, log L, topology T) retrieved through association, for example: Candidate root cause set C: Hypotheses generated based on topology and knowledge base. like .
[0047] Evidence rule base: stores each Corresponding set of evidence rules and their weights .
[0048] The specific steps of the algorithm processing are as follows: Step 1: Feature extraction. Standardize and analyze the input set of related data, extract anomaly types and values from time series indicators, extract error pattern keywords from logs, extract interruption nodes from call chains, etc., and generate a structured feature vector F. Step 2: Candidate set generation. Based on the service dependency topology T and historical failure modes, dynamically generate or retrieve the candidate root cause set C related to this anomaly from the knowledge base. Step 3: Evidence Matching and Scoring. For Each candidate root cause : Initialize its score ; Traverse all its rules of evidence and corresponding weights ; call The function calculates the matching degree. For example, if the rule is "connection rejection error occurred" and F contains this feature, the matching degree can be 1.0; otherwise, it is 0. It also calculates the weighted contribution. And accumulate to .
[0049] Step 4: Sorting and Outputting. After scoring all candidate root causes, sort the output by... Candidate root causes are sorted from highest to lowest value.
[0050] Root cause list sorted in descending order of confidence level: Using this algorithm, the system can accurately calculate the confidence probability of each potential fault point (i.e., The scores are calculated, and the root causes of the faults are sorted from high to low confidence probability. The sorted root causes and their corresponding reasoning are then output to ensure the traceability and credibility of the root causes.
[0051] S05: Based on the identified root cause of the fault, combined with the preset dedicated network topology model of the power master station, the fault propagation path is traced through graph computing technology and the impact range of the fault on related power services is deduced, resulting in fault location results including the root cause of the fault, propagation path, fault impact score, business impact range and handling priority.
[0052] In a preferred embodiment of this example, based on the identified root cause of the fault and combined with a preset power station-specific network topology model, graph computing technology is used to trace the fault propagation path and deduce the impact range of the fault on related power services, resulting in a fault location result that includes the root cause of the fault, propagation path, fault impact score, service impact range, and handling priority. Specifically: The business impact assessment mechanism in this embodiment is based on an extensible plug-in framework. This framework defines standardized data access interfaces specifically for importing and parsing structured power grid business model data. The system executes pre-defined association mapping rules in the rule engine to deeply integrate and calculate the tagged real-time monitoring data stream from the unified data processing layer with the power grid business model, outputting structured business impact assessment data objects. This achieves accurate mapping from technical indicators to specific business impact items, providing standardized input for operation and maintenance decisions.
[0053] The business data access layer supports two standard access methods: RESTful API and configuration files. It receives and parses external business data. The accessed data is structured using a pre-defined JSON or XML schema, and its core includes three key types of data: first, a power grid topology model, which clearly describes the physical connections of substations, feeders, transformers, and other equipment using a graph data structure; second, a business service mapping table, establishing the correspondence between application service IDs and business functional units; and third, an importance level configuration, which specifies the priority weight of each business system in key-value pairs. In addition, the accessed data also includes business rules and constraints that standardize fault propagation paths and impact scope assessment. These business data collectively constitute the knowledge base for business impact analysis.
[0054] The specific execution flow of the business impact analysis engine is as follows: ① Technical Fault Mapping: The engine receives standardized alarm events from the correlation analysis layer. These events contain core information such as [fault entity identifier, timestamp, and fault type]. The engine queries the business service mapping table and, by matching the fault entity identifier with the technical service identifier field in the mapping table, accurately obtains one or more corresponding business function codes and the identifier of the associated power grid equipment, thus completing the mapping conversion from technical faults to power grid business entities.
[0055] ② Impact Propagation Analysis: The engine uses the power grid equipment identifiers output in step ① as the starting vertices and loads the power grid topology model into the memory graph computation engine. In this embodiment, the core of impact propagation analysis and quantification relies on graph computation technology based on directed graph traversal, and its detailed logic is as follows: This technology abstracts the power grid business model into a directed graph: .
[0056] V (Vertex Set): Represents substations, transformers, feeders, or service function units in the power grid.
[0057] E (Edge Set): Represents the physical connection relationships or business logic dependencies between devices. This indicates the direction of influence propagation from vertex u to vertex v.
[0058] W (weight function): for each vertex Assign a business importance weight The weight values are usually derived from business configuration.
[0059] The algorithm input consists of three key parts: first, the starting vertex, which is the business model node obtained by mapping the technical fault entity through the business service mapping table; second, the business topology graph, which is a graph structure model containing all power grid equipment and business functional units and their connection relationships; and third, the fault severity coefficient, which is a positive real number obtained from the technical fault analysis report and is used to uniformly adjust the impact base corresponding to the business weight.
[0060] The algorithm's processing steps are as follows: First, initialization: create an empty set AffectedSet to store affected vertices, an empty queue Queue to manage vertices to be visited, and an empty set Visited to record visited vertices to prevent duplication. Add the starting vertex vfault to Queue and Visited. Second, breadth-first traversal: when Queue is not empty, dequeue the current vertex vcurrent and add it to AffectedSet. Traverse all edges (vcurrent, vnext) starting from vcurrent. Add vnext that is not in Visited to Visited and Queue until Queue is empty, completing the traversal of all downstream reachable vertices. Through this process, a complete fault propagation path is formed, clarifying the trajectory of fault spread from the starting vertex downstream.
[0061] ③ Business Value Assessment: The engine calls a pre-defined quantitative assessment function, taking the affected vertex set AffectedSet obtained in step ② as input, traversing each vertex v in the set, and querying the importance level configuration based on its corresponding business function code to obtain the weight. According to the formula Calculate the overall impact score of the fault (where s is the fault severity coefficient), which directly reflects the overall impact of the fault on the business.
[0062] ④ Priority Calculation: The engine maintains a global fault handling queue, inserting fault events generated in this analysis and their comprehensive impact scores into the queue. The queue has a mechanism for automatically sorting by comprehensive impact score in descending order, and can automatically generate fault handling priority suggestions based on the sorting results, ensuring that high-impact faults related to core business are handled first; Based on the above analysis process, the system encapsulates the output of the business impact analysis engine into a standardized data interface. The data objects output by this interface specifically include: a list of affected business systems, statistical values of the scope and number of affected customers, abnormal states of key business indicators, and predicted fault recovery times based on historical data. These data objects can directly drive the front-end visual dashboard for display, allowing operations and maintenance personnel to intuitively grasp the specific impact of faults on business operations.
[0063] By organically combining business data access, intelligent analysis engines, and visual dashboards, this system has established a complete business impact assessment system for technical failures. This system can not only accurately determine the actual impact of technical failures on business operations, but also provide quantitative basis for prioritizing failure handling decisions, realizing the transformation from traditional technical operation and maintenance to business value assurance, and effectively improving the operation and maintenance support capabilities and business continuity management level of the power station system.
[0064] like Figure 3 As shown, Figure 3 The presentation showcases a technical process centered on a three-layer logical architecture: a unified data acquisition layer, a unified data processing layer, and a unified correlation analysis layer. The bottom layer, the unified data acquisition layer, uses three components—a distributed acquisition agent, an integrated log collection system, and business plugins—to acquire infrastructure operational metrics (CPU, memory, etc.), business process logs, and business operational status data, achieving full coverage of multi-source data. The middle layer, the unified data processing layer, relies on a distributed message bus to achieve asynchronous data decoupling. After data cleaning, standardization transformation, and multi-dimensional tagging processing by a stream processing engine, the data is persisted to a distributed time-series database. The top layer, the unified correlation analysis layer, uses a configurable rule engine and intelligent algorithms (based on a directed graph-based business impact propagation algorithm and a multi-source evidence weighted root cause reasoning algorithm) to identify root causes of faults and assess business impact.
[0065] like Figure 4 As shown, Figure 4This application's overall framework diagram (or a core flowchart of real-time regional power grid dispatching / passive overload energy storage call for power grid components, adapted to specific scenarios) clearly presents the technical chain and data flow with "data acquisition - data processing - correlation analysis / decision execution" as the core logic: The bottom layer acquires power station component operation data and energy storage resource data (or power grid static model, topology, and real-time operating parameters) through distributed acquisition agents and monitoring equipment. After asynchronous decoupling via a distributed message bus, the stream processing engine cleans, standardizes, and tags the data before storing it in a time-series database; the middle layer relies on... It can be configured with a rule engine and a multi-objective co-evolutionary algorithm (or a power deficit calculation model) to achieve scenario similarity matching and initial population optimization (or overload condition judgment and load control calculation). The top layer uses graph computing technology and a multi-source evidence weighted root cause reasoning algorithm (or energy storage discharge / load shedding command issuance logic) to complete fault root cause identification and propagation path tracing (or overload continuous control), and finally outputs fault location results (or real-time power distribution dispatch / energy storage call commands). It comprehensively solves the problems of data silos and inefficient decision-making in traditional power monitoring, and supports the stable operation and precise control of the power master station system.
[0066] In summary, this application achieves comprehensive data collection from infrastructure to business logic by acquiring multi-source data categorized by business topology in a distributed environment, including infrastructure operation indicators, business process logs, and business operation status data. Asynchronous decoupling of multi-source data using a distributed message bus effectively solves the synchronization bottleneck problem in data transmission, improving the system's real-time performance and stability. Data cleaning and unified labeling are performed using power master station business rules and business scenario tagging systems, achieving standardization and enhanced correlation of heterogeneous data, providing a high-quality data foundation for subsequent analysis. By inputting the processed multi-source data into a correlation analysis model trained on power master station business dependencies and historical fault data, cross-dimensional fault root cause identification is achieved, overcoming the limitations of traditional monitoring systems that rely on a single data source. Furthermore, based on the power master station's dedicated network topology model, graph computing technology is used to trace fault propagation paths and infer the impact range of faults on related power businesses, generating fault location results that include fault root causes, propagation paths, fault impact scores, business impact range, and handling priorities. This process not only accurately pinpoints the root cause of the fault but also quantitatively assesses the actual impact of the fault on operations, providing a comprehensive and quantitative basis for operational and maintenance decisions. This application effectively solves the problem that existing technologies cannot accurately and efficiently locate faults in power substation operations.
[0067] Example 2 Please refer to Figure 5 This is a fault location device for power station services provided in the embodiments of this application.
[0068] In this embodiment, the fault location device for power station services includes an acquisition module 10, a decoupling module 20, a processing module 30, an identification module 40, and a location module 50.
[0069] Module 10 is used to acquire multi-source data classified by business topology level in a distributed environment; The decoupling module 20 is used to asynchronously decouple the multi-source data through a preset distributed message bus; The processing module 30 is used to perform data cleaning and unified identification processing on the asynchronously decoupled multi-source data in sequence based on the preset power master station business rules and the preset tag system based on the power master station business scenario, so as to obtain the processed multi-source data. The identification module 40 is used to input the processed multi-source data into the association analysis model trained based on the power station business dependency relationship and historical fault data, so that the association analysis model can perform cross-dimensional association mining on the processed multi-source data and identify the root cause of the fault. The location module 50 is used to trace the fault propagation path and deduce the impact range of the fault on related power services based on the identified fault root cause and the preset power master station dedicated network topology model, through graph computing technology, to obtain fault location results including fault root cause, propagation path, fault impact score, service impact range and handling priority.
[0070] For ease of description and brevity, the embodiments of the device of the present invention include all the implementation methods in the above embodiments of the fault location method for power station services, and will not be repeated here.
[0071] Example 3: This application provides a computer-readable storage medium, which includes a stored computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the fault location method for power station services. The fault location method for power station operations, when implemented as a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments can also be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0072] Example 4 This embodiment provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any of the fault location methods for power station services as described in Embodiment 1.
[0073] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A fault location method for power station operations, characterized in that, include: Acquire multi-source data categorized by business topology hierarchy in a distributed environment; The multi-source data is asynchronously decoupled through a pre-defined distributed message bus; Based on the preset power master station business rules and combined with the preset tag system based on the power master station business scenario, the asynchronously decoupled multi-source data is sequentially cleaned and uniformly identified to obtain the processed multi-source data. The processed multi-source data is input into an association analysis model trained based on the business dependency relationship of the power master station and historical fault data, so that the association analysis model can perform cross-dimensional association mining on the processed multi-source data and identify the root cause of the fault. Based on the identified root cause of the fault, combined with the pre-set dedicated network topology model of the power master station, the fault propagation path is traced through graph computing technology and the impact range of the fault on related power services is deduced, resulting in fault location results including the root cause of the fault, propagation path, fault impact score, business impact range, and handling priority.
2. The fault location method for power station operations according to claim 1, characterized in that, The asynchronous decoupling of the multi-source data through a preset distributed message bus specifically involves: The collected multi-source data is categorized by data type and business importance level and then connected to the distributed message bus through a preset topic partitioning mechanism. After storing multi-source data through a multi-replica mechanism of a distributed message bus, the multi-source data is asynchronously decoupled so that the asynchronously decoupled multi-source data can run independently at the acquisition end and the processing end.
3. The fault location method for power station operations according to claim 1, characterized in that, Based on preset power station business rules and a preset tagging system for power station business scenarios, the asynchronously decoupled multi-source data is sequentially cleaned and uniformly identified to obtain processed multi-source data, specifically: By loading preset power station business rules through the stream processing engine, the asynchronously decoupled multi-source data is cleaned and standardized to obtain the first multi-source data. The first multi-source data is input into a preset rule engine, so that the rule engine matches the corresponding data center and service information based on the device identifier and IP address in the first multi-source data to obtain the second multi-source data; Based on the preset power master station business scenario labeling system, a multi-dimensional unified identifier of cluster identifier, service name and business unit is added to each piece of data in the second multi-source data; The multi-source data, after being added with a unified identifier, is stored in a distributed time-series database to obtain the processed multi-source data.
4. The fault location method for power station operations according to claim 1, characterized in that, The process involves inputting the processed multi-source data into a correlation analysis model trained based on the business dependencies of power master stations and historical fault data. This allows the correlation analysis model to perform cross-dimensional correlation mining on the processed multi-source data and identify the root causes of faults. Specifically, this process includes: The processed multi-source data is monitored in real time. When an indicator that violates a preset rule is detected in the processed multi-source data, the processed multi-source data is input into the correlation analysis model so that the correlation analysis model can retrieve the corresponding historical indicator data, error logs and call chain data in parallel based on the unified identifier and time window in the processed multi-source data. Extract various multi-source features from the retrieved multi-source correlation data; the multi-source features include anomaly type, numerical deviation degree, duration, topological dependency relationship, and time series fluctuation pattern; Based on the evidence-weighted inference rules trained on the business dependencies of the power master station and historical fault data, the spatiotemporal correlation between various multi-source features is analyzed, and the confidence probability of each fault point is calculated by combining the preset evidence weights in the rule template library. Sort the confidence probabilities from high to low and output the sorted root causes of the fault.
5. The fault location method for power station operations according to claim 1, characterized in that, Based on the identified root cause of the fault, combined with a pre-set dedicated network topology model for the power station, graph computing technology is used to trace the fault propagation path and deduce the impact range of the fault on related power services. This yields a fault location result including the root cause, propagation path, fault impact score, service impact range, and handling priority. Specifically: An extensible plug-in framework is integrated into a preset power grid business model, so that the power grid business model, based on a preset business service mapping table, matches the fault entity corresponding to the identified fault root cause with the technical service identifier in the mapping table, and outputs the corresponding power grid equipment identifier and the business function code to which it belongs. The dedicated network topology model of the power master station is loaded into the memory graph calculation engine, and the vertex corresponding to the power grid equipment identifier is determined as the starting vertex for fault propagation analysis; Based on graph computing technology, traversal analysis is performed. Starting from the initial vertex, all reachable downstream vertices are traversed sequentially according to the connection direction of the edges in the topology model to form a fault propagation path. The set of downstream business nodes affected by the fault is selected based on the fault propagation path. The fault impact score is calculated by summing the importance weights of each business node based on the preset importance weights of each node. Based on the fault impact score, the affected nodes are sorted, and combined with the business function scope associated with each node and the constraint rules in the power grid business model, the impact range of the fault on the associated power business is deduced, generating fault location results that include fault root cause, propagation path, fault impact score, business impact range and handling priority.
6. The fault location method for power station operations according to claim 5, characterized in that, The graph-based computation technique performs traversal analysis, starting from the initial vertex and sequentially traversing all reachable downstream vertices according to the connection direction of the edges in the topological model, to form a fault propagation path, specifically as follows: Initialize the set of affected nodes, the queue of nodes to be visited, and the set of visited nodes to an empty set, and add the starting vertex to the queue of nodes to be visited and the set of visited nodes. When the queue of nodes to be visited is not empty, retrieve the current vertex from the queue of nodes to be visited and add it to the set of affected nodes; Iterate through all downstream vertices corresponding to the current vertex, and add downstream vertices that have not been added to the visited node set to the visited node set and the unvisited node queue. The fault propagation path is obtained when the queue of nodes to be visited is empty.
7. A fault location device for power station operations, characterized in that, include: The acquisition module is used to acquire multi-source data categorized by business topology hierarchy in a distributed environment; A decoupling module is used to asynchronously decouple the multi-source data through a preset distributed message bus; The processing module is used to perform data cleaning and unified identification processing on the asynchronously decoupled multi-source data in sequence, based on the preset power master station business rules and the preset tag system based on the power master station business scenario, to obtain the processed multi-source data. The identification module is used to input the processed multi-source data into an association analysis model trained based on the business dependency relationship of the power master station and historical fault data, so that the association analysis model can perform cross-dimensional association mining on the processed multi-source data and identify the root cause of the fault. The location module is used to trace the fault propagation path and deduce the impact range of the fault on related power services based on the identified root cause of the fault and the preset dedicated network topology model of the power master station through graph computing technology. The result is a fault location result that includes the root cause of the fault, the propagation path, the fault impact score, the scope of the business impact, and the handling priority.
8. The fault location device for power station operations according to claim 7, characterized in that, The identification module includes: The retrieval unit is used to monitor the processed multi-source data in real time. When the processed multi-source data is found to contain indicators that violate preset rules, the processed multi-source data is input into the correlation analysis model so that the correlation analysis model can retrieve the corresponding historical indicator data, error logs and call chain data in parallel based on the unified identifier and time window in the processed multi-source data. The extraction unit is used to extract various multi-source features from the retrieved multi-source associated data; the multi-source features include anomaly type, numerical deviation degree, duration, topological dependency relationship and time series fluctuation pattern; The analysis unit is used to analyze the spatiotemporal correlation between various multi-source features based on the evidence-weighted inference rules trained on the power station business dependency relationship and historical fault data, and to calculate the confidence probability of each fault point by combining the preset evidence weights in the rule template library. The sorting unit is used to sort the confidence probabilities from high to low and output the sorted root causes of the faults.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the fault location method for power station services as described in any one of claims 1 to 6.
10. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the fault location method for power station services as described in any one of claims 1 to 6.