Wind power plant network equipment fault accurate positioning method based on multi-data fusion
By establishing equipment information and network topology models, and combining multi-dimensional fault characteristics and rule engine analysis, the system achieves accurate location of faults in wind farm network equipment, solving the problems of unclear fault root cause identification and inaccurate impact range assessment in existing technologies, and improving fault diagnosis efficiency and operation and maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-13
AI Technical Summary
The existing monitoring and fault handling methods of wind farm production control networks lack the ability to comprehensively analyze multiple types of data, making it difficult to accurately locate the root cause of the fault and its scope of impact. Especially in the centralized operation and maintenance scenario with multiple wind farms and multiple monitoring units, the fault investigation cycle is long and the recovery efficiency is low.
By establishing device information models and network topology models, collecting multi-dimensional fault characteristics, analyzing alarm events using a fault location rule engine, and combining time window grouping and device operation data, the root cause device and affected devices are automatically identified, generating accurate fault location results and displaying them in the topology view.
It significantly improves the automation and accuracy of root cause analysis of network equipment failures, reduces reliance on human experience judgment, shortens troubleshooting time, and improves operation and maintenance efficiency and operational reliability.
Smart Images

Figure CN121664618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection technology for network equipment, specifically to a method for accurate fault location of wind farm network equipment based on multi-data fusion. Background Technology
[0002] Currently, large-scale wind farms generally adopt a hierarchical and zoned production control network structure. Multiple subsystems, such as the wind turbine monitoring system, the booster station monitoring system, and the site operation management system, rely on network equipment such as field servers, switches, optical transceivers, firewalls, and forward isolation devices to achieve data acquisition, monitoring, and remote control. As the installed capacity of wind farms continues to increase, the same enterprise often manages multiple wind farms in different regions. Each site is connected to the central control center via dedicated lines or private networks. The overall network structure is characterized by the coexistence of multiple sites, multiple network segments, and multiple types of equipment, and the network scale and topology are becoming increasingly complex.
[0003] To ensure the safe and stable operation of the wind farm's production control network, existing technologies typically employ network management systems or simple inspection scripts to perform connectivity checks and performance monitoring on critical network devices. For example, this involves periodically performing ICMP probes and SNMP polls on core switches, front-end servers, firewalls, and other devices to record basic indicators such as device online status, CPU utilization, and memory utilization. Simultaneously, certain alarm thresholds are configured locally or on the centralized control side. When indicators exceed the thresholds or devices become unreachable, alarms are generated, and the relevant information is displayed on the monitoring interface in the form of a list or topology view. Some systems also support alarm classification, confirmation, and simple statistical analysis to assist maintenance personnel in assessing network operating conditions.
[0004] However, in the power production control scenario of wind farms, existing network monitoring and fault handling methods still have significant shortcomings: First, most systems focus on single-dimensional equipment connectivity or a few performance indicators, lacking the ability to comprehensively analyze multiple types of data such as equipment static information, operating performance, power status, network topology, and site-level information. This results in unclear correlations between alarms, making it difficult to identify the true root cause of the fault from a large number of alarms in a timely manner. Second, existing general network management systems are mostly designed for information or communication networks, without fully considering the business relationships between wind turbine units, booster station units, and the site and control center. When a network fault occurs, it can often only locate the anomaly in a certain network segment or equipment, and cannot pinpoint the specific link, port, and its corresponding wind turbine or booster station unit, resulting in inaccurate assessment of the impact range. Third, due to the remote geographical location of wind farms, the limited number of on-duty personnel, and the adoption of forward isolation and security zoning measures in the production control network, the ability to transmit fault information across sites and provide unified operation and maintenance support is limited. Existing technologies still rely heavily on manual experience in alarm aggregation, remote analysis, and rapid location, resulting in long fault investigation cycles and low recovery efficiency.
[0005] In summary, existing monitoring and maintenance solutions for wind farm production control networks are insufficient to fully utilize multi-source data for automated and accurate fault location of network equipment. This is especially true in centralized maintenance scenarios involving multiple wind farms and multiple monitoring units. There is a lack of technical solutions that can simultaneously provide the root cause of the fault and its impact at both the network topology and the site's business level, which requires further improvement. Summary of the Invention
[0006] This invention provides a method for accurate fault location of wind farm network equipment based on multi-data fusion.
[0007] The present invention solves the above-mentioned technical problems through the following technical solution:
[0008] A method for accurate fault location of wind farm network equipment based on multi-data fusion is applied to a monitoring system including a wind farm production control network, comprising:
[0009] (1) Obtain network equipment information and inter-equipment connection relationships in the wind farm production control network, and establish equipment information model and network topology model;
[0010] (2) Collect operating data from the network device according to a preset period, and generate alarm events associated with the network device based on a preset threshold and / or historical operating data, and store them in the alarm event database;
[0011] (3) Group alarm events within the same time window according to the time window, and associate each alarm event with the corresponding operating data, the equipment information model, the network topology model and the wind farm and monitoring unit information to form multi-dimensional fault characteristics for each network device;
[0012] (4) Input the multidimensional fault characteristics and network topology relationship into the fault location rule engine, analyze the alarm event groups in each time window, and determine the root cause device and / or root cause link and fault type according to the pre-configured fault location rules.
[0013] (5) Based on the network topology model, determine the set of affected network devices from the root cause device and / or root cause link in the downstream direction, map the root cause device and / or root cause link and the affected network devices to the wind farm network topology view and output the fault location result containing the fault location, fault type and impact range to the operation and maintenance terminal.
[0014] In a specific embodiment, the operational data collected in step (2) includes at least device connectivity indicators, operational performance indicators and device status information. The device connectivity indicators include at least one of packet loss rate and latency, and the operational performance indicators include at least one of CPU utilization, memory utilization, storage space utilization and port traffic.
[0015] In a specific embodiment, step (2) uses a combination of static threshold and dynamic baseline to generate alarm events. The dynamic baseline is calculated based on historical operating data within a preset statistical period. When the current operating data deviates from the dynamic baseline by more than a preset multiple and continues for a preset time, an alarm event is triggered.
[0016] In a specific embodiment, the multidimensional fault features in step (3) include at least: time features representing the alarm occurrence time and duration, topological location features representing the hierarchy and upstream and downstream relationships of network devices in the topology, device attribute features representing device type and network role, operational performance features representing the degree of abnormality of operational indicators, and site hierarchy features representing the wind farm and monitoring unit to which the device belongs.
[0017] In a specific embodiment, in step (4), the fault location rule engine scores the candidate root cause devices and / or candidate root cause links. The scoring is based at least on the severity of the device status, the topology level weight, and the alarm level and duration of the corresponding alarm event, and selects the candidate with the highest score as the root cause device and / or root cause link.
[0018] In a specific embodiment, the fault location rule includes at least an upstream link fault rule, which is used to determine a network device and / or its connected link as the root cause when the following conditions are met: the network device and its direct downstream network devices are in an unreachable state, while other network devices connected to its upstream network device but not through the network device are in a reachable state.
[0019] In a specific embodiment, the fault location rule further includes a performance bottleneck fault rule, which is used to determine a target network device as the root cause when the following conditions are met: the CPU utilization and / or port bandwidth utilization of the target network device continuously exceed a preset threshold in multiple collection cycles, and the latency of multiple downstream network devices connected to the target network device increases significantly but connectivity remains available.
[0020] In a specific embodiment, determining the set of affected network devices in step (5) includes: traversing downstream of the network topology from the root cause device and adding network devices that have alarm events and / or abnormal operating data within the current time window and whose service paths pass through the root cause device and / or root cause link to the set of affected network devices.
[0021] In a specific embodiment, the method further includes: writing the identifier of the root cause device and / or root cause link and the fault location result into the alarm event database, and establishing an association with the corresponding alarm event, so as to merge and display multiple alarms caused by the same root cause in the operation and maintenance terminal.
[0022] This invention establishes equipment information models and network topology models, and performs unified collection and fusion modeling of multi-source data such as equipment connectivity, operating performance, status information, and wind farm and monitoring unit hierarchical information. Compared with existing monitoring methods that mainly rely on single connectivity detection or a few performance indicators, this invention can automatically analyze the correlation between alarms and operating data within the system, thereby significantly improving the automation and accuracy of network equipment fault root cause analysis and reducing reliance on human experience judgment.
[0023] This invention utilizes network topology models and wind farm / monitoring unit hierarchical information. Based on the determination of root cause devices and / or root cause links using a rule engine, it identifies the set of affected network devices from the root cause node along the downstream direction of the topology and visualizes it in the topology view. This allows fault location to be refined from the traditional "network segment or device level" to specific links, ports, and corresponding wind turbine units or substation units, accurately providing the fault location and its impact range. This solves the problems of unclear impact range assessment and insufficient location granularity in existing technologies.
[0024] This invention groups alarm events by time windows, scores and filters candidate root causes by combining pre-configured fault location rules and multi-dimensional fault characteristics, and writes the root cause results back to the alarm event database to achieve alarm merging and display. In scenarios with centralized operation and maintenance of multiple wind farms and multiple monitoring units, it can effectively suppress a large number of duplicate alarms caused by the same fault, help operation and maintenance personnel quickly focus on key fault points, shorten fault investigation and recovery time, and improve the overall operation and maintenance efficiency and operational reliability of the wind farm production control network. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in this utility model or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this utility model. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the overall process steps of an embodiment of the present invention.
[0027] Figure 2 This is a global topology view of an embodiment of the present invention.
[0028] Figure 3 This is a diagram showing the relationship between key elements of a station topology in one embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of this utility model will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this utility model, and not all embodiments. Based on the embodiments of this utility model, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this utility model.
[0030] Example 1
[0031] This embodiment provides a method for accurate fault location of wind farm network equipment based on multi-data fusion, which is applied to a monitoring system that includes production control networks of multiple wind farms. This monitoring system can be deployed according to the customer's existing network model. For example, one intelligent inspection server can be deployed at each wind farm site, and maintenance terminals and data display servers can be deployed at the centralized control center. Data upload and display are achieved through a dedicated network and forward isolation devices.
[0032] A wind farm production control network typically includes: wind turbine monitoring units, substation monitoring units, and integrated automation systems. Each unit is equipped with various network devices such as front-end servers, database servers, process layer and bay layer switches, optical transceivers, firewalls, and forward isolation devices. These devices are connected by optical fibers or cables to form a multi-level star or ring network topology.
[0033] The method in this embodiment is mainly executed by the intelligent inspection server on the site side, and its specific steps are as follows.
[0034] Step (1): Equipment information collection and model building
[0035] During the system deployment phase, maintenance personnel first import the equipment list of each wind farm into the intelligent inspection server. The equipment list can be maintained via Excel spreadsheets or configuration files, and each record must include at least:
[0036] Unique Device Identifier (ID);
[0037] Equipment Name;
[0038] Equipment types (core switches, aggregation switches, access switches, firewalls, servers, optical transceivers, forward isolation devices, etc.);
[0039] Network roles (core layer, aggregation layer, access layer, security boundary, business server, etc.);
[0040] The identifier of the wind farm is field_id;
[0041] Type of monitoring unit (wind turbine monitoring unit, substation monitoring unit, etc.);
[0042] Data center name, rack number, and USB port information;
[0043] Manage IP address mgmt_ip;
[0044] Equipment manufacturer and model information;
[0045] Device port resource information, such as the number of ports and their speed.
[0046] The intelligent inspection server writes the above information into the device information table (e.g., T_DEVICE) to form a device information model.
[0047] To establish a network topology model, operations and maintenance personnel further configure the physical and logical connections between devices. Each connection relationship includes at least:
[0048] Source device ID, source port number;
[0049] Target device ID, target port number;
[0050] Link type (fiber optic link, cable link, logical tunnel, etc.);
[0051] Link bandwidth.
[0052] The intelligent inspection server writes the connection relationships into the topology link table (e.g., T_TOPO_LINK) and generates network topology data based on the device information table and the link table. Optionally, the system also assigns display coordinates in the topology view to each network device, forming a topology node table (e.g., T_TOPO_NODE).
[0053] Based on the aforementioned equipment information model and network topology model, the intelligent inspection server can generate a wind farm network topology view on the operation and maintenance terminal, displaying nodes such as core switches, firewalls, front-end servers, wind turbine monitoring units, and booster station monitoring units, as well as their connection relationships, according to the wind farm dimension.
[0054] Step (2): Run data acquisition and alarm generation
[0055] Data collection cycle and protocol configuration: Operation and maintenance personnel configure a unified inspection cycle on the intelligent inspection server, such as polling all registered network devices every 5 minutes.
[0056] For devices such as switches, firewalls, and optical transceivers that support the SNMP protocol, the SNMPv3 protocol is used to read the performance and status information in the device MIB;
[0057] For some server devices, information such as system load and disk usage can be obtained by installing an agent or executing monitoring scripts using the SSH protocol;
[0058] For all network devices, basic connectivity and latency are detected using ICMPping.
[0059] Operational data content: For each inspection, the intelligent inspection server generates an operational data record for each device and writes it to the operational data table (e.g., T_INSPECTION). The main fields include:
[0060] timestamp inspect_time;
[0061] Device ID;
[0062] Equipment connectivity metrics:
[0063] Ping packet loss rate (%)
[0064] Ping average round-trip time (RTT, ms);
[0065] Operating performance indicators:
[0066] CPU utilization (%), such as obtained via SNMPOID;
[0067] Memory utilization (%)
[0068] Storage space utilization rate (%)
[0069] Inbound / outbound traffic (bps) and bandwidth utilization (%) of key ports;
[0070] Device status information:
[0071] Critical port up / down status;
[0072] Equipment operating status (normal, alarm, down);
[0073] Power status (Main power supply normal / abnormal, backup power supply normal / abnormal).
[0074] Optional environmental conditions (temperature, fan status, etc.).
[0075] Static threshold configuration: The system sets default static thresholds for different types of devices, for example:
[0076] Core Switch:
[0077] CPU utilization >80% and continues for 3 inspection cycles;
[0078] Single-port bandwidth utilization >90% and continues for 3 inspection cycles;
[0079] Access switch:
[0080] Packet loss rate > 20%;
[0081] server:
[0082] Memory utilization > 85%;
[0083] Storage space utilization >90%;
[0084] All equipment:
[0085] Multiple consecutive ping attempts that fail (e.g., 3 in a row) are considered as the device being unreachable.
[0086] Dynamic baseline calculation: To adapt to the actual load characteristics of different wind farms, the system can also enable the dynamic baseline function. Taking CPU utilization as an example:
[0087] Select the most recent 7 days or 30 days as the statistical period, and calculate the mean μ and standard deviation σ of the CPU utilization of each device by hour or by inspection cycle;
[0088] μ+k×σ (k is an empirical coefficient, such as 2 or 3) is used as the dynamic baseline of this index;
[0089] When the deviation between the current collected value and the dynamic baseline exceeds a preset multiple (e.g., current value > μ + 3σ) and the duration exceeds a preset duration (e.g., 10 minutes), a dynamic baseline alarm is triggered.
[0090] Alarm Event Generation: After completing an inspection, the intelligent inspection server performs both static threshold and dynamic baseline checks on each running data record. If either condition is met, an alarm event is generated and written to the alarm event database (e.g., T_ALARM), with fields including:
[0091] Alarm ID;
[0092] Device ID;
[0093] The field_id of the wind farm;
[0094] The monitoring unit to which it belongs (wind turbine monitoring unit, substation monitoring unit, etc.);
[0095] alarm time (alarm_time);
[0096] Alarm levels (urgent, important, minor, alert);
[0097] Alarm types (such as device offline alarm, link interruption alarm, high CPU utilization alarm, high port bandwidth utilization alarm, power failure alarm, etc.).
[0098] Trigger metric types (such as RTT, packet loss rate, CPU utilization, etc.);
[0099] Current indicator value: current_value;
[0100] The corresponding threshold or dynamic baseline value is threshold_value;
[0101] Alarm status (active / recovered).
[0102] Step (3): Time window grouping and multidimensional fault feature construction
[0103] Time window segmentation: To perform correlation analysis on simultaneously occurring alarms, the intelligent inspection server groups alarm events according to a fixed-length time window. The time window length can be set according to the inspection cycle; for example, if the inspection cycle is 5 minutes, the time window can be set to 10 minutes or 15 minutes. Each window contains all alarm events whose alarm times fall within that time period.
[0104] Alarm event association: For each time window, the system selects a set of all alarm events falling within that window from the alarm event database. For each alarm event in the set, it associates it with the device ID and operating data, device information model, network topology model, and wind farm / monitoring unit information, including:
[0105] The device ID is used to obtain the device type, network role, wind farm to which it belongs, monitoring unit to which it belongs, and the location of the equipment room and cabinet in the device information model;
[0106] Retrieve the topology level, upstream device set, and downstream device set of a device in the topology model using the device ID;
[0107] Query the device's operating data in the operating data table by timestamp (e.g., one collection cycle before or after) to obtain the changes in indicators such as CPU utilization, port traffic, packet loss rate, and latency before and after the alarm.
[0108] Multidimensional Fault Feature Construction: The system constructs multidimensional fault features for each device within the current time window. The feature vector can include the following:
[0109] Time characteristics: the time of the first alarm, the time of the last alarm, the duration of the alarm, and the number of alarms within the window;
[0110] Topology location characteristics: the topology level of the device (core, aggregation, access, etc.), the number of upstream devices, the number of downstream devices, and the proportion of alarms occurring among upstream and downstream devices;
[0111] Device attribute characteristics: device type (switch, firewall, server, etc.), network role (core layer, access layer, etc.), whether it is a single point of critical device;
[0112] Performance characteristics: maximum, average, and sudden increase in CPU / memory / port traffic within the window; whether packet loss rate and latency increase significantly, etc.
[0113] Site hierarchy characteristics: the wind farm to which the equipment belongs, the monitoring unit to which it belongs (such as which wind turbine monitoring unit, which booster station), and the type of business system to which it belongs.
[0114] Through the above processing, a set of multi-dimensional fault feature vectors based on equipment can be obtained within each time window, providing input for subsequent root cause analysis.
[0115] Step (4): Root Cause Analysis by the Rule Engine
[0116] Rule Engine Framework: The intelligent inspection server is internally configured with a fault location rule engine. Rules are stored in a structured manner, such as recording rule condition expressions and action expressions in a rule table. Rules include upstream link fault rules, performance bottleneck rules, power failure rules, etc., and different rules can be assigned different priorities.
[0117] Candidate root cause screening and scoring: For the set of alarm events within each time window, the system performs the following processing based on multi-dimensional fault characteristics and topology models:
[0118] (1) Filter the candidate set according to the severity of equipment condition.
[0119] The equipment status is classified into levels such as "normal, single indicator abnormal, multiple indicators severely abnormal, and unreachable";
[0120] Devices with a status of "unreachable" or "multiple indicators severely abnormal" were initially selected as candidate root cause devices.
[0121] (2) Analysis based on topology hierarchy and alarm distribution
[0122] If an upstream device malfunctions and most of its downstream devices also malfunction, while the upstream device further upstream is functioning normally, then according to the upstream link failure rules, the upstream device or the link between it and the upstream device is added to the candidate set.
[0123] If a device is located in the core layer or security boundary layer (such as a firewall or forward isolation device), set a higher topology layer weight for it.
[0124] (3) Scoring candidate root causes: For each device or link in the candidate set, calculate a comprehensive score, considering at least the following:
[0125] Equipment status severity rating: for example, unreachable scores the highest, followed by multiple abnormal indicators;
[0126] Topology hierarchy weights: Core devices and critical links are assigned higher weights;
[0127] Corresponding alarm levels and durations: Emergency alarms and alarms with long durations receive higher scores.
[0128] The scoring can be done using a linear weighting method.
[0129] Score = a × Severity score + b × Topology weight + c × Alarm level score + d × Duration score
[0130] Where a, b, c, and d are weighting coefficients set based on experience.
[0131] (4) Examples of specific rules for application
[0132] Upstream link failure rule: When a first switch B and its downstream access switches C and D are all unreachable, but its upstream core switch A is still reachable, and other branch devices connected to A but not through B are still reachable, the rule engine determines that the A-B link or device B is the root cause, and the failure type is link interruption or intermediate node failure.
[0133] Performance bottleneck fault rule: When the CPU utilization of a certain aggregation switch E is greater than 90% in the last 3 inspection cycles, and the latency of multiple downstream access switches and servers of E is significantly increased but still can be pinged, the rule engine determines that device E has a performance bottleneck and the fault type is performance congestion.
[0134] Based on the matching results of each rule and the comprehensive score, the system selects the device and / or link with the highest score from the candidate root cause set as the root cause device and / or root cause link of the alarm set in the current time window, and gives the fault type and confidence level.
[0135] Step (5): Calculation of the scope of influence and display of results
[0136] Determining the set of affected devices: After identifying the root cause device or root cause link, the intelligent inspection server, based on the network topology model, traverses all reachable devices downstream from the root cause device.
[0137] If a device has an alarm event within the current time window, or its operating data indicators show obvious abnormalities (such as a sudden increase in latency or an increase in packet loss rate), and its business path passes through the root cause device or root cause link, then add the device to the set of affected devices.
[0138] The traversal process can be limited to monitoring units that have business relationships with the root cause device, such as traversing down from the core switch of the booster station to the booster station server, front-end machine, etc.
[0139] Fault location result generation: The system organizes the relevant information of the root cause device (or root cause link) and affected devices into structured fault location results, including:
[0140] Fault location: wind farm, monitoring unit, equipment room, cabinet, equipment name, and port information;
[0141] Fault types: link interruption, equipment failure, performance bottleneck, power failure, etc.
[0142] Scope of impact: List of affected network devices, corresponding wind turbine units or substation units, and potentially affected business systems;
[0143] Time of fault occurrence, time of detection, and current status;
[0144] Confidence level of fault location results.
[0145] Topology view visualization: In the topology view interface of the operations and maintenance terminal:
[0146] Highlight the root cause device and its associated link in a striking color (such as red);
[0147] The affected device set is marked with a different color (such as orange);
[0148] Display the fault type, brief description, and alarm level next to the device or link icon;
[0149] Maintenance personnel can click on any device to view its operating data curves and alarm records during the current fault event.
[0150] Result writing and alarm merging: The intelligent inspection server writes the identifier of the root cause device or root cause link and the fault location result into the alarm event database, and marks multiple alarms caused by the same root cause within the same time window with a unified root cause identifier. When displaying the alarm list on the operation and maintenance terminal, it can merge and display alarms according to the root cause dimension, that is, fold multiple alarms corresponding to the same root cause into a single comprehensive alarm, reducing the number of alarms and making it easier for operation and maintenance personnel to quickly focus on key issues.
[0151] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0152] The above description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for accurate fault location of wind farm network equipment based on multi-data fusion, applied to a monitoring system including a wind farm production control network, characterized in that, include: (1) Obtain network equipment information and inter-equipment connection relationships in the wind farm production control network, and establish equipment information model and network topology model; (2) Collect operating data from the network device according to a preset period, and generate alarm events associated with the network device based on a preset threshold and / or historical operating data, and store them in the alarm event database; (3) Group alarm events within the same time window according to the time window, and associate each alarm event with the corresponding operating data, the equipment information model, the network topology model and the wind farm and monitoring unit information to form multi-dimensional fault characteristics for each network device; (4) Input the multidimensional fault characteristics and network topology relationship into the fault location rule engine, analyze the alarm event groups in each time window, and determine the root cause device and / or root cause link and fault type according to the pre-configured fault location rules. (5) Based on the network topology model, determine the set of affected network devices from the root cause device and / or root cause link in the downstream direction, map the root cause device and / or root cause link and the affected network devices to the wind farm network topology view and output the fault location result containing the fault location, fault type and impact range to the operation and maintenance terminal.
2. The method as described in claim 1, characterized in that, The operational data collected in step (2) includes at least device connectivity indicators, operational performance indicators and device status information. The device connectivity indicators include at least one of packet loss rate and latency, and the operational performance indicators include at least one of CPU utilization, memory utilization, storage space utilization and port traffic.
3. The method as described in claim 1 or 2, characterized in that, In step (2), an alarm event is generated by combining static thresholds and dynamic baselines. The dynamic baseline is calculated based on historical operating data within a preset statistical period. An alarm event is triggered when the current operating data deviates from the dynamic baseline by more than a preset multiple and continues for a preset time.
4. The method as described in claim 3, characterized in that, The multidimensional fault features mentioned in step (3) include at least the following: time features representing the time of alarm occurrence and duration, topological location features representing the hierarchy and upstream and downstream relationships of network devices in the topology, device attribute features representing device type and network role, operational performance features representing the degree of abnormality of operational indicators, and site hierarchy features representing the wind farm and monitoring unit to which the device belongs.
5. The method as described in claim 4, characterized in that, In step (4), the fault location rule engine scores the candidate root cause devices and / or candidate root cause links. The scoring is based at least on the severity of the device status, the topology level weight, and the alarm level and duration of the corresponding alarm event, and selects the candidate with the highest score as the root cause device and / or root cause link.
6. The method as described in claim 5, characterized in that, The fault location rules include at least upstream link fault rules, which are used to determine a network device and / or its connected links as the root cause when the following conditions are met: the network device and its direct downstream network devices are in an unreachable state, while other network devices connected to its upstream network devices but not through the network devices are in a reachable state.
7. The method as described in claim 6, characterized in that, The fault location rules also include device power failure rules, which are used to determine a target network device as the root cause when the following conditions are met: the power status of the target network device is abnormal and unreachable, and the power status of other network devices in the same cabinet is normal.
8. The method as described in claim 7, characterized in that, The fault location rules also include performance bottleneck fault rules, which are used to determine a target network device as the root cause when the following conditions are met: the CPU utilization and / or port bandwidth utilization of the target network device continuously exceed a preset threshold in multiple collection cycles, and the latency of multiple downstream network devices connected to the target network device increases significantly but connectivity remains available.
9. The method as described in claim 8, characterized in that, Step (5) involves determining the set of affected network devices by traversing downstream of the network topology from the root cause device and adding network devices that have alarm events and / or abnormal operating data within the current time window and whose service paths pass through the root cause device and / or root cause link to the set of affected network devices.
10. The method as described in claim 9, characterized in that, Also includes: The identifiers of the root cause device and / or root cause link, along with the fault location results, are written into the alarm event database and associated with the corresponding alarm events, so that multiple alarms caused by the same root cause can be merged and displayed in the operation and maintenance terminal.