Fault diagnosis and adaptive reconstruction method for communication network of power distribution network

By monitoring the status of the distribution network communication network, extracting key indicators, formulating hierarchical partitioning strategies, identifying and isolating faults, formulating and verifying reconstruction plans, and dynamically adjusting reconstruction intervals, the real-time, accuracy and adaptive challenges of fault diagnosis and adaptive reconstruction in the distribution network communication network are solved, and network stability and reliability are improved.

CN120050159APending Publication Date: 2025-05-27FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID +1

Patent Information

Application Number
CN202510139818.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

There are real-time, accuracy and adaptive challenges in fault diagnosis and adaptive reconstruction in the distribution network communication network, and the impact of reconstruction on the business needs to be minimized, and seamless coordination between diagnosis and reconstruction is required.

Method used

By monitoring network status, extracting key performance indicators, formulating hierarchical partitioning strategies, identifying and isolating faults, formulating reconstruction plans, updating network parameters in real time, verifying reconstruction performance, and dynamically adjusting reconstruction intervals.

Benefits of technology

It improves the stability and reliability of the distribution network, optimizes network performance, ensures efficient and safe operation of the power system, and realizes efficient and accurate fault diagnosis and adaptive reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050159A_ABST
    Figure CN120050159A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution network communication network fault diagnosis and self-adaptive reconstruction method, which comprises the steps of formulating corresponding layering and partitioning strategies aiming at different network scales and service types, including modular management of a large-scale multi-layer network and centralized response of a small-scale network; the faults of various networks can be effectively positioned and isolated; recognizing a reconstruction object after the fault positioning is completed, and formulating a network reconstruction scheme according to the recognized reconstruction object and the current state of the network, including but not limited to replacing a fault node, reconstructing a connection or modifying a route by starting a standby resource; after the reconstruction scheme is executed, the data collection and analysis period is adjusted according to the fault frequency, and the reconstruction interval is controlled to reduce redundant network structure adjustment and potential network fluctuation, including active detection and passive monitoring of the power distribution network and bidirectional monitoring of the network condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a method for fault diagnosis and adaptive reconstruction of a distribution network communication network. Background Art

[0002] Although fault diagnosis and adaptive reconstruction in a distribution network communication network play an important role in improving network robustness, there are still many technical challenges and contradictions in practical applications. Fault diagnosis needs to quickly discover abnormal patterns in a large amount of network data and accurately locate the fault point, which poses extremely high requirements for the real-time performance and accuracy of the diagnosis algorithm. At the same time, the diagnosis algorithm also needs to adapt to the diversity and dynamics of the distribution network communication network and be able to work stably under different network architectures, protocol stacks, and service scenarios. Secondly, adaptive reconstruction needs to minimize the impact on services while ensuring network continuity. This requires the reconstruction to accurately evaluate the fault impact range and quickly converge to the optimal network topology at the lowest cost. In addition, the frequency and time of the reconstruction process also need to be dynamically adjusted according to the tolerance of the service, avoiding network oscillations caused by overly frequent reconstructions and service interruptions caused by reconstruction lags. How to achieve efficient and accurate fault diagnosis and adaptive reconstruction in a distribution network communication network, while ensuring the real-time performance, adaptability, minimizing service impact, and realizing seamless coordination between the diagnosis and reconstruction links, is an urgent problem to be solved. Summary of the Invention

[0003] The present invention provides a method for fault diagnosis and adaptive reconstruction of a distribution network communication network, mainly including:

[0004] Monitor the distribution network communication network, obtain network status information, extract key network performance indicators to establish a baseline for normal network operation, and identify potential abnormal fluctuations, including hardware failures, software failures, or configuration errors, by comparing the current network status with the baseline in real time;

[0005] Develop corresponding hierarchical partitioning strategies for different network scales and service types, including modular management for large-scale multi-level networks and centralized response methods for small-scale networks, to effectively locate and isolate faults in various networks;

[0006] Screen according to preset fault modes, predict and identify complex or unknown faults, centrally locate the fault range to refine the fault impact range, and determine the specific location where the fault occurs, including nodes or connection points, through analysis of the network topology structure and fault propagation mode, so as to implement distributed isolation for the local distribution network;

[0007] After the fault location is completed, identify the reconstruction object, and formulate a network reconstruction plan based on the identified reconstruction object and the current network state, including but not limited to replacing the faulty node by starting the standby resource, reconstructing the connection or modifying the route;

[0008] Update the network parameters in real time, including adjusting the channel allocation and power control, suppressing the new network conflicts and performance degradation caused by the reconstruction, testing the performance of the reconstructed network through a network simulator, verifying whether the adjusted configuration meets the design requirements, and execute the reconstruction plan on the premise of meeting the real-time and reliability requirements;

[0009] After executing the reconstruction plan, adjust the data collection and analysis period according to the fault frequency, and control the reconstruction interval to reduce the redundant network structure adjustment and potential network fluctuations, including actively detecting and passively listening to the distribution network, and performing two-way monitoring of the network status.

[0010] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0011] The present invention discloses a fault diagnosis and adaptive reconstruction method for a distribution network communication network. The method obtains network status information, extracts key network performance indicators, establishes a baseline for normal network operation, and identifies potential abnormal fluctuations by comparing the current network state with the baseline in real time; different methods are adopted for different-scale networks, which helps to effectively locate and isolate faults in various networks; this application improves the stability and reliability of the distribution network, optimizes the network performance, and ensures the efficient and safe operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flowchart of a fault diagnosis and adaptive reconstruction method for a distribution network communication network of the present invention.

[0013] Figure 2 It is a schematic diagram of a fault diagnosis and adaptive reconstruction method for a distribution network communication network of the present invention.

[0014] Figure 3 It is another schematic diagram of a fault diagnosis and adaptive reconstruction method for a distribution network communication network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0016] Such as Figures 1-3 , a fault diagnosis and adaptive reconstruction method for a distribution network communication network in this embodiment may specifically include:

[0017] Step S101, monitor the communication network of the distribution network, obtain network status information, extract key network performance indicators to establish a baseline for the normal operation of the network; by comparing the current network status with the baseline in real time, identify potential abnormal fluctuations, including hardware failures, software failures, or configuration errors.

[0018] Obtain the status information of the communication network of the distribution network, conduct correlation analysis on the network performance indicators to obtain the correlation and influence rules between different indicators; based on the correlation analysis results of the network performance indicators, construct a network anomaly detection model, train and optimize the network anomaly detection model; deploy the optimized network anomaly detection model to network monitoring, analyze the network status information in real time, and when an abnormal fluctuation of the network performance indicators is detected, trigger an alarm and notify the network management personnel to conduct a troubleshooting and handling; represent and store the network fault diagnosis results and handling experiences in a structured manner to form a fault diagnosis knowledge base for the communication network of the distribution network. When a network anomaly occurs, assist the network management personnel in fault location and handling by retrieving relevant cases.

[0019] Exemplarily, a network monitoring device is used to comprehensively monitor the distribution network communication network, obtain network status information, and extract key network performance indicators to establish a baseline for the normal operation of the network. By comparing the current network status with the baseline in real time, potential abnormal fluctuations are identified, including potential problems such as hardware failures, software failures, or configuration errors. According to the distribution network communication network status information obtained by the network monitoring device, methods such as association rule mining and time series analysis are used to deeply mine and conduct correlation analysis on network performance indicators to determine the correlation and influence laws between different indicators. Through the correlation analysis of network performance indicators, key indicators that have a significant impact on the abnormal network status are identified, and these key indicators are used as input features for the anomaly detection model. At the same time, according to the correlation analysis results, the weight relationship between different indicators is determined to guide the feature selection and parameter adjustment of the anomaly detection model. Based on the correlation analysis results, a network anomaly detection model is constructed, and machine learning algorithms such as support vector machine (SVM), random forest (RandomForest), or neural network (NeuralNetwork) are selected. Taking the random forest as an example, multiple decision trees are constructed by randomly selecting samples and features, and through the method of ensemble learning, the prediction results of multiple decision trees are integrated to improve the accuracy and robustness of anomaly detection. During the training process of the anomaly detection model, supervised learning is carried out using the labeled historical network status data to continuously optimize the parameters of the anomaly detection model and improve the generalization ability of the model. The optimized network anomaly detection model is deployed to network monitoring to analyze network status data in real time. When abnormal fluctuations in network performance indicators are detected, an alarm is triggered in a timely manner to notify network administrators to conduct inspections and handling. The anomaly detection results are applied to fault diagnosis, and the network fault diagnosis and handling experience are structurally represented and stored to form a fault diagnosis knowledge base for the distribution network communication network. When constructing the fault diagnosis knowledge base, the core concepts and relationships in the fault diagnosis field are first defined, such as fault types, fault causes, fault phenomena, diagnostic steps, etc. Then, entities and relationships are extracted from historical fault handling records to construct a knowledge graph. In the knowledge graph, fault instances are used as nodes and fault attributes are used as edges to form a structured knowledge representation. When a new network anomaly occurs, relevant fault cases are quickly retrieved through semantic search and reasoning in the knowledge graph, and network administrators are guided to locate and handle faults according to the diagnostic steps and solutions in the cases. At the same time, after dealing with a new fault case, it is added to the knowledge graph to continuously expand and improve the knowledge base. Through continuous network monitoring and anomaly analysis, the network anomaly detection model and the fault diagnosis knowledge base are continuously improved and updated. The anomaly detection model can be fine-tuned and optimized by periodically retraining and evaluating it using new network status data to improve the adaptability and accuracy of the model.The fault diagnosis knowledge base can expand the knowledge coverage, improve the efficiency and accuracy of fault diagnosis by continuously accumulating new fault cases and handling experiences. Through the collaborative optimization of anomaly detection and fault diagnosis, a self-learning and optimization mechanism for the distribution network communication network is formed to improve the intelligent level of network operation and maintenance and ensure the stable operation of the distribution network communication network. In the monitoring of the distribution network communication network, network monitoring devices such as network probes and traffic analyzers are deployed to collect key performance indicators such as the CPU utilization rate, memory occupancy rate, network delay, and packet loss rate of network devices at a sampling frequency of 100 times per second. By statistically analyzing the network status data for 30 consecutive days, the mean and standard deviation of each indicator are calculated to establish a baseline for normal network operation. During real-time monitoring, anomaly detection algorithms such as IsolationForest or One-ClassSVM are used to slide and analyze the difference between the current network status and the baseline with a 30-second time window. When it is detected that a certain indicator deviates from the baseline by more than 3 standard deviations, an anomaly alarm is triggered to notify the network administrator for inspection. At the same time, association rule mining algorithms such as Apriori or FP-growth are used to perform association analysis on network performance indicators to discover frequent patterns and association rules between indicators. For example, when the network delay exceeds 100 ms, the packet loss rate will also exceed 1%. The results of association analysis are used to guide the feature selection of the anomaly detection model to improve the accuracy of anomaly detection. When constructing the anomaly detection model, the random forest algorithm is selected. By randomly selecting 80% of the sample data and sqrt(n) features, 100 decision trees are generated, and combined with a voting mechanism, the network status is classified and predicted. During the model training process, 5-fold cross-validation is used to evaluate the accuracy, recall rate, and F1 value of the model, and the model hyperparameters such as the maximum depth of the tree and the minimum number of samples in the leaf nodes are optimized through grid search. When constructing the fault diagnosis knowledge base, 20 core concepts such as router failures and fiber optic interruptions and 30 relationships such as fault causes and influence scopes are extracted from 500 historical fault handling records to construct a knowledge graph containing 500 nodes and 1000 edges. Through graph-based shortest path algorithms such as the Dijkstra algorithm, rapid retrieval and reasoning of fault cases are realized, and the average retrieval time is less than...

[0020] Step S102: For different network scales and service types, formulate corresponding hierarchical and partition strategies, including modular management for large-scale multi-level networks and centralized response methods for small-scale networks, so as to effectively locate and isolate faults in various networks.

[0021] According to different network scales and service types, identify the devices, links, and interconnection relationships in the distribution network communication network to generate a network topology diagram; by analyzing the connectivity and device attributes in the network topology diagram,

[0022] Combine the service traffic to divide different management domains and subnets; for large-scale multi-level networks, adopt modular management; for small-scale networks, implement centralized management; trigger an alarm when a certain indicator deviates from the baseline range, including analyzing the alarm data, locating the fault point, and locating the fault point includes analyzing the suspected fault points to narrow down the troubleshooting scope; according to the fault location result, isolate the fault area to prevent the fault from spreading further; match the predefined typical fault scenarios and disposal plans, give repair suggestions, and guide the administrator to handle the fault.

[0023] Exemplarily, according to the network scale and service type, network topology discovery technology is adopted to automatically identify the devices, links and interconnection relationships in the network, generate a complete network topology map, and lay a foundation for hierarchical partitioning. By analyzing the connectivity and device attributes in the network topology map, combined with service requirements and management objectives, different management domains and subnets are divided. For large-scale multi-level networks, a modular management strategy is adopted to divide the network into multiple independent management domains, and autonomous management is realized within each management domain. Between management domains, by defining a data exchange format based on XML or JSON and using standardized interfaces such as RESTful API or SOAP for data transmission, cross-domain information interaction is achieved. At the same time, by deploying a distributed resource management system such as Apache Mesos or Kubernetes, cross-domain resource abstraction, allocation and scheduling are realized; within each management domain, subnets are further divided according to the network scale and service characteristics. For key subnets such as the core layer and aggregation layer, a highly reliable design is adopted, redundant devices and links are introduced to ensure that a single point of failure will not affect the communication of the entire subnet. For edge subnets such as the access layer, flexible management strategies are implemented according to service requirements and user distribution. For subnets with high requirements for service real-time performance, such as voice and video services, low-latency and high-bandwidth transmission technologies are adopted, and QoS strategies are configured, such as the Differentiated Services (DiffServ) mechanism based on traffic marking and priority mapping, to ensure the priority transmission of key services. For subnets with ordinary data services, a best-effort transmission strategy is adopted to reduce management complexity and costs. Different-scale networks need to adopt different management strategies. For large-scale networks, hierarchical partitioning and modular management are used to cope with complexity and heterogeneity; while for small-scale networks, a relatively simplified centralized management method can be adopted. A complete network baseline is established. Through the device operation status data collected in real time, such as key indicators such as CPU utilization rate, memory occupancy rate, traffic, and latency, the network performance is comprehensively perceived. When it is monitored that a certain indicator deviates from the baseline range, an alarm mechanism is triggered, and a rule engine based on CEP (Complex Event Processing), such as Esper or Drools, is used to analyze the alarm event stream in real time. By defining failure modes and association rules, potential root causes of failures and the scope of influence are identified, and the failure information is presented to the administrator through a visual interface to assist in decision-making. At the same time, an automated inspection tool is called, adopting a distributed architecture based on Agents. Lightweight collection Agents are deployed on network devices to regularly perform tasks such as connectivity tests and configuration verification, and the results are summarized to the central server for analysis. Common network inspection tools include Nagios, Zabbix, etc., which can be selected and customized according to the network scale and management requirements to narrow the scope of fault troubleshooting. According to the fault location results, relevant policy configurations such as routing and ACL are dynamically adjusted to quickly isolate the fault area and prevent the fault from spreading further.Match predefined typical fault scenarios and disposal plans, give repair suggestions, and guide the administrator to handle faults. After the handling is completed, trace back the cause of the fault, optimize the network architecture and management strategies, such as adjusting the topology structure, upgrading the device software version, improving the security protection measures, etc., to improve the reliability and maintainability of the network and reduce the recurrence of similar faults. During the network topology discovery process, adopt a detection method based on the SNMP protocol. By regularly polling the MIB libraries of network devices, such as IF-MIB, IP-MIB, etc., with a period of 30 seconds, collect key information such as link status, IP addresses, and device names, and use the Dijkstra shortest path algorithm to calculate the connectivity between devices and generate a network topology map. Use the community discovery algorithm to divide the network into different management domains and subnets, such as the core layer, aggregation layer, access layer, etc., according to attributes such as the physical location of devices, VLAN configuration, and routing protocols. Each management domain contains no more than 500 network devices. Considering that the LPA algorithm can adaptively spread labels to form communities without presetting the number of communities, this makes it suitable for initial exploratory analysis. By regarding network devices as nodes and the connections between devices, such as physical proximity, VLAN membership, routing interaction, etc., as edges, the LPA algorithm can be applied to discover naturally formed management domains. First, use LPA for preliminary community division to obtain a general outline of the network structure. Then apply the GN algorithm or modularity optimization methods, such as Louvain, to further refine, identify and strengthen the boundaries of the core, aggregation, and access layers, ensure that the number of devices in each management domain does not exceed 500, and at the same time keep the network efficient and logically clear. Evaluate the output results, such as measuring the quality of the community structure by calculating the modularity Q value, and manually adjust the division results according to the actual network operation and maintenance requirements if necessary. Use an XML-based data format to synchronize the configuration data and status information between management domains at a frequency of 100 times per minute through the RESTful API, and use Apache Mesos to implement cross-domain resource scheduling. With the goal of increasing the resource utilization rate by 20%, dynamically allocate and recycle computing, storage and other resources. Regarding the high reliability of key subnets, deploy core switches and aggregation switches in a primary and standby mode, and implement device-level redundant backup through the VRRP protocol to ensure that when a single point of failure occurs, the backup device can complete the switch within 100 milliseconds. At the same time, implement port authentication based on the 802.1X protocol and MAC address binding to strictly control access to the access terminals, and deploy intrusion detection to monitor network traffic in real time and identify potential attack behaviors. In terms of the QoS policy configuration of the edge subnet, adopt a differential service mechanism based on traffic marking and priority mapping to preferentially forward real-time service traffic such as voice and video, and ensure that the end-to-end delay is less than 50 milliseconds and the jitter is less than 20 milliseconds.Meanwhile, for different types of data services, corresponding queue scheduling strategies are configured. For example, the Low Latency Queue (LLQ) is adopted for critical applications, and the Weighted Fair Queue (WFQ) is adopted for ordinary services. In terms of network baseline management, unified management of the distribution network is carried out. Key indicators such as CPU utilization rate, memory occupancy rate, interface traffic, and packet loss rate of network devices are collected every minute. Using the Statistical Process Control (SPC) method and combining with Six Sigma theory, quantitative evaluation and trend prediction of network performance indicators are carried out to timely detect abnormal fluctuations and potential risks. When a certain indicator deviates from the baseline range by more than 3 standard deviations, a yellow alarm is triggered; when it exceeds 6 standard deviations, a red alarm is triggered to notify the administrator to intervene and handle. An association analysis algorithm based on decision trees is adopted. By analyzing a large amount of alarm log and performance indicator data, fault modes and root causes are mined to generate a diagnostic rule library. When a fault occurs, through rule matching and reasoning, a preliminary judgment on the fault cause and impact scope is given within 5 minutes to assist the administrator in quickly locating and isolating the fault. Meanwhile, an automated inspection tool is called to perform connectivity and reachability detection on the devices of the entire network. ICMP (Internet Control Message Protocol) and SNMP (Simple Network Management Protocol) are used as detection means, and then a network health status report is generated to narrow down the scope of investigation for potential fault locations. Combining with the historical fault case library and expert experience, similar fault scenarios are matched, and repair operation suggestions are given to guide the administrator in handling the fault.

[0024] Step S103, screen according to the preset fault mode, predict and identify complex or unknown faults, centrally locate the fault scope to precisely define the fault impact scope, and through the analysis of the network topology structure and fault propagation mode, determine the specific location where the fault occurs, including nodes or connection points, so as to implement distributed isolation for the local distribution network.

[0025] A fault knowledge base is established, including typical fault scenarios of hardware faults, software faults, and configuration errors, and the characteristic attributes and discrimination rules of each fault mode are defined to form a fault knowledge base; historical fault data is trained to establish a fault prediction model, and by extracting the change trends and abnormal patterns of key indicators before the fault occurs, potential faults are predicted and early warnings are given; for complex or unknown faults, through the correlation analysis of fault symptoms and impact scope, by reasoning and tracing the fault propagation path on the network topology diagram, the root cause and key nodes of the fault are identified; the fault alarm information and abnormal indicator data are summarized and analyzed to determine the location where the fault occurs, and the impact degree of the fault on network performance and business continuity is evaluated to generate a fault impact report and handling suggestions; according to the network topology structure and the logical connection relationship between devices, the fault impact scope is divided, and the specific location where the fault occurs is determined. The specific location where the fault occurs includes faulty devices, ports, and links.

[0026] Exemplarily, based on historical fault data and expert experience, a preset fault mode library is established, including typical fault scenarios such as common hardware faults, software faults, configuration errors, etc., and the characteristic attributes and discrimination rules of each fault mode are defined, such as the device type where the fault occurs, alarm information, abnormal indicators, etc., to form a structured fault knowledge base, providing a basis for fault screening and matching. Constructing a structured fault knowledge base includes collecting and sorting historical fault data, including information such as the fault occurrence time, faulty device, fault phenomenon, fault cause, etc., cleaning and annotating these data to form a fault case library. Secondly, analyze and summarize the fault cases, extract the common characteristics and laws of the faults, such as the device type where the fault occurs, fault mode, fault trigger conditions, etc., to form a fault feature vector. Then, according to the fault feature vector, design a fault knowledge representation method, such as a fault tree, a fault map, etc., to structurally represent and store the knowledge in the fault cases, and construct the basic framework of the fault knowledge base. Use machine learning algorithms, such as support vector machine (SVM), random forest, etc., to train the historical fault data, establish a fault prediction model, and realize the prediction and early warning of potential faults by extracting the change trend and abnormal pattern of key indicators before the fault occurs, improving the accuracy and real-time performance of fault identification. At the same time, for complex or unknown faults, through the correlation analysis of fault symptoms and the scope of influence, using graph theory algorithms, infer and trace the fault propagation path on the network topology map, identify the root cause and key nodes of the fault, and provide decision-making support for fault location. When inferring and tracing the fault propagation path, the following graph theory algorithms can be used, including mapping the network topology relationship into a directed weighted graph, taking devices as nodes, taking the connections between devices as edges, and setting the weight of the edge as the probability or influence degree of fault propagation; using the shortest path algorithm, such as Dijkstra algorithm or Floyd-Warshall algorithm, to calculate the shortest path from the fault node to other nodes and find out the key path of fault propagation. At the same time, using the critical path algorithm, such as the critical path method (CPM) or the critical chain method (CCM), identify the key nodes and key activities that affect fault propagation, and determine the key scope of influence of the fault. In addition, community discovery algorithms, such as the Louvain algorithm or the Infomap algorithm, can also be used to divide the network topology into different communities or modules, analyze the propagation characteristics of faults within and between different communities, and predict the scope of influence of faults. Combining the correlation analysis results of fault symptoms and the scope of influence and the analysis results of the propagation characteristics of faults within and between different communities, generate a visual map of fault propagation, intuitively display the fault propagation path, scope of influence and key nodes, and provide a decision-making basis for fault diagnosis and disposal.Adopt a centralized fault location method. After aggregating fault alarm information and abnormal index data, conduct comprehensive analysis and judgment through a rule engine and an expert system to locate the exact location where the fault occurs, evaluate the impact of the fault on network performance and service continuity, and generate a fault impact report and disposal suggestions. The comprehensive analysis and judgment through the rule engine and the expert system include extracting the empirical knowledge and rules of fault diagnosis to form a structured knowledge base, such as a fault - cause knowledge base, a fault - solution knowledge base, etc. Then, use a rule representation language, such as IF - THEN rules, decision trees, etc., to transform the knowledge in the fault diagnosis knowledge base into executable rules and build a rule base. During fault diagnosis, the fault information and monitoring data are used as inputs to trigger the rules in the rule engine. Through forward reasoning or backward reasoning, automatically match and execute the corresponding diagnostic rules to obtain the fault cause and disposal suggestions. According to the results of fault location, adopt network topology analysis technology to accurately judge and divide the fault impact scope. By analyzing the network topology structure and the logical connection relationship between devices, construct a network connectivity graph, and use graph search algorithms, such as depth - first search (DFS), breadth - first search (BFS), etc., to traverse the connected area of the faulty device, identify the affected devices, ports, and links, and form a topological view of the fault impact scope to provide a precise target area for fault isolation. In the fault diagnosis of the distribution network, fully consider the radial topology structure and power supply mode of the distribution network and adopt a distributed fault isolation strategy. When a fault occurs, automatically divide the isolation area according to the fault location and impact scope, and realize the rapid segmentation of the fault area and the healthy area by controlling the actions of switches and circuit breakers to minimize the impact of the fault on the distribution network automation system. When constructing the fault knowledge base, extract 20 common fault modes from 500 historical typical fault cases, such as optical cable breakage, equipment overheating, etc., and summarize 10 key features of each mode to form a 200 - dimensional fault feature vector. Adopt a fault knowledge representation method based on decision trees to map the fault feature vector to the non - leaf nodes of the decision tree, and map the fault cause and disposal plan to the leaf nodes to construct a fault decision tree containing 500 nodes and 2000 edges. When conducting fault prediction, select 100 key nodes in the network as monitoring objects, collect KPI indicators every 5 minutes, and use an anomaly detection algorithm, such as One - Class SVM, to judge whether the KPI indicators exceed the normal range. When an anomaly is detected, use the abnormal indicators as inputs and use a random forest algorithm to judge the probability of a fault occurring. If the fault probability exceeds 80%, trigger an early warning to detect potential faults 5 to 10 minutes in advance. After a fault occurs, through network topology analysis, construct a network connectivity graph containing 500 nodes and 1000 edges, and use the Dijkstra shortest path algorithm to calculate the shortest path from the faulty node to other nodes to generate a key path map of fault propagation.Meanwhile, through the Critical Path Method (CPM), identify the key nodes and bottleneck links for fault propagation, predict the impact scope of the fault within the next 1 hour, and use the Louvain community detection algorithm to divide 5 key regions for fault propagation. When performing fault diagnosis, randomly extract 100 fault diagnosis rules from the knowledge base, and through forward reasoning, quickly lock the fault cause based on the device status and alarm information, with an average diagnosis time of less than 1 minute. Then, use the Apriori-based association rule mining algorithm to extract 20 most relevant fault solutions from 1000 historical fault handling records, and combine with the reasoning mechanism in the expert system to optimize and combine the solutions to form a comprehensive fault handling strategy. Generate a fault analysis report and present it to the operation and maintenance personnel in a visual way to assist them in making fault decisions. When performing fault isolation, according to the topology of the distribution network, use the depth-first search algorithm to quickly search for the primary connected domain of the fault node to form a fault isolation area. At the same time, use the greedy algorithm to find the optimal action plan for the disconnector, minimizing the power outage scope while ensuring fault isolation. Through power flow calculation and optimization algorithms, considering constraints such as line capacity and load importance, optimize the network reconstruction plan to make the network after fault isolation as much as possible meet the N-1 security criterion and maximize the recovery degree of important loads. Through the evaluation algorithm, quantify the impact of the fault on the system, calculate the network performance indicators before and after the fault occurs, such as SAIDI, SAIFI, etc., evaluate the effectiveness of the fault isolation and recovery plan, and generate a fault handling evaluation report.

[0027] Step S104, when the fault location is completed, identify the reconstruction object, and formulate a network reconstruction plan according to the identified reconstruction object and the current network state, including but not limited to replacing the fault node by starting the standby resource, reconstructing the connection or modifying the routing.

[0028] After determining the specific location of the fault, identify the reconstruction object, including analyzing the type, location, and functional attributes of the fault node, as well as the status of the links and adjacent nodes connected to it; based on the role and importance of the fault node in the network, judge whether reconstruction is required; calculate the connectivity and availability of the network topology, evaluate the impact degree of the fault on the network service quality, and according to the change of the network performance indicators, formulate a reconstruction plan, including: selecting a standby node to replace the fault node according to the configuration and real-time status of the standby resource, and adjusting the number and distribution of standby nodes according to the requirements of load balancing and disaster tolerance protection; when reconstructing the connection of the fault node, perform link reconstruction and traffic switching between the fault node and the adjacent nodes; adjust the forwarding path and priority of the data packet according to the service quality requirements of different services and the real-time status of network resources.

[0029] Exemplarily, after fault location is completed, the reconfiguration object recognition process is immediately triggered. By analyzing the attributes of the fault node such as its type, location, function, etc., as well as the status of the connected links and adjacent nodes, the role and importance of the fault node in the network are automatically judged. And according to the pre-configured reconfiguration strategy, it is determined whether to start the network reconfiguration process, as well as the priority and target of the reconfiguration. Graph theory algorithms and simulation technologies are used to model and analyze the network topology. When modeling and analyzing the network topology, first, through depth-first search (DFS) or breadth-first search (BFS) algorithms, the network topology graph is traversed, the fault nodes and affected links are marked, and the connectivity and availability of the network are determined. Then, a minimum spanning tree algorithm, such as Kruskal's algorithm or Prim's algorithm, is used to construct a minimum-cost spanning tree that connects all nodes among the remaining available nodes and links as the backbone topology for network reconfiguration. Next, a shortest path algorithm, such as Dijkstra's algorithm or Bellman-Ford algorithm, is used to calculate the shortest paths between each pair of nodes on the backbone topology, and the performance metrics such as network delay and jitter are evaluated according to parameters such as path length and link bandwidth. Finally, through a network flow algorithm, such as the maximum flow minimum cut algorithm, the critical links and bottleneck resources in the network are analyzed to optimize the network load balancing and resource utilization. When conducting simulation verification, professional network simulation tools such as NS-3 and OPNET are used to build a network topology model, set the attribute parameters of nodes and links, simulate the network behavior under fault scenarios, and evaluate the feasibility and effectiveness of the reconfiguration plan through the simulation results. After a fault occurs, quickly calculate the connectivity and availability of the network, evaluate the impact degree of the fault on the network service quality, and dynamically adjust the reconfiguration plan according to the change of network performance metrics to ensure the rapid recovery of the network and the stability of the service quality. According to the configuration and real-time status of the backup resources, through a resource scheduling algorithm, the backup nodes with high similarity to the fault node and low switching cost are preferentially selected as replacements, and the number and distribution of the backup nodes are dynamically adjusted according to the requirements of load balancing and disaster tolerance protection to improve the redundancy and survivability of the network. The intelligent resource scheduling algorithm can be implemented based on the following ideas: First, collect the resource usage conditions of each node in the network, including CPU utilization rate, memory occupancy, storage space, etc., to form a resource status matrix. Then, use a clustering algorithm, such as the K-means algorithm or hierarchical clustering algorithm, to divide the resource status matrix into several resource utilization patterns, and each pattern represents a class of similar resource usage behaviors. Next, use an association rule mining algorithm, such as the Apriori algorithm or FP-growth algorithm, to discover the association relationships and frequent item sets between different resource utilization patterns for predicting the resource demand trends of nodes. Then, combined with factors such as the business importance and historical reliability of nodes, a resource scheduling optimization model is constructed, and the goal is to minimize the overhead and delay of resource scheduling under the premise of meeting business continuity.Finally, heuristic search algorithms such as genetic algorithms and ant colony algorithms are used to solve the optimal solution for resource scheduling, dynamically allocate and adjust spare resources, and realize intelligent resource scheduling. When rebuilding the connection of the faulty node and modifying the route, the flow table delivery technology based on software defined network (SDN) is adopted. The centralized controller quickly delivers forwarding rules to achieve link reconstruction and traffic switching between the faulty node and the adjacent node, ensuring that the network can converge and recover quickly when a fault occurs. The routes in the network are optimized and dynamically adjusted in real time. According to the service quality requirements of different businesses and the real-time status of network resources, the forwarding path and priority of the data packet are adaptively adjusted to reduce network congestion and delay, and improve the transmission efficiency and reliability of the network. When reconstructing object recognition, the key attributes in the network topology, such as the degree centrality and betweenness centrality of the node, are first extracted and used as input features. The node importance evaluation model is trained using the random forest algorithm. Through learning 2000 historical fault cases, 95% node role discrimination accuracy is achieved on the test set. At the same time, community discovery algorithms, such as label propagation algorithms, are used to divide the network into 10 functional areas, and the proportion of faulty nodes in each area is calculated. If the proportion exceeds 5%, the area is automatically included in the reconstruction scope. In the simulation verification phase, a 20% node failure rate is set for a typical network topology of 300 nodes and 500 links. Through 500 Monte Carlo simulations, the effect of the reconstruction scheme on network performance improvement is evaluated. The results show that the average fault recovery time is shortened from the original 10 minutes to less than 2 minutes, and the network throughput and latency are optimized by 30% and 50% respectively. In terms of intelligent resource scheduling, six resource indicators such as CPU and memory of 500 nodes in the network are collected every 5 minutes to form a 500x6 resource status matrix. The K-means algorithm is used to cluster them into 5 resource utilization modes, and the transition probabilities between different modes are discovered through association rule mining. A Markov prediction model is constructed to predict the node resource demand for the next hour, and the average prediction error is controlled within 10%. In the resource scheduling optimization model, three objective functions, including fault recovery time, switching overhead, and load balancing, are set, and the NSGA-II multi-objective genetic algorithm is used to solve the problem. The Pareto optimal solution set is found among 200 alternatives, achieving a trade-off between the shortest fault recovery time, the smallest switching overhead, and the most balanced load. In terms of routing optimization, differentiated routing cost models are constructed for different service types. For example, for VoIP services with high real-time requirements, higher weights are given to latency and jitter, and for financial services with high reliability requirements, higher weights are given to packet loss rate. Using the deep reinforcement learning algorithm DQN, with routing cost as the environment state and routing change as the action, the optimal routing strategy is found through continuous trial and error and learning. After 5,000 rounds of training, the network latency is reduced by 20% and the packet loss rate is reduced.

[0030] Analyze the availability of backup resources, and screen available backup node and link resources; according to the results of fault impact assessment, identify the boundary of the fault impact, search for alternative paths outside the fault impact area, and use the redundancy of the network structure to find the best alternative data transmission paths.

[0031] Combined with the hardware configuration, software version, and load status of the backup nodes, build an analysis model for the availability of backup resources, calculate the availability score of each backup node, and screen out the backup nodes with high availability as candidate switching objects according to the preset threshold. At the same time, evaluate the performance indicators of the backup link, such as bandwidth, delay, and packet loss rate, and determine the set of available links that meet the data transmission requirements; take the faulty node as the source point and the backup node as the sink point to obtain the maximum transmission capacity from the faulty node to each backup node, and evaluate the switching priority of the backup node based on this; identify the boundary area affected by the fault, and extract the topological relationship of the boundary nodes and links; start the search for alternative paths outside the external area with the fault impact boundary as the starting point, combine multiple constraint conditions such as the bandwidth, delay, and reliability of the link, find multiple alternative transmission paths from the boundary node to the target node, and sort them according to the path performance indicators, and preferentially select the path with the best indicators as the main path for data transmission, and other paths as backups.

[0032] Exemplarily, build an analysis model for the availability of backup resources, comprehensively consider factors such as the hardware configuration, software version, and load status of the backup nodes, and use the fuzzy comprehensive evaluation method to calculate the availability score of each backup node. The fuzzy comprehensive evaluation method can be implemented through the following steps: first, determine the evaluation index system, select the key indicators reflecting the availability of the backup node, and build a hierarchical index system. Then, use the analytic hierarchy process (AHP) to determine the index weights, compare the indicators pairwise, calculate the relative importance of each indicator, and ensure the rationality of the weight distribution through consistency testing. Next, conduct membership degree evaluation, and according to the actual values of the indicators, use triangular membership functions, trapezoidal membership functions, etc. to determine its membership to "

[0033] Then, a fuzzy comprehensive evaluation is performed, and the membership and weight of each indicator are weighted and summed to obtain the comprehensive membership of the backup node at each fuzzy level. The level with the largest membership is selected as the availability evaluation result of the node. Nodes are screened according to the preset threshold, and nodes with a comprehensive membership higher than the threshold are included in the set of available nodes as candidate switching objects. At the same time, performance indicators such as bandwidth, delay, and packet loss rate of the backup link are evaluated to determine the set of available links that meet the data transmission requirements. After a fault occurs, the boundary area affected by the fault is identified through network topology analysis and route tracking, the topological relationship between boundary nodes and links is extracted, and the fault impact subgraph is constructed. For complex For network topology, a community discovery algorithm, such as the Louvain algorithm, can be used to divide the network into multiple closely connected communities, and impact analysis can be performed on a community basis to reduce computational complexity. When using the community discovery algorithm to divide the fault-affected area, the network topology graph is first converted into an undirected weighted graph, with node connection relationships as edges and performance indicators such as link bandwidth and latency as edge weights. Then, the Louvain algorithm is used to divide the network into communities, with the goal of maximizing modularity. By continuously iteratively optimizing the communities to which the nodes belong, a closely connected node cluster is obtained. Next, based on the community where the faulty node is located, the initial area affected by the fault is determined, and considering that the fault may spread across multiple communities, adjacent communities are expanded and merged. Get the complete impact range. Finally, extract the boundary nodes and key links of the fault-affected area as the starting point and constraints for subsequent path search and optimization. Starting from the fault-affected boundary, start the backup path search in its external area, and use classic algorithms such as Dijkstra's shortest path algorithm and Yen'sk-shortestpathsalalgorithm. Consider multiple constraints such as link bandwidth, delay, and reliability to find multiple alternative transmission paths from the boundary node to the target node, and sort them according to the path performance indicators. Prioritize the path with the best indicator as the main path for data transmission, and other paths as backup. In the path search process, make full use of the structural redundancy of the network and tap into potential transmission capacity. On the one hand, loop detection algorithms, such as the Tarjan algorithm, are used to find all loops in the network and use them as a candidate set of alternative paths; on the other hand, disjoint path search algorithms, such as the Suurballe algorithm, are used to solve edge-disjoint paths or node-disjoint paths between nodes, so as to maximize the use of network resources and improve the parallelism of data transmission. For large-scale complex networks, a hierarchical path search strategy can be adopted to divide the network into multiple logical layers, such as the core layer, aggregation layer, and access layer, and perform path calculation and optimization in each layer. At the same time, virtual links or tunnels are established between different layers to achieve cross-layer path connectivity and resource scheduling, thereby improving the efficiency and flexibility of path search.A virtual link refers to an end-to-end connection logically established on top of a physical network according to service requirements and routing policies. By combining multiple physical links into one virtual link, the complexity of the underlying network can be masked, simplifying the path calculation and traffic scheduling of upper-layer services. A tunnel is a virtual link implementation method based on encapsulation technology. By encapsulating an additional header outside the original data packet, a logical channel is established between the tunnel entrance and exit to achieve transparent transmission of data between different network layers or protocol stacks. Common tunnel technologies include GRE, VXLAN, IPSec, etc. When searching for paths, virtual links can be constructed separately for different levels of network views to form a multi-level logical topology. Independent routing protocols and optimization strategies are used within each level to achieve hierarchical autonomy. Between levels, vertical connections are established through tunnel technology to map upper-layer virtual links to lower-layer physical links, realizing end-to-end data transmission. To cope with the dynamically changing network environment, an adaptive path optimization mechanism is introduced. By regularly collecting real-time performance data such as link traffic and latency, and using time series prediction algorithms such as the AutoRegressive Integrated Moving Average model (ARIMA) or Long Short-Term Memory (LSTM) neural network, the trend of link status is predicted, and the path selection strategy is dynamically adjusted. When link congestion or failure risk is detected, path recalculation and switching are automatically triggered to ensure the continuity and reliability of data transmission. At the same time, reinforcement learning algorithms such as Q-Learning are used to dynamically optimize the path scoring function according to network status and transmission feedback, achieving adaptive optimization of path selection. In the analysis of standby resource availability, 10 key indicators are selected, such as CPU utilization rate, memory occupancy rate, hard disk IO, etc. The weights of each indicator are calculated by the AHP method to construct a fuzzy comprehensive evaluation matrix. Taking CPU utilization rate as an example, three fuzzy levels of 0-30% for low load, 30%-60% for medium load, and 60%-90% for high load are set, and membership degrees of 0.2, 0.5, and 0.8 are assigned respectively to calculate the comprehensive availability score of the backup node. When the comprehensive score exceeds 0.7, the node is included in the candidate resource pool. In the analysis of fault impact, the network topology relationship and link attributes are extracted to construct a fault impact subgraph containing 1000 nodes and 2000 edges. The Louvain algorithm is used for community division, and the influence factor of the fault node is introduced into the modularity function. Through 10 rounds of iteration, it converges to a stable community structure to identify the key areas and boundary nodes affected by the fault.In the backup path search, for 100 boundary nodes and 500 alternative links, the Yen's algorithm is used to calculate the TOP-10 shortest paths. With bandwidth, delay, and packet loss rate as constraints, the Dijkstra algorithm is used to solve the constrained shortest path problem, and a set of backup paths that meet the conditions is selected. At the same time, the Tarjan algorithm is used to search for loops in the alternative links, and 20 independent loops are found as candidate paths. In the construction of virtual links, with the service initiator and receiver as endpoints, 5 core nodes are abstracted as intermediate nodes of the virtual links. The minimum spanning tree algorithm such as Kruskal is used to calculate the topology of the virtual links, and the GRE tunneling technology is used to map the virtual links to the underlying physical links to achieve network layering and end-to-end connectivity. In the process of path optimization, link status data is collected every 5 minutes, and the LSTM network is used for traffic prediction. The link weights are dynamically adjusted according to the prediction results. When the predicted traffic exceeds 80% of the link bandwidth, the backup path is automatically switched. At the same time, the Q-Learning algorithm is used for online learning of path selection strategies. Through the balance of exploration and exploitation, the path scoring function is continuously optimized. After 1000 rounds of iteration, the adaptive convergence of path selection is achieved, the average delay is reduced by 20%, and the packet loss rate is reduced.

[0034] Combined with the reconstruction speed of the connection point, find new links or relay nodes for the disconnected connection, while optimizing the flexibility of route modification and dynamically updating the network routing table to adapt to the new network structure.

[0035] Combined with the processing capacity of the connection point, the bandwidth of the backup link, and the type of routing protocol, estimate the reconstruction time of the connection point, obtain the probability distribution of the reconstruction time, and filter out the alternative connection points and links whose reconstruction speed meets the requirements according to the set delay threshold; for the disconnected connection, by analyzing the location of the link break point and the connection requirements, select the link with the minimum cost and the optimal path in the set of alternative links as the new connection channel; if the alternative link resources are insufficient or do not meet the requirements, introduce relay nodes, and through the forwarding ability of the relay nodes, form a composite link that spans multiple hops; search for the optimal combination of relay nodes in the network topology, evaluate the performance of the composite link, weigh the transmission delay and reliability, and select the optimal relay scheme.

[0036] Exemplarily, considering factors such as the processing capacity of the connection point, the bandwidth of the backup link, and the type of routing protocol, based on machine learning algorithms such as support vector machine (SVM) or neural network (NN), by constructing a network reconnection prediction model, the reconstruction time of the connection point is estimated. The construction of the network reconnection prediction model can adopt the following steps: First, collect historical network fault and connection point reconstruction data, including the processing capacity of the connection point, the bandwidth of the backup link, the type of routing protocol, the reconstruction time, etc., and clean and preprocess the data to remove outliers and missing values. Then, select a suitable machine learning algorithm, such as SVM or NN, and determine the hyperparameters and architecture of the model according to the characteristics of the data and the performance of the model. Next, divide the data set into a training set, a validation set, and a test set, use the training set to train the model, and continuously adjust the parameters of the model through the backpropagation algorithm and the gradient descent method to minimize the prediction error. During the training process, use the validation set to evaluate the generalization ability of the model, and select the optimal model parameters and structure through methods such as cross-validation; finally, evaluate the prediction performance of the model on the test set, calculate metrics such as mean square error (MSE) and mean absolute error (MAE), and judge the prediction accuracy of the model. Through continuous iteration and optimization, a network reconnection prediction model with stable performance and strong generalization ability is obtained, and according to the set delay threshold, alternative connection points and links whose reconstruction speed meets the requirements are selected. For the disconnected connection,

[0037] Perform link repair. By analyzing the location of the link break point and connection requirements, using graph theory algorithms such as the minimum spanning tree or shortest path algorithm, select the link with the lowest cost and optimal path in the set of alternative links as the new connection channel. At the same time, consider factors such as the capacity and reliability of the link to ensure that the new link meets the quality of service requirements for service transmission. The network reconstruction scheme can include deploying an SDN controller, using the OpenFlow protocol to issue forwarding rules, and performing real-time optimization on the newly established link. According to metrics such as link bandwidth utilization and delay jitter, dynamically adjust the queue scheduling strategy and traffic shaping parameters to ensure efficient utilization of the link and load balancing, and minimize the impact of link switching on services. If the alternative link resources are insufficient or do not meet the requirements, consider introducing relay nodes. Through the forwarding ability of the relay nodes, establish a multi-hop composite link. Use heuristic algorithms such as the Ant Colony Optimization (ACO) algorithm to search for the optimal combination of relay nodes in the network topology. When selecting relay nodes and constructing the composite link, the specific steps include abstracting the network topology graph into a weighted directed graph, where nodes represent network devices, edges represent links, and weights represent the cost or performance metrics of the links. Then, initialize the parameters of the ant colony algorithm, such as the number of ants, pheromone concentration, heuristic function, etc., and set the termination condition of the algorithm. In each iteration, starting from the source node, release a certain number of ants. Each ant selects the next-hop node according to the pheromone concentration and heuristic function with a certain probability until it reaches the target node or reaches the maximum hop count limit. During the ant traversal process, record the path and performance metrics of each ant, such as delay and bandwidth, and calculate the fitness value of each path. After the iteration ends, update the pheromone concentration on each edge according to the fitness value. The higher the pheromone concentration, the better the quality of the path where the edge is located, and the greater the probability of being selected. Repeat the above steps until the termination condition of the ant colony algorithm is reached, such as the iteration count reaching the upper limit or the fitness value converging. Finally, select the path with the highest fitness value as the optimal solution from all the paths obtained in the iterations. Use the intermediate nodes on this path as relay nodes to construct the composite link, evaluate the end-to-end performance of the composite link, weigh the transmission delay and reliability, and select the optimal relay scheme to ensure the reachability and quality of service of the end-to-end connection. After completing the link reconstruction, adaptively adjust the routing strategy according to the new network topology structure to optimize the convergence speed and calculation efficiency of the routing algorithm. During the routing convergence process, use reinforcement learning algorithms such as Q-learning or SARSA and other specific methods to adaptively adjust the parameters and weights of the routing strategy according to the network state and service feedback, and continuously optimize the accuracy and stability of routing selection.For large-scale networks, a hierarchical routing mechanism is introduced to divide the network into multiple regions. Inside the regions, fast-converging link-state routing protocols such as OSPF (Open Shortest Path First) or IS-IS (Intermediate System to Intermediate System) are used. Between the regions, scalable distance-vector routing protocols such as BGP (Border Gateway Protocol) are adopted to achieve reachability transfer and policy control between routing domains. For the complex and changeable network environment, for adaptive routing control, according to the dynamic changes of network traffic, the time intervals for routing calculation and update are adjusted in real time. In the peak traffic period, a fast-convergence mode is adopted to shorten the routing update cycle. In the stable traffic period, a smooth-convergence mode is adopted to reduce routing oscillation and jitter. At the same time, traffic engineering methods are used to optimize the objective function of routing calculation, taking into account multiple performance indicators such as load balancing, link utilization rate, and delay, to improve the utilization efficiency of network resources. A semantic association model among network devices, links, and routing policies is constructed, and an inference engine is used to realize the automatic configuration and optimization of routing policies, improving the intelligent level and flexibility of routing management; the specific implementation steps include data on network devices, links, routing policies, etc., including device types, link bandwidths, routing protocols, policy rules, etc., and the data is subjected to structured processing and semantic annotation. Then an ontology model of the network routing knowledge graph is constructed, defining basic concepts such as network entities, attributes, and relationships, forming an ontology architecture that conforms to the characteristics of the network routing field. Next, natural language processing technologies such as named entity recognition and relation extraction are used to automatically extract entities and relationships from network configuration files and log data and map them to the ontology model, continuously enriching and improving the knowledge graph. After the knowledge graph is constructed, graph database technologies such as Neo4j are used to store the knowledge graph as a graphical data structure to support efficient graph query and reasoning operations. For specific routing management tasks such as routing policy optimization and fault diagnosis, inference rules and query algorithms based on the knowledge graph are designed, and technologies such as graph pattern matching and shortest path search are used to obtain the required information and knowledge from the knowledge graph. During the reasoning process, an ontology reasoning engine such as Apache Jena is used to automatically generate new knowledge and conclusions according to predefined reasoning rules to assist in the decision-making and optimization of routing management. In the construction of the network reconnection prediction model, 10 key features such as the CPU utilization rate of connection points, memory occupancy, spare link bandwidth, and routing protocol type are extracted from 1000 historical network fault records. The support vector regression (SVR) algorithm is adopted, and the model parameters are determined through 5-fold cross-validation to train a reconnection time estimation model with an average prediction error within 10%.When a connection interruption occurs, the minimum spanning tree algorithm is used to select the link with the lowest latency and the highest bandwidth among the five alternative links as the new connection channel. The forwarding rules are issued through the OpenFlow controller to optimize the queue scheduling strategy of the new link, ensuring that the latency jitter is less than 5 ms. When there are insufficient alternative links, the ant colony algorithm is used to search for relay nodes. With latency and node reliability as the optimization objectives, the optimal relay path is found through 500 iterations, and the end-to-end latency is controlled within 50 ms. During the routing convergence process, the Q-learning reinforcement learning algorithm is adopted, with network throughput and latency as the states and the routing update strategy as the action. The optimal routing update strategy is obtained through 1000 rounds of training, and the average convergence time is shortened from 30 seconds to 10 seconds. For a network with a scale of 200 nodes, the OSPF and BGP protocols are used to implement hierarchical routing. By introducing a traffic-aware adaptive routing update mechanism, the update time is dynamically adjusted according to the link utilization rate, shortened to 1 minute during peak periods and relaxed to 5 minutes during idle periods, effectively suppressing routing jitter. 500 device entities, 1000 link entities, and 2000 routing policy entities are automatically extracted from 10000 configuration records to construct a network routing knowledge graph containing 5000 triples. Through a graph reasoning engine based on SPARQL, the automatic optimization of routing policies is realized, and the matching degree between the generated policy rules and expert configurations reaches over 95%.

[0038] Based on the monitoring of the current network state, the real-time operation data of the network is retrieved, including the load conditions of each node, the bandwidth utilization rate, and the stability indicators of the links, to evaluate the executability of the reconstruction plan.

[0039] Monitor the current network state and retrieve the real-time operation data of the network; combine the indicators of packet loss rate, delay jitter, and bit error rate of the links to evaluate the link stability, calculate the health score of each link, and set a health threshold. When the link score is lower than the health threshold, trigger link reconstruction or backup link switching; construct a simulation model similar to the actual network topology and configuration parameters. By inputting the real-time monitoring data, simulate the network behavior and performance after the execution of the reconstruction plan, evaluate the impact of the reconstruction on the network service quality, and select the optimal reconstruction plan based on the evaluation results.

[0040] Exemplarily, monitoring agents that can be deployed on each node of the network periodically collect key performance indicator data such as CPU utilization, memory usage, and network card traffic of the nodes through the SNMP protocol, and can also be transmitted in real time to the central monitoring platform through message queue middleware such as RabbitMQ. Using a stream computing engine such as Apache Flink or Spark Streaming, the data is cleaned, transformed, and aggregated to calculate key indicators such as the real-time load conditions and bandwidth utilization of each node, and by setting reasonable thresholds, real-time alarms for network anomalies and bottlenecks are realized, providing a basis for network reconstruction. Based on real-time data analysis, intelligent analysis and anomaly detection are performed on network monitoring data. By establishing a baseline model of network performance and using algorithms such as support vector machine (SVM) and isolation forest, it is determined in real time whether the network state deviates from the normal range. In the training stage of the anomaly detection model, the initial anomaly threshold is determined through the cross-validation method as a hyperparameter of the model; in the online prediction stage of the model, for each newly collected data point, its anomaly score is calculated and compared with the current threshold to determine whether it is abnormal. At the same time, the anomaly scores of each data point are accumulated and statistically analyzed to calculate the anomaly score distribution within a certain period of time, including mean, variance, quantiles, etc. According to the change trend of the anomaly score distribution, the anomaly threshold is adaptively adjusted. For example, when the mean and variance of the anomaly score increase significantly, the threshold is increased; when the mean and variance of the anomaly score decrease significantly, the threshold is decreased. During the threshold adjustment process, a smoothing factor and a decay factor are introduced, and methods such as exponentially weighted moving average (EWMA) are used to smoothly update the threshold to avoid overly drastic changes in the threshold, resulting in unstable detection results. In addition, the upper and lower limits and step size of the threshold adjustment can be set to prevent the threshold from being too high or too low, affecting the sensitivity and accuracy of detection. Finally, the dynamically learned threshold is fed back to the anomaly detection model to realize the online update and optimization of the model, continuously improving the adaptability and robustness of anomaly detection. For the evaluation of link stability, a link health score model is introduced, comprehensively considering multiple key indicators such as packet loss rate, delay jitter, and bit error rate of the link, and methods such as weighted average or analytic hierarchy process (AHP) are used to calculate the health score of each link. Specifically, the AHP method can be used. By means of expert scoring, the weights of each indicator are determined, an indicator hierarchy structure is constructed, and then the indicators of each link are normalized, and the weighted average value is calculated to obtain the comprehensive health score of the link. A reasonable health threshold is set. When the link score is lower than the threshold, link reconstruction or backup link switching is triggered to ensure the reliability and stability of the network.When evaluating the executability of the reconstruction plan, network simulation technology is adopted, and professional network simulation tools such as NS-3 and OPNET are used to build a simulation model similar to the actual network topology and configuration parameters. By injecting real-time monitoring data, the network behavior and performance after the execution of the reconstruction plan are simulated, and the impact of the reconstruction on the network service quality is evaluated. For the key parameters in the simulation model, such as link bandwidth, delay, packet loss rate, etc., a set of parameter combinations are generated by designing orthogonal experiments or Latin hypercube sampling methods to form an experimental sample space. For each parameter combination, the simulation model is run to evaluate the performance indicators of the reconstruction plan, such as throughput, delay, packet loss rate, etc., and a set of simulation results of the performance indicators are obtained. By methods such as principal component analysis (PCA) or partial least squares regression (PLSR), a sensitivity model between parameters and performance indicators is established to quantify the influence degree and direction of each parameter on the performance indicators. According to the results of the sensitivity analysis, the key parameters that have the greatest impact on the performance of the reconstruction plan are identified, and optimization methods such as gradient descent and evolutionary algorithms are used to search for the optimal parameter configuration. In the process of parameter optimization, the idea of multi-objective optimization is introduced, and multiple performance indicators are comprehensively considered, such as the trade-off between throughput and delay. By constructing a weighted objective function or Pareto front method, the globally optimal parameter configuration is sought. For the optimized parameter configuration, sensitivity analysis is carried out to evaluate its robustness and adaptability to ensure that good performance can still be maintained within a certain range of parameter perturbations. Finally, the optimized parameter configuration is applied to the actual network reconstruction plan to guide the execution and optimization of the reconstruction, and improve the efficiency and reliability of the reconstruction. Before the reconstruction is executed, configuration management tools such as RANCID are used to back up the current configuration of the network devices and save it as a complete configuration snapshot as the basis for rollback. During the execution of the reconstruction, by setting checkpoints and transaction mechanisms, the reconstruction task is divided into multiple atomic operations. After each operation is completed, a configuration submission and confirmation are carried out to ensure the atomicity and consistency of the reconstruction. If an exception or error occurs during the execution of the reconstruction, the rollback mechanism is immediately triggered to restore the configuration of the network devices to the previous stable version to ensure the availability and stability of the network. In the deployment of the network monitoring agent, 10 core nodes and 50 edge nodes are selected, and the SNMPV3 protocol is adopted to collect 12 key indicators such as CPU utilization, memory usage, and network card traffic of the nodes at a 5-second interval and transmit them to the central monitoring platform in real time with a 10ms delay through the ZeroMQ message queue. The Flink stream computing engine is used to process the monitoring data in real time. Through operations such as data cleaning and normalization, the average node load and bandwidth utilization are calculated in a 30-second window, and the monitoring data with a 60-second granularity is stored in the Prometheus time series database, retaining 30 days of historical data.In the anomaly detection model, the SVM algorithm is selected to build the performance baseline. The optimal parameters C = 10 and gamma = 0.01 are determined through 5-fold cross-validation. The anomaly score is calculated for each newly collected data point and compared with the dynamic threshold. The initial value of the threshold is set to 2.5, and it is adaptively adjusted every 10 minutes based on the anomaly score distribution in the recent 1 hour. The adjustment step size is 0.1, the upper limit is 3.0, and the lower limit is 2.0. The EWMA smoothing coefficient alpha = 0.6 is used. In the link health score, the weights of packet loss rate, delay jitter, and bit error rate are determined to be 0.5, 0.3, and 0.2 respectively through the AHP method. The Min-Max normalization method is used to dimensionless process each index. The link health score range is [0,1], and the threshold is set to 0.6. In the reconstruction scheme evaluation, the NS-3 simulation tool is used to build a simulation model with 100 nodes and 500 links. 200 groups of parameter combinations are generated through Latin Hypercube sampling. Each simulation runs 10 times, and the average value is taken as the performance index result. The PLSR algorithm is used to establish a sensitivity model between parameters and performance indicators, and it is found that bandwidth, delay, and packet loss rate are the key parameters affecting throughput and delay. The NSGA-II genetic algorithm is used for multi-objective optimization to balance the Pareto front of throughput and delay and evaluate the robustness of parameter configuration. The finally selected parameter configuration has a performance degradation of no more than 5% within a 10% bandwidth change range. RANCID combined with Ansible is used to realize the automatic backup and rollback of network device configurations. A configuration snapshot is taken before the reconstruction execution, and configuration confirmation is performed after each atomic operation. In case of an anomaly, the rollback can be completed within 30 seconds.

[0041] Step S105: Update network parameters in real time, including adjusting channel allocation and power control to suppress new network conflicts and performance degradation caused by reconstruction; test the performance of the reconstructed network through a network simulator to verify whether the adjusted configuration meets the design requirements, and execute the reconstruction scheme on the premise of meeting real-time and reliability.

[0042] Based on the status information of network devices and according to the requirements of network reconstruction, adjust the configuration parameters of the devices, including the routing table and access control list, to minimize the impact of reconstruction on services; through spectrum sensing and channel state prediction, optimize channel allocation and power control, dynamically adapt to changes in network topology and propagation environment, and suppress co-channel interference and performance degradation caused by reconstruction; conduct impact analysis on the reconstruction plan, and formulate corresponding risk response strategies and rollback plans; use a network simulation platform to evaluate and verify the performance of the reconstructed network, including key indicators such as throughput, latency, packet loss rate, and fairness; by comparing the performance differences before and after reconstruction, optimize the reconstruction plan, and through simulation tests, verify whether the reconstructed network meets the design requirements and service needs; if the reconstructed network meets the design requirements and service needs, then implement the reconstruction plan; when implementing network reconstruction, divide the reconstruction process into multiple stages, and only adjust some network devices or links in each stage to reduce the complexity and risk of reconstruction.

[0043] Exemplarily, based on the collected network device status information, key performance indicators such as link bandwidth, latency, packet loss rate, etc. are extracted, and according to the requirements of network reconstruction, the configuration parameters of the devices are dynamically adjusted, such as routing tables, access control lists, etc., to ensure a smooth transition and continuous operation of the network during the reconstruction process and minimize the impact of the reconstruction on services. Through machine learning algorithms, such as supervised learning algorithms like support vector machine (SVM), decision tree, neural network, etc., the network status is classified and predicted to achieve real-time detection and early warning of abnormal situations such as network congestion and faults. At the same time, unsupervised learning algorithms, such as K-means clustering, principal component analysis (PCA), etc., are used to analyze and mine network traffic and user behavior to discover internal laws and correlations such as traffic patterns and user groups, providing data support for network optimization and resource allocation. In addition, reinforcement learning algorithms, such as Q-learning, policy gradient, etc., can be used to adaptively optimize network control and resource management. By continuously trial and error and exploration, the optimal routing strategy, traffic scheduling scheme, etc. are learned to dynamically adapt to the changes in network status and improve the intelligence level and self-regulation ability of the network. For the wireless network environment, cognitive radio technology is introduced. Through spectrum sensing and channel state prediction, the channel allocation and power control strategies are optimized in real time to dynamically adapt to the changes in network topology and propagation environment, suppress co-channel interference and performance degradation caused by reconstruction, and improve the reliability and stability of the wireless link; the specific implementation includes first, through spectrum sensing technology, the occupancy situation and signal quality of the wireless channel are monitored in real time, and methods such as energy detection, matched filtering, and cyclic stationary feature detection are used to quickly and accurately identify idle channels and interference sources; then, according to the results of spectrum sensing, combined with the radio environment map and historical data, the channel state is predicted. According to the real-time spectrum monitoring and prediction results, the working frequency and bandwidth of the wireless device are dynamically adjusted, and idle channels with less interference are preferentially selected to improve spectrum utilization efficiency and transmission reliability. Distributed power control algorithms, such as game theory, particle swarm optimization, etc., are used to adaptively adjust the transmission power according to factors such as the location of the wireless device and interference situation, reduce co-channel interference, and improve network capacity and coverage quality. Before implementing network reconstruction, a risk assessment and early warning mechanism for network reconstruction is established, and a comprehensive impact analysis of the reconstruction plan is carried out, including the impact on aspects such as network topology, routing protocol, and service model. Through methods such as fault tree analysis, the reconstruction risk is quantified, and corresponding risk response strategies and rollback plans are formulated to ensure the safety and controllability of the reconstruction process. Fault tree analysis is a top-down logical reasoning method used to analyze the reliability and safety of complex systems. The specific steps include determining the top event of the fault tree according to the goals and scope of network reconstruction, that is, the final state of network reconstruction failure or causing serious service interruption. Then, using expert experience and historical data, various intermediate events and basic events that cause the top event to occur are identified, such as device failures, configuration errors, link interruptions, etc., and the branch structure of the fault tree is constructed according to the causal logical relationship.During the construction of the fault tree, Boolean logic gates such as AND gates and OR gates are used to represent the combination relationships between different events, and the failure probabilities of each logic gate are calculated based on the conditional probabilities of event occurrences. Through quantitative analysis techniques of the fault tree, such as Binary Decision Diagrams (BDDs) and importance analysis, the occurrence probability of the top event is calculated, the overall risk level of network reconstruction is evaluated, and key risk factors and weak links are identified. Based on the results of the fault tree analysis, targeted risk mitigation and emergency response plans are formulated, such as adding redundant backups, optimizing configuration processes, strengthening monitoring and early warning, etc., to improve the reliability and security of network reconstruction. On the basis of the fault tree analysis, other risk analysis methods such as attack trees and event trees can also be introduced to evaluate the security risks of network reconstruction from different perspectives and dimensions, improving the comprehensiveness and accuracy of risk assessment. Using network simulation platforms such as OPNET and NS-3, a comprehensive performance evaluation and verification of the reconstructed network are carried out, including key indicators such as throughput, latency, and packet loss rate. By comparing the performance differences before and after reconstruction, the reconstruction plan is optimized and improved, and through simulation tests, it is verified whether the reconstructed network meets the design requirements and business needs. When implementing network reconstruction, a progressive migration strategy is adopted, dividing the reconstruction process into multiple stages. Only part of the network devices or links are adjusted in each stage. Through configuration updates and switches, the complexity and risk of reconstruction are reduced, and by setting rollback points and health check mechanisms, the controllability and reversibility of the reconstruction process are ensured. In SDN network reconstruction, by deploying an OpenFlow controller and a flow table collection agent, 20 key indicators such as traffic and latency of 100 network nodes are collected at a 1-second interval, and real-time alarms for network anomalies are realized using anomaly detection algorithms such as Double Exponential Smoothing (DES), with an alarm accuracy rate reaching 95%. In anomaly localization, a decision tree algorithm is used to classify the network state. By training 2000 labeled data, the accuracy rate of fault root cause localization is 90%. Using reinforcement learning algorithms, for intention-driven network optimization scenarios, by designing 20 network control primitives, a reward and punishment function for network state and optimization actions is established. After 500 rounds of training, the network utilization rate is increased from 60% to 85%. A network device, configuration, and policy semantic library containing 1000 nodes and 5000 edges is constructed, and the Graph Attention Network (GAT) algorithm based on the attention mechanism is introduced. By updating the knowledge graph through online learning, the accuracy rate of network configuration recommendation reaches 92%. In cognitive radio network reconstruction, a spectrum sensing algorithm based on energy detection is adopted. By setting an energy threshold of -95 dBm, the frequency band from 900 MHz to 2.4 GHz is scanned at a 10-ms interval, and the missed detection rate of idle channel detection is less than 1%. Combining the Kalman filter and the Support Vector Regression (SVR) algorithm, the channel occupancy rate is predicted, and the prediction error is controlled within 5%.In dynamic channel allocation, the Markov decision process (MDP) is used for modeling. By defining 6 channel states and 10 allocation actions, the optimal policy is solved, which increases the channel utilization rate by 15%. In power control, a distributed algorithm based on game theory is adopted. By defining the utility function and considering the reciprocity degree among nodes, it converges to the Nash equilibrium after 20 rounds of games, and the system throughput is increased by 25% compared with the fixed power allocation. In the risk assessment of network reconstruction, the fault tree analysis (FTA) method is used. For the failure scenarios of network architecture adjustment, 10 key risk events are identified through expert evaluation, a fault tree containing 50 logic gates is constructed, and the Monte Carlo simulation is used to calculate the fault probability. Plans are made for high-risk events, and the reconstruction failure rate is controlled below 1%. In the simulation verification, OPNET is used to model the reconstruction scheme, a simulation scenario with 500 nodes is built, and through parameter scanning experiments, 10 key performance indicators such as delay, jitter, and packet loss rate are evaluated. Combining with the confidence interval analysis, the link bandwidth and cache configuration are optimized to make the network performance after reconstruction meet the requirements of the service level agreement (SLA).

[0044] Step S106, after executing the reconstruction scheme, adjust the data collection and analysis period according to the fault frequency, and control the reconstruction interval to reduce redundant network structure adjustment and potential network fluctuations, including actively detecting and passively listening to the distribution network, and conducting two-way monitoring of the network status.

[0045] Combined with the topological structure, asset status, and environmental factors of the distribution network, conduct fault prediction and risk assessment, including simulating the network behavior and influence range under various fault scenarios, evaluating the fault risk level and loss degree; by analyzing the frequency, interval, and trend of fault occurrence, adjust the time window and granularity of data collection and analysis; according to the results of fault prediction and risk assessment, dynamically optimize the reconstruction scheme and execution cycle of the distribution network. In high-fault areas and important load nodes, correspondingly shorten the reconstruction interval to quickly respond to potential fault hazards and abnormal situations; in low-fault areas and secondary load nodes, correspondingly extend the reconstruction interval to reduce the reconstruction cost and the impact on network stability; actively detect the monitoring blind spots of the distribution network and the uncertainty of fault prediction, and conduct non-contact detection of assets to actively discover hidden defects; at the same time, through the active injection and backhaul analysis of fault signals, locate the position and type of the fault.

[0046] Exemplarily, intelligent sensing units can be embedded in primary and secondary power distribution equipment to online monitor and analyze multi-dimensional signals such as load fluctuations, harmonic content, and partial discharge signals. By extracting and fusing features from multi-source heterogeneous monitoring data of the distribution network, including load curves, power quality, environmental factors, etc., a unified feature vector is formed. Feature extraction can adopt signal processing methods such as wavelet transform and Fourier transform, and feature fusion can adopt dimensionality reduction techniques such as principal component analysis (PCA) and independent component analysis (ICA). Machine learning algorithms, such as support vector machines (SVM), random forests, etc., are used to classify and detect anomalies in the feature vector, and potential fault symptoms such as abnormal load fluctuations and excessive harmonics are identified. Algorithm training can adopt supervised learning, semi-supervised learning, etc., and abnormal samples are labeled through historical data and expert knowledge. Based on anomaly detection, spatio-temporal data mining techniques such as spatio-temporal clustering and spatio-temporal association rules are used to analyze the correlation and propagation law of abnormal events in time and space, and the chain reaction and influence scope of faults are identified. Spatio-temporal clustering can adopt methods such as density-based DBSCAN algorithm and hierarchical clustering, and spatio-temporal association rules can adopt algorithms such as Apriori and FP-growth. According to the results of spatio-temporal association analysis, causal reasoning of abnormal events is carried out to mine the root causes and triggering factors of faults. Causal reasoning can adopt probabilistic graphical models such as Bayesian networks and Markov logic networks, and through probability calculation and conditional dependence analysis of causal relationships, early warning of fault symptoms and multi-factor traceability are realized. In terms of active detection, devices such as intelligent inspection robots and drones can be introduced to regularly conduct non-contact detections such as infrared thermal imaging and ultraviolet photoelectric detection on key assets such as lines and poles, and actively discover hidden defects such as conductor sag and insulation aging. At the same time, through active injection and backhaul analysis of fault signals, such as impedance method and traveling wave time difference method, the location and type of faults are accurately located, aiming at the monitoring blind spots of the distribution network and the uncertainty of fault prediction, and the accuracy and reliability of fault location are improved. According to the historical fault data recorded by the distribution automation system, time series analysis methods such as ARIMA model or Prophet algorithm are used to model and predict the time series data of network faults. By analyzing the frequency, interval, and trend of fault occurrences, the time window and granularity of data collection and analysis are adaptively adjusted, reducing unnecessary calculation and storage overhead while ensuring analysis accuracy, and improving the real-time performance and efficiency of fault prediction. Combining with the physical model and simulation of the distribution network, fault prediction and risk assessment of the distribution network are carried out, including constructing the topological model of the distribution network, including main equipment such as lines, transformers, switches, etc., and their connection relationships. Data structures such as node admittance matrix or adjacency list can be used to represent the network topology. Then, a power flow model of the distribution network is established, considering factors such as line parameters, load types, and distributed power generation output, to form a power flow equation set. Commonly used power flow algorithms include Newton-Raphson method, fast decoupling method, etc.Based on the power flow model, state estimation of the distribution network is carried out. Using real-time measurement data and the power flow model, the voltage amplitude and phase angle of each node in the network, as well as the power flow distribution of the lines, are estimated. Algorithms such as the least squares method and the weighted least squares method can be used for state estimation. According to the results of state estimation, the static security of the distribution network is evaluated, including indicators such as the voltage qualification rate and the line overload rate, to judge whether the network meets the operation constraint conditions. Methods such as the continuation power flow method and the optimal power flow method can be used for static security analysis. On the basis of static analysis, further dynamic simulation of the distribution network is carried out to simulate the network dynamic response process under emergencies such as faults and load fluctuations, and to evaluate the fault impact range and duration. Technologies such as time-domain simulation and harmonic analysis can be used for dynamic simulation. Combining the physical model of the distribution network with machine learning algorithms, a fault prediction and risk assessment model is trained through a large amount of simulation data to improve the generalization ability and robustness of the model. By comprehensively modeling the topological structure, asset status, environmental factors, etc. of the distribution network, the network behavior and influence range under various fault scenarios are simulated, the fault risk level and loss degree are evaluated, and differentiated risk prevention and control strategies and emergency plans are formulated to improve the operation safety and reliability of the distribution network. According to the results of fault prediction and risk assessment, the reconfiguration scheme and execution cycle of the distribution network are dynamically optimized. In high-fault areas and important load nodes, the reconfiguration interval is appropriately shortened to increase the frequency and accuracy of network architecture adjustment, and quickly respond to potential fault hazards and abnormal conditions; in low-fault areas and secondary load nodes, the reconfiguration interval is appropriately extended to reduce unnecessary network structure adjustments and reduce the reconfiguration cost and the impact on network stability. By adaptively adjusting the reconfiguration cycle, while ensuring power supply reliability, the network fluctuations and energy losses caused by frequent reconfiguration are minimized, and the dynamic optimization and intelligent scheduling of distribution network reconfiguration are realized. In the all-round perception of the distribution network, by deploying edge computing gateways in 500 distribution terminals and 1000 user intelligent meters, 30 operating parameters such as voltage, current, and power are collected at a cycle of 100 ms. The time series data is decomposed into 6 levels by wavelet transform to extract fault feature vectors, which are reduced to 10 dimensions by combining with the independent component analysis (ICA) algorithm, and the random forest algorithm is used for fault classification, achieving an identification rate of over 95% for 10 typical faults. In anomaly detection, the DBSCAN spatio-temporal clustering algorithm is used. With a spatial radius of 100 meters and a time window of 1 minute, clustering analysis is carried out on the collected massive data to identify the spatio-temporal propagation trajectory of the fault chain reaction. Using the Apriori association rule mining algorithm, correlation analysis is carried out on fault events in different regions and at different times, and 100 association rules are mined from 10,000 historical data to construct a causal graph of fault propagation. In fault tracing and diagnosis, a Bayesian network is used to construct a causal inference model. By setting 80 causal nodes, based on multi-source information such as network topology, equipment parameters, and environmental factors, the fault probability distribution is calculated, and the root cause of the fault is comprehensively judged.In terms of fault prediction, the ARIMA model is used to extrapolate the trend of the fault sequence of the distribution network in the past three years. The model parameters are fitted through maximum likelihood estimation to predict the fault occurrence time in the next month, and the average error rate is controlled within 5%. Combining the fault prediction results, the time granularity of data collection is adaptively adjusted. During the early warning period, the data collection frequency is dynamically adjusted from 1 minute to 10 seconds to improve the timeliness of data and the capture probability of fault symptoms. At the same time, the Prophet algorithm is used to predict the fault duration, and the accuracy rate reaches 90%, providing a decision-making basis for emergency repair and maintenance plans. According to the equipment health status and load importance, with the fault risk and maintenance cost as the optimization objectives, the genetic algorithm is used to reconstruct the distribution network, solve the optimal switch combination, and on average reduce the power outage range by 20% and save 1 million yuan in maintenance costs per year.

[0047] Analyze the rationality of the reconstruction time interval; if frequent reconstruction is found, reduce the reconstruction interval time to avoid oscillation; if reconstruction lag is found, increase the reconstruction interval time to avoid service interruption.

[0048] By analyzing historical reconstruction data, identify abnormal fluctuations and change trends in the reconstruction frequency, and adaptively adjust the length of the reconstruction interval time to reduce unnecessary reconstruction operations while ensuring network stability; based on network state monitoring data, calculate network stability indicators, including voltage qualification rate and frequency deviation. When the network state is monitored to oscillate frequently or deviate from the steady state for a long time, dynamically shorten the reconstruction interval to quickly respond to network anomalies and suppress the further deterioration of oscillations; based on the timeliness requirements of network reconstruction, dynamically evaluate the reconstruction effect. By calculating indicators such as the improvement amplitude of network performance and the improvement degree of user-side voltage quality before and after reconstruction, judge whether the reconstruction is timely and effective. When the reconstruction effect is poor for several consecutive times, increase the reconstruction time interval to avoid service interruption and service quality degradation caused by frequent reconstruction.

[0049] Exemplarily, the time series analysis method, such as the moving average method, is used to analyze the reconstructed time interval data in the past year. It is found that the average value of the reconstruction interval is 30 days and the standard deviation is 5 days. The reconstruction interval in the most recent month has been shortened to 25 days, deviating from the mean by more than one standard deviation, indicating a decline in network stability. It is necessary to further shorten the reconstruction interval to 20 days to promptly respond to possible network anomalies. Based on the real-time monitored network status data, including the voltage data of network nodes collected in real time by deploying voltage sensors, the voltage qualification rate is calculated with a 10-minute cycle. When the voltage qualification rate is lower than 95% for three consecutive cycles, the reconstruction interval is shortened from 1 hour to 30 minutes to quickly respond to the voltage over-limit problem and avoid further deterioration of network stability. Regarding the timeliness requirement of network reconstruction, the network loss rate within one hour before and after reconstruction is compared. The average loss rate before reconstruction is 6%, and the average loss rate after reconstruction is reduced to 5%, with the loss rate reduced by 25%, achieving the expected reconstruction effect. However, the reduction amplitude of the loss rate in three consecutive reconstructions is less than 10%, and the reconstruction effect is limited. The reconstruction interval is automatically increased from 1 hour to 2 hours to avoid the impact of frequent reconstruction on business continuity.

[0050] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for fault diagnosis and adaptive reconstruction of a distribution network communication network, characterized in that: The method comprises: Monitor the distribution network communication network, obtain network status information, and extract key network performance indicators to establish a baseline for normal network operation. By comparing the current network status with the baseline in real time, potential abnormal fluctuations, including hardware failures, software failures, or configuration errors, can be identified; Develop corresponding hierarchical zoning strategies for different network scales and service types, including modular management for large-scale multi-layer networks and centralized response for small-scale networks, so as to effectively locate and isolate faults in various networks; Screening is performed based on preset fault modes, complex or unknown faults are predicted and identified, the fault range is centrally located to accurately determine the fault impact range, and the specific location of the fault, including nodes or connection points, is determined through network topology and fault propagation mode analysis to implement distributed isolation of the local distribution network; After the fault location is completed, the reconstruction object is identified, and a network reconstruction plan is formulated according to the identified reconstruction object and the current state of the network, including but not limited to replacing the faulty node by starting a backup resource, reestablishing the connection, or modifying the route; Update network parameters in real time, including adjusting channel allocation and power control, suppress new network conflicts and performance degradation caused by reconstruction, test the performance of the reconstructed network through a network simulator, verify whether the adjusted configuration meets the design requirements, and execute the reconstruction plan under the premise of meeting real-time and reliability requirements; After executing the reconstruction plan, the data collection and analysis cycle is adjusted according to the fault frequency, and the reconstruction interval is controlled to reduce redundant network structure adjustments and potential network fluctuations, including active detection and passive monitoring of the distribution network, and two-way monitoring of the network status.

2. The method according to claim 1, wherein: The monitoring of the distribution network communication network, obtaining network status information, and extracting key network performance indicators to establish a baseline for normal network operation, and identifying potential abnormal fluctuations, including hardware failures, software failures, or configuration errors, by comparing the current network status with the baseline in real time, include: Obtain the communication network status information of the distribution network, conduct correlation analysis on network performance indicators, and obtain the correlation and influence rules between different indicators; Based on the correlation analysis results of the network performance indicators, a network anomaly detection model is constructed, and the network anomaly detection model is trained and optimized; Deploy the optimized network anomaly detection model to network monitoring, analyze network status information in real time, and when abnormal fluctuations in network performance indicators are detected, trigger an alarm and notify network management personnel to conduct investigation and processing; The network fault diagnosis results and processing experience are structured and stored to form a distribution network communication network fault diagnosis knowledge base. When a network anomaly occurs, relevant cases can be retrieved to assist network management personnel in locating and processing the fault.

3. The method according to claim 1, wherein: The above-mentioned layered and partitioned strategies are formulated for different network scales and service types, including modular management for large-scale multi-layered networks and centralized response for small-scale networks, so as to effectively locate and isolate faults of various networks, including: According to different network scales and business types, identify the equipment, links and interconnection relationships in the distribution network communication network and generate a network topology diagram; By analyzing the connectivity and device attributes in the network topology diagram and combining the service traffic, different management domains and subnets are divided; For large-scale multi-layer networks, modular management is adopted; For small-scale networks, centralized management is implemented; When a monitored indicator deviates from the baseline range, an alarm is triggered, including analyzing the alarm data and locating the fault point. Locating the fault point includes analyzing the suspected fault point and narrowing the scope of investigation; Based on the fault location results, isolate the fault area to prevent the fault from spreading further; Match pre-defined typical fault scenarios and disposal plans, provide repair suggestions, and guide administrators to handle faults.

4. The method according to claim 1, wherein: The method of screening according to preset fault modes, predicting and identifying complex or unknown faults, centrally locating the fault range to accurately determine the fault impact range, and determining the specific location of the fault, including nodes or connection points, through network topology and fault propagation mode analysis, so as to implement distributed isolation of the local distribution network, includes: Establish a fault knowledge base, including typical fault scenarios of hardware faults, software faults, and configuration errors, and define the characteristic attributes and discrimination rules of each fault mode to form a fault knowledge base; Train historical fault data to establish a fault prediction model, extract key indicator change trends and abnormal patterns before a fault occurs, and predict and warn of potential faults in advance; For complex or unknown faults, the root cause and key nodes of the fault can be identified through correlation analysis of fault symptoms and impact range, and by reasoning and tracing the fault propagation path on the network topology diagram; After the fault alarm information and abnormal indicator data are aggregated, the fault location is analyzed, and the impact of the fault on network performance and business continuity is evaluated, and a fault impact report and handling suggestions are generated; According to the network topology and the logical connection relationship between devices, the fault impact range is divided and the specific location of the fault is determined. The specific location of the fault includes the faulty device, port and link.

5. The method according to claim 1, wherein: After the fault location is completed, the reconstruction object is identified, and a network reconstruction plan is formulated according to the identified reconstruction object and the current state of the network, including but not limited to replacing the faulty node by starting a backup resource, reestablishing the connection or modifying the route, including: After determining the specific location of the fault, the reconstruction object is identified, including by analyzing the type, location and functional attributes of the faulty node, as well as the status of the links and adjacent nodes connected to it; Determine whether reconstruction is necessary based on the role and importance of the faulty node in the network; Calculate the connectivity and availability of the network topology, evaluate the impact of faults on network service quality, and develop a reconstruction plan based on changes in network performance indicators, including: selecting backup nodes to replace failed nodes based on the configuration and real-time status of backup resources, and adjusting the number and distribution of backup nodes based on load balancing and disaster recovery protection requirements; When reestablishing the connection of a failed node, the link between the failed node and the adjacent nodes is reestablished and traffic is switched; Adjust the forwarding path and priority of data packets based on the service quality requirements of different businesses and the real-time status of network resources; It also includes: analyzing the availability of backup resources, screening available backup nodes and link resources; identifying the boundaries of the fault impact based on the results of the fault impact assessment, searching for backup paths outside the fault impact area, and finding the best data transmission alternative path by using the redundancy of the network structure; finding new links or relay nodes for disconnected connections based on the speed of connection point reconstruction, optimizing route modification flexibility, and dynamically updating the network routing table to adapt to the new network structure; based on the monitoring of the current state of the network, retrieving the real-time operation data of the network, including the load of each node, bandwidth utilization and link stability indicators, to evaluate the feasibility of the reconstruction plan; The method comprises: analyzing the availability of backup resources, screening available backup nodes and link resources; identifying the boundary of the fault impact according to the result of the fault impact assessment, searching for backup paths outside the fault impact area, and finding the best data transmission alternative path by using the redundancy of the network structure, specifically including: building a backup resource availability analysis model in combination with the hardware configuration, software version, and load status of the backup node, calculating the availability score of each backup node, and screening out backup nodes with high availability as candidate switching objects according to a preset threshold, and evaluating the performance indicators of the bandwidth, delay, and packet loss rate of the backup link at the same time, and determining the set of available links that meet the data transmission requirements; taking the fault node as the source point and the backup node as the sink point, obtaining the maximum transmission capacity from the fault node to each backup node, and using this as the basis for evaluating the switching priority of the backup node; identifying the boundary area affected by the fault, and extracting the topological relationship between the boundary nodes and the links; starting from the fault impact boundary, starting the backup path search in its external area, and finding multiple alternative transmission paths from the boundary node to the target node in combination with multiple constraints such as the bandwidth, delay, and reliability of the link, and sorting them according to the path performance indicators, and preferentially selecting the path with the best indicator as the main path for data transmission, and other paths as backups; The method combines the reconstruction speed of the connection point to find a new link or relay node for the disconnected connection, optimizes the flexibility of route modification, and dynamically updates the network routing table to adapt to the new network structure, specifically including: combining the processing capacity of the connection point, the bandwidth of the backup link, and the type of routing protocol, estimating the reconstruction time of the connection point, obtaining the probability distribution of the reconstruction time, and screening out the candidate connection points and links that meet the reconstruction speed requirements according to the set delay threshold; for the disconnected connection, by analyzing the location of the link disconnection point and the connection requirements, selecting the link with the minimum cost and the best path from the candidate link set as the new connection channel; if the resources of the candidate link are insufficient or do not meet the requirements, introducing a relay node, and forming a composite link across multiple hops through the forwarding capability of the relay node; searching for the optimal relay node combination in the network topology, evaluating the performance of the composite link, weighing the transmission delay and reliability, and selecting the optimal relay solution; The method monitors the current state of the network and retrieves the real-time operation data of the network, including the load of each node, bandwidth utilization and link stability indicators, to evaluate the feasibility of the reconstruction plan, specifically including: monitoring the current state of the network and retrieving the real-time operation data of the network; combining the link's packet loss rate, delay jitter, and bit error rate indicators to evaluate link stability, calculate the health score of each link, and set a health threshold. When the link score is lower than the health threshold, link reconstruction or backup link switching is triggered; constructing a simulation model similar to the actual network topology and configuration parameters, simulating the network behavior and performance after the reconstruction plan is executed by inputting real-time monitoring data, evaluating the impact of reconstruction on network service quality, and selecting the optimal reconstruction plan based on the evaluation results.

6. The method according to claim 1, wherein: The real-time updating of network parameters includes adjusting channel allocation and power control, suppressing new network conflicts and performance degradation caused by reconstruction, testing the performance of the reconstructed network through a network simulator, verifying whether the adjusted configuration meets the design requirements, and executing the reconstruction plan under the premise of meeting real-time and reliability requirements, including: Based on the status information of network devices and according to the requirements of network reconstruction, adjust the configuration parameters of devices, including routing tables and access control lists, to minimize the impact of reconstruction on services; Through spectrum sensing and channel state prediction, channel allocation and power control are optimized, dynamically adapting to changes in network topology and propagation environment, and suppressing co-channel interference and performance degradation caused by reconstruction; Conduct impact analysis on the reconstruction plan and formulate corresponding risk response strategies and rollback plans; Use the network simulation platform to evaluate and verify the performance of the reconstructed network, including key indicators such as throughput, latency, packet loss rate, and fairness; By comparing the performance differences before and after reconstruction, we can optimize the reconstruction plan and verify through simulation tests whether the reconstructed network meets the design requirements and business needs. If the reconstructed network meets the design requirements and business needs, the reconstruction plan is implemented; When implementing network reconstruction, the reconstruction process is divided into multiple stages, and only some network devices or links are adjusted in each stage to reduce the complexity and risk of reconstruction.

7. The method according to claim 1, wherein: After the reconstruction scheme is executed, the data collection and analysis cycle is adjusted according to the fault frequency, and the reconstruction interval is controlled to reduce redundant network structure adjustment and potential network fluctuations, including active detection and passive monitoring of the distribution network, and two-way monitoring of the network status, including: Combine the topology, asset status and environmental factors of the distribution network to conduct fault prediction and risk assessment, including simulating network behavior and impact scope under various fault scenarios, and assessing the fault risk level and loss extent; Adjust the time window and granularity of data collection and analysis by analyzing the frequency, interval and trend of failures; Based on the results of fault prediction and risk assessment, the reconstruction plan and execution cycle of the distribution network are dynamically optimized. In areas with high fault incidence and important load nodes, the reconstruction interval is shortened to quickly respond to fault hazards and abnormal situations. In areas with low fault incidence and nodes with secondary load, the reconstruction interval is extended to reduce the reconstruction cost and the impact on network stability. Actively detect the monitoring blind spots and fault prediction uncertainties of the distribution network, and conduct non-contact inspection of assets to proactively discover hidden defects; At the same time, through active injection and feedback analysis of fault signals, the location and type of fault can be located; It also includes: analyzing the rationality of the reconstruction time interval; if frequent reconstruction is found, reducing the reconstruction interval to avoid oscillation; if reconstruction lag is found, increasing the reconstruction interval to avoid business interruption, specifically including: by analyzing historical reconstruction data, identifying abnormal fluctuations and changing trends in the reconstruction frequency, and adaptively adjusting the length of the reconstruction interval to reduce unnecessary reconstruction operations while ensuring network stability; based on network status monitoring data, calculating network stability indicators, including voltage qualification rate and frequency deviation, and when it is monitored that the network status frequently oscillates or deviates from the steady state for a long time, dynamically shortening the reconstruction interval to quickly respond to network anomalies and suppress further deterioration of oscillation; based on the timeliness requirements of network reconstruction, dynamically evaluating the reconstruction effect, and judging whether the reconstruction is timely and effective by calculating the improvement in network performance before and after reconstruction and the improvement degree of user-side voltage quality indicators, when the effect of multiple consecutive reconstructions is poor, increasing the reconstruction time interval to avoid business interruption and service quality degradation caused by frequent reconstruction.

Citation Information

Patent Citations

  • Method and device for detecting network node abnormity

    CN110166271A

  • Intelligent fault sensing analysis processing method suitable for special communication system

    CN114710391A

  • Historical station data redundancy multi-process acquisition method and system

    CN118509134A

  • Intelligent diagnosis and isolation device and method for line fault of power distribution network

    CN118607390A

  • Power distribution network fault positioning method based on artificial intelligence and storage medium

    CN118884129A

Cited By

  • Planning method and system of power distribution communication access network

    CN120301761A

  • Communication route optimization method, system and equipment suitable for power distribution network

    CN120358567A

  • Emergency rescue data acquisition system and method suitable for underground engineering

    CN120430589A

  • Network self-healing device based on node state data

    CN120434108A

  • A network self-healing device based on node state data

    CN120434108B