Network fault repairing method and device, nonvolatile storage medium and electronic equipment

Through dynamic graph neural network and multi-agent collaborative system, the network topology diagram is updated in real time, and fault repair methods are automatically detected and verified, which solves the problem of low network fault detection and repair efficiency in the existing technology, and achieves efficient and secure network operation and maintenance.

CN120474895APending Publication Date: 2025-08-12CHINA TELECOM CORP LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510560886.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, the detection and repair of network failures relies on manual inspection, resulting in inefficiency and the lack of real-time and cross-vendor collaboration capabilities of automated repair tools.

Method used

By collecting the status data and topological connection relationships of network equipment, the dynamic graph neural network is used to update the topological graph in real time, and combining the multi-agent collaborative system, it automatically detects faults and verifies the repair method, including timing embedding modules, multi-scale graph attention layer and topological reconstruction modules, to achieve accurate positioning and repair of faults.

Benefits of technology

It realizes automated detection and repair of network failures, improves fault detection accuracy, ensures the safety of repair operations, and supports seamless collaboration across manufacturers' equipment, improving network operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474895A_ABST
    Figure CN120474895A_ABST
Patent Text Reader

Abstract

The invention discloses a network fault repairing method and device, a nonvolatile storage medium and electronic equipment. The method comprises the steps that state data of all network devices in a system and topological connection relations among all the network devices are collected, a network topological graph is updated according to the state data and the topological connection relations, nodes in the network topological graph correspond to the network devices, and edges in the network topological graph correspond to connection links among the network devices; determining an abnormal probability predicted value of each node in the network topological graph and a health state score of each edge; determining whether a fault occurs or not according to the abnormal probability predicted value and the health state score, and determining a repair mode corresponding to the fault after determining that the fault occurs; and verifying the repair mode through a system simulation model corresponding to the system, and repairing the fault by adopting the repair mode after verification is passed. The technical problem of low network fault repair efficiency caused by manual inspection in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network operation and maintenance, and specifically to a network fault repair method, device, non-volatile storage medium and electronic device. Background Art

[0002] The stable operation of networks plays a vital role in supporting social and economic development. However, due to the complexity of network equipment and network structure, the detection and repair of network faults has become a difficult problem in current network operations and maintenance. In most cases, fault detection and repair rely on manual labor, consuming a large amount of manpower and material resources, and resulting in low network operation and maintenance efficiency.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a network fault repair method, apparatus, non-volatile storage medium, and electronic device to at least solve the technical problem of low network fault repair efficiency caused by manual inspection in related technologies.

[0005] According to one aspect of an embodiment of the present application, a network fault repair method is provided, including: collecting status data of each network device in the system, and the topological connection relationship between each network device, and updating a network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; determining an abnormality probability prediction value of each node in the network topology map, and a health status score of each edge; determining whether a fault has occurred based on the abnormality probability prediction value and the health status score, and after determining that a fault has occurred, determining a repair method corresponding to the fault; verifying the repair method through a system simulation model corresponding to the system, and using the repair method to repair the fault after the verification is passed.

[0006] Optionally, determining the abnormality probability prediction value of each node in the network topology diagram and the health status score of each edge includes: determining the local characteristics of each node and the corresponding global characteristics of the system based on the status data, wherein the local characteristics include the device status corresponding to the node's adjacent nodes, and the global characteristics include the global load distribution status of the system; determining the local feature weights corresponding to the local features and the global feature weights corresponding to the global features; determining the abnormality probability prediction value of each node and the health status score of each edge based on the local characteristics and the local feature weights, and the global characteristics and the global feature weights.

[0007] Optionally, determining whether a fault has occurred based on the abnormal probability prediction value and the health status score includes: determining historical fault data, wherein the historical fault data includes historical abnormal probability prediction values and historical health status scores; determining a first distribution of historical abnormal probability prediction values, and a second distribution of historical health status scores; determining a first threshold based on the first distribution, and determining a second threshold based on the second distribution; comparing the abnormal probability prediction value of each node with the first threshold, and determining whether there is a faulty node based on the comparison result, and comparing the health status score of each edge with the second threshold, and determining whether there is a faulty edge based on the comparison result.

[0008] Optionally, determining the repair method corresponding to the fault includes: determining the fault type of the fault, as well as the status data and link status of each network device in the system, and determining the general fault repair instruction corresponding to the fault; determining the intermediate representation instruction corresponding to the general fault repair instruction, wherein the intermediate representation instruction is a fault repair instruction that is independent of the device type, and the intermediate representation instruction includes a fault repair parameter for indicating the fault repair method; in the case where the fault type is a device fault, determining the device type of the faulty device, and determining the instruction template corresponding to the device type; determining the device fault repair instruction based on the intermediate representation instruction and the instruction template, wherein the fault repair instruction is used to indicate the fault repair method.

[0009] Optionally, determining the repair method corresponding to the fault includes: determining multiple alternative repair methods corresponding to the fault, and a set of evaluation values of the alternative repair methods, wherein the evaluation value set includes a safety evaluation value, a performance evaluation value, and a cost evaluation value; determining a weight set of the system, wherein the weight set includes a safety weight, a performance weight, and a cost weight; and determining a target repair method corresponding to the fault from multiple alternative repair methods based on the weight set and the evaluation value sets corresponding to each alternative repair method.

[0010] Optionally, determining the weight set of the system includes: determining a network status of the system, and adjusting weights in the weight set according to the network status, wherein the network status includes at least one of the following: encountering a security incident, and business load conditions.

[0011] Optionally, determining the repair method corresponding to the fault includes: determining the repair method corresponding to the fault through preset expert rules; or, determining the repair method corresponding to the fault through a deep learning network, wherein the deep learning network is used to determine the repair instructions corresponding to the fault in the action space based on the state space for indicating the repair method, the state space includes the network topology state of the system, the equipment load condition and the fault type of the fault, and the action space includes a set of executable alternative repair instructions.

[0012] Optionally, verifying the repair method through the system simulation model corresponding to the system includes: performing routing loop detection on the repair method through the system simulation model to determine whether there is a loop path in the system after the repair method is executed; performing bandwidth congestion detection on the repair method through the system simulation model to determine whether there is a link whose utilization exceeds a preset threshold after the repair method is executed; when the detection result is that there is no loop path and no link whose utilization exceeds the preset threshold, it is determined that the repair method has passed the verification.

[0013] According to another aspect of an embodiment of the present application, a network fault repair device is also provided, including: a first processing module, used to collect status data of each network device in the system, and the topological connection relationship between each network device, and update the network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; a second processing module, used to determine the abnormality probability prediction value of each node in the network topology map, and the health status score of each edge; a third processing module, used to determine whether a fault occurs based on the abnormality probability prediction value and the health status score, and after determining that a fault occurs, determine the repair method corresponding to the fault; a fourth processing module, used to verify the repair method through the system simulation model corresponding to the system, and use the repair method to repair the fault after the verification is passed.

[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is running, the device where the non-volatile storage medium is located is controlled to execute a network fault repair method.

[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the network fault repair method is executed when the program is run.

[0016] According to another aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which implements a network fault repair method when executed by a processor.

[0017] In an embodiment of the present application, the status data of each network device in the collection system and the topological connection relationship between each network device are collected, and the network topology map is updated based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; the abnormal probability prediction value of each node in the network topology map and the health status score of each edge are determined; whether a fault occurs is determined based on the abnormal probability prediction value and the health status score, and after determining that a fault occurs, the corresponding repair method of the fault is determined; the repair method is verified by the system simulation model corresponding to the system, and the repair method is used to repair the fault after the verification is passed. By constructing a multi-agent collaborative system, the failure probability of network devices and connection links is determined based on the network topology, and then whether a fault occurs and the repair method is determined and verified, the purpose of automated fault detection, repair, and verification is achieved, thereby achieving the technical effect of improving fault detection accuracy, ensuring the safety of repair operations and seamless collaboration of cross-vendor equipment, and thus solving the technical problem of low efficiency of network fault repair caused by manual inspections in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 is a structural diagram of a computer terminal provided according to an embodiment of the present application;

[0020] Figure 2 This is a flowchart of a network fault repair method provided in accordance with an embodiment of the present application;

[0021] Figure 3 is a schematic diagram of a network fault repair system provided according to an embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of a process for detecting an intelligent agent according to an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of the organization of a multi-agent association according to an embodiment of the present application;

[0024] Figure 6 It is a structural diagram of a network fault repair device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0028] Dynamic Graph Neural Network (Dynamic GNN): A neural network model that can update the network topology graph in real time and capture temporal changes.

[0029] Multi-Agent Collaboration: Multiple agents work together to complete complex tasks.

[0030] Digital Twin: A virtual simulation environment built based on physical network parameters, used to pre-verify the safety of repair operations.

[0031] Temporal Embedding Module: A module based on recurrent neural networks (such as LSTM) to capture temporal changes.

[0032] Multi-Scale Graph Attention Layer: A graph neural network layer that fuses local connection states with global network load characteristics to dynamically adjust node and edge weights.

[0033] In related technologies, network fault detection and repair mainly rely on the following technical solutions:

[0034] (1) Manual operation and maintenance mode

[0035] Implementation method: Operation and maintenance personnel receive alarm information through the monitoring system, manually log in to the device to execute diagnostic commands, and write repair scripts based on experience.

[0036] limitation:

[0037] (1) High latency: Manual response results in a mean time to repair (MTTR) of several hours;

[0038] (2) High cost: A professional team is required to be on duty 24 hours a day, and labor costs account for more than 40% of the total operation and maintenance expenses.

[0039] (2) Single detection model based on traditional machine learning

[0040] Implementation: Use supervised learning algorithms (such as random forests and support vector machines) or time series analysis models (such as ARIMA) to train fault classification models based on historical data.

[0041] limitation:

[0042] (1) Static rule dependence: The model relies on fixed feature engineering and cannot adapt to dynamic topological changes;

[0043] (2) Detection and repair are disconnected: only alarm information is output, and there is a lack of automated repair capabilities.

[0044] Although the above technologies have improved network operation and maintenance efficiency to a certain extent, the following core issues still exist:

[0045] (1) Insufficient adaptability of dynamic topology

[0046] Existing technologies rely on static topology maps and are unable to capture dynamic changes in network device status (such as traffic and latency) in real time, making it difficult to detect hidden faults (such as cascading interruptions).

[0047] (2) Repair operations lack security verification

[0048] Existing automated repair tools directly operate on physical devices, which may cause secondary failures (such as routing loops and configuration overlays) due to command conflicts and lack a pre-verification mechanism.

[0049] (3) Difficulty in cross-manufacturer equipment collaboration

[0050] Heterogeneous devices from different manufacturers use their own proprietary protocols, and repair instructions cannot be used across manufacturers, making automated repair difficult to achieve.

[0051] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0052] According to an embodiment of the present application, a method embodiment of a network fault repair method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0053] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a network fault repair method. Figure 1 As shown, the computer terminal 10 may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0054] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0055] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the network fault repair method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned network fault repair method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0056] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0057] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0058] In the above operating environment, the embodiment of the present application provides a network fault repair method, such as Figure 2 As shown, the method includes the following steps:

[0059] Step S202: collect status data of each network device in the system and the topological connection relationship between each network device, and update the network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices and the edges correspond to the connection links between the network devices.

[0060] As an optional implementation method, by detecting the status data of each network device in the collection system, as well as the topological connection relationship between each network device and updating the topology map, the detection agent mainly realizes real-time update of the network topology map based on the dynamic GNN model, identifies hidden faults that are difficult to detect with traditional threshold detection, and combines local connection status (such as adjacent device traffic) with global network load (such as full network bandwidth utilization) to improve the comprehensiveness and accuracy of fault location. Dynamic GNN continuously updates the network topology map by collecting network device status data (such as traffic, delay, packet loss rate) and topological connection relationships in real time. Dynamic GNN is mainly divided into three modules: Temporal Embedding Module, Multi-Scale Graph Attention Layer and Topology Reconstruction Module.

[0061] Among them, the timing embedding module is used to capture the timing changes of the network topology (such as device status fluctuations and link connection / disconnection). The timing embedding module uses a recurrent neural network to perform timing modeling on the device status data and generate dynamic node features. For example, if the switch port traffic suddenly drops from 100Mbps to 0, the LSTM (long short-term memory network) module in the timing embedding module will identify this abnormal timing pattern and update the feature vector of the node.

[0062] Feature vectors play a core role in the time series embedding module. They are the output of a recurrent neural network (such as an LSTM) after performing time series analysis on network device states. They represent the multidimensional state information of a device at a specific point in time. Feature vectors capture dynamic changes in device states and convert these changes into machine-understandable numerical values, providing critical data support for fault detection, network analysis, and even predictive decision-making.

[0063] A feature vector can be considered a dynamic representation of a node (i.e., a device in a network). Each node has a unique feature vector that contains various device attributes, such as traffic flow, latency, packet loss rate, CPU utilization, and memory usage. These attributes form a comprehensive description of the device's state. Through processing in the LSTM module, the feature vector can further include the time-varying trends of the device's state, i.e., time series information.

[0064] When an abnormal event occurs, such as a sudden drop in switch port traffic from 100 Mbps to 0, the LSTM module can keenly identify this temporal pattern and, through its internal memory cells and gating mechanism, incorporate this change into the feature vector. The updated feature vector not only reflects the latest traffic values but, more importantly, contains the context of the sudden change—when it occurred, how drastic the change was, and whether it constitutes an anomaly. This information is crucial for subsequent fault detection and remediation decisions.

[0065] Dynamic updates of feature vectors ensure that the dynamic graph neural network can capture subtle changes in device status in real time, enabling rapid response to early faults. This is crucial for automated network operations and maintenance. By continuously analyzing and updating feature vectors, the time series embedding module provides a more accurate and comprehensive description of network status for the network fault self-healing system, helping the system identify and repair faults promptly, avoiding potential network performance degradation and service interruptions.

[0066] Among them, the multi-scale graph attention layer is used to fuse local connection status and global network load characteristics. Local features focus on the status of the node's direct neighbors (such as the traffic load of adjacent switches); global features analyze the load distribution of the entire network (such as the bandwidth utilization of core routers). The multi-scale graph attention layer dynamically allocates weights through the attention mechanism to determine the contribution ratio of local and global features.

[0067] Among them, the Topology Reconstruction Module is used to generate a real-time updated topology map based on the dynamic node features and connection relationships captured by the timing embedding module, and dynamically adjust the edge weights based on the similarity of node features (such as increasing the edge weights between devices with similar traffic); if it is detected that the traffic on a link has returned to zero, the edge is marked as "disconnected".

[0068] Step S204: determining the abnormal probability prediction value of each node in the network topology graph and the health status score of each edge.

[0069] In the technical solution provided in step S204, determining the abnormal probability prediction value of each node in the network topology diagram and the health status score of each edge includes: determining the local characteristics of each node and the global characteristics corresponding to the system based on the status data, wherein the local characteristics include the device status corresponding to the adjacent nodes of the node, and the global characteristics include the global load distribution status of the system; determining the local feature weights corresponding to the local features and the global feature weights corresponding to the global features; determining the abnormal probability prediction value of each node and the health status score of each edge based on the local characteristics and the local feature weights, and the global characteristics and the global feature weights.

[0070] Optionally, the timing embedding module extracts local node features from state data (including traffic, latency, and packet loss rate). This includes the device status of the node's immediate neighbors, such as the traffic load, latency, or packet loss rate of adjacent switches, among other key performance indicators. At the same time, it focuses on the system's global characteristics, which typically include the load distribution of the entire network, such as the bandwidth utilization of core routers and the traffic inflow and outflow ratio of data centers, reflecting overall network health and performance.

[0071] The multi-scale graph attention layer utilizes an attention mechanism to determine the local feature weights for local features and the global feature weights for global features. The core of the attention mechanism lies in its ability to filter the most relevant information from a large amount of information. By calculating the correlation between local and global features, it dynamically assigns weights to ensure that both local and global features receive appropriate attention under varying network conditions. For example, in situations where network traffic fluctuates significantly, global feature weights may be higher, as this may indicate a network-wide event. Conversely, when local device status is abnormal, local feature weights are emphasized to help the system focus on the detailed conditions in that area. When abnormal bandwidth utilization is detected in a core router, the attention mechanism automatically increases the local feature weights for switches directly connected to that router, enabling a more detailed analysis of the status of these nodes and providing a more accurate basis for fault location and repair strategy formulation.

[0072] Then, based on local features and their weights, as well as global features and their corresponding weights, we can determine the abnormality probability prediction value for each node and the health status score for each edge. Specifically, the abnormality probability prediction value for a node is obtained by weightedly fusing local and global features, and then transforming it through an activation function (such as Softmax), resulting in a probability value between 0 and 1, indicating the probability of the node being abnormal. The health status score of an edge is also calculated based on the status of the connecting link and the overall network load, using a scoring function to reflect the current health level and potential risk of the link.

[0073] As an optional implementation, local features and their weights, as well as global features and their corresponding weights, can also be input into the evaluation model to obtain the abnormal probability prediction value of each node and the health status score of each edge output by the evaluation model. The evaluation model is obtained through historical data and label training. Specifically, first, operating data including indicators such as device traffic, latency, and CPU utilization are collected. At the same time, based on past operation and maintenance records, each data point is labeled. For node data, the label reflects whether the device is in a faulty state; for edge data, the health level of the link is marked. During the training process, the model gradually optimizes the parameters by comparing the prediction results with the actual labels to improve the prediction accuracy. After training is completed, the evaluation model can output the abnormal probability prediction value of each node (between 0 and 1, the higher the value, the greater the possibility of failure) and the health status score of each edge (also between 0 and 1, the lower the value, the worse the link quality) based on real-time network data.

[0074] Implementation of this process relies not only on accurate state data collection but also on close collaboration among the various modules of the dynamic graph neural network: the time-series embedding module provides information on the temporal changes in node states, the topology reconstruction module ensures real-time updates of the topology graph, and the multi-scale graph attention layer is responsible for dynamically adjusting the contribution ratio of local and global features based on this information. In this way, the system can capture subtle changes in network device status in real time, improving the accuracy and timeliness of fault detection.

[0075] Step S206 , determining whether a fault occurs based on the abnormal probability prediction value and the health status score, and after determining that a fault occurs, determining a repair method corresponding to the fault.

[0076] In the technical solution provided in step S206, determining whether a fault has occurred based on the abnormal probability prediction value and the health status score includes: determining historical fault data, wherein the historical fault data includes historical abnormal probability prediction values and historical health status scores; determining a first distribution of historical abnormal probability prediction values, and a second distribution of historical health status scores; determining a first threshold based on the first distribution, and determining a second threshold based on the second distribution; comparing the abnormal probability prediction value of each node with the first threshold, and determining whether there is a faulty node based on the comparison result, and comparing the health status score of each edge with the second threshold, and determining whether there is a faulty edge based on the comparison result.

[0077] Optionally, based on a pre-set sliding time window, historical fault data (e.g., 5 minutes) for the last N minutes (including historical abnormal probability prediction values and historical health status scores) is selected, and the mean μ and standard deviation σ of the historical abnormal probability prediction values (i.e., the first distribution situation) are calculated, thereby determining a first threshold μ+3σ based on the first distribution situation, and the mean μ and standard deviation σ of the historical health status scores (i.e., the second distribution situation) are calculated, thereby determining a second threshold μ-3σ based on the second distribution situation. Comparing the abnormal probability prediction value of each node with the first threshold, and determining whether there is a faulty node based on the comparison result includes: if the abnormal probability prediction value is greater than the first threshold, determining that a faulty node exists; comparing the health status score of each edge with the second threshold, and determining whether there is a faulty edge based on the comparison result includes: if the health status score is less than the second threshold, determining that there is a faulty edge.

[0078] As an optional implementation, the sliding time window is adjusted based on the verification results of the repair method by the subsequent system simulation model. For example, in the case of a false positive, it usually indicates that the current sliding window may be too large, causing the model to focus too much on the long-term trend of historical data and ignore recent dynamic changes. To reduce false positives, the system will reduce the window size so that the model pays more attention to recent data and increases sensitivity to temporary or sudden state changes. If a false negative occurs, it usually means that the current window may be too small and the model cannot effectively capture the long-term trend that is sufficient to form a failure mode. In this case, the system will increase the window size to ensure that sufficient information can be extracted from a longer time series to reduce the possibility of false negatives.

[0079] As an optional implementation, an alarm threshold can also be set to issue a fault alarm when the number of consecutive failures exceeds the threshold, and the alarm threshold can be adjusted based on the verification results of the repair method according to the system simulation model: in the case of a false alarm, the alarm threshold is increased, and in the case of a missed alarm, the alarm threshold is lowered.

[0080] In the technical solution provided in step S206, determining the repair method corresponding to the fault includes: determining the fault type of the fault, as well as the status data and link status of each network device in the system, and determining the general fault repair instruction corresponding to the fault; determining the intermediate representation instruction corresponding to the general fault repair instruction, wherein the intermediate representation instruction is a fault repair instruction that is independent of the device type, and the intermediate representation instruction includes a fault repair parameter for indicating the fault repair method; in the case where the fault type is a device fault, determining the device type of the faulty device and determining the instruction template corresponding to the device type; determining the device fault repair instruction based on the intermediate representation instruction and the instruction template, wherein the fault repair instruction is used to indicate the fault repair method.

[0081] Optionally, after a fault is identified, the system retrieves the state information of the faulty node or edge and uses this information to determine the fault type. For example, if the detection agent, through dynamic graph neural network analysis, reports a significant increase in the predicted abnormality probability for a core router, this could indicate a device overload, misconfiguration, or hardware failure. The system then determines the specific fault type based on historical data, combined with the device's abnormality patterns and relevant parameter indicators such as CPU utilization, memory usage, and network traffic.

[0082] Optionally, the protocol semantic parser is used to determine the intermediate representation instruction corresponding to the general fault repair instruction, including parsing the repair instruction into a device-independent intermediate representation through semantic parsing.

[0083] Optionally, the manufacturer protocol adapter is used to determine the corresponding instruction template based on the device type (such as the manufacturer) of the faulty device, and the intermediate representation instruction and the instruction template determine the device fault repair instruction, extract the fault repair parameters from the intermediate representation instruction and fill them into the instruction template.

[0084] In the technical solution provided in step S206, determining the repair method corresponding to the fault includes: determining multiple alternative repair methods corresponding to the fault, and a set of evaluation values of the alternative repair methods, wherein the evaluation value set includes a safety evaluation value, a performance evaluation value, and a cost evaluation value; determining a weight set of the system, wherein the weight set includes a safety weight, a performance weight, and a cost weight; and determining a target repair method corresponding to the fault from multiple alternative repair methods based on the weight set and the evaluation value sets corresponding to each alternative repair method.

[0085] As an optional implementation, determining the weight set of the system includes: determining the network status of the system, and adjusting the weights in the weight set according to the network status, wherein the network status includes at least one of the following: encountering a security incident, business load conditions.

[0086] Optionally, after determining the fault's corresponding alternative repair methods through pre-set expert rules or a deep learning network, multiple alternative repair methods are identified. Each alternative repair method is assigned a set of evaluation values, including a safety evaluation value, a performance evaluation value, and a cost evaluation value. The safety evaluation value reflects the impact of the repair operation on network stability and data security; the performance evaluation value describes the performance change of the system after executing the repair method, that is, the performance change of the system after executing the repair method; and the cost evaluation value takes into account the resource consumption and economic cost of executing the operation.

[0087] In the process of determining the system weight set, the system will make intelligent adjustments based on the current network status to ensure that the repair method selected in a specific scenario can give priority to meeting the most urgent needs. For example, when suffering a security incident (such as a DDoS attack), the system security weight will be significantly increased, which means that when choosing a repair method, the security evaluation value will become the decisive factor. Even if other repair methods are better in performance or cost, they will be excluded if there are security risks. During business peaks, the system may automatically adjust the performance weight to a high position to ensure high network availability and response speed. Even if some repair methods may be more expensive or have slightly greater security risks, in order to maintain business continuity and user experience, the performance evaluation value will be given a higher weight.

[0088] Based on the set of weights and the corresponding evaluation values for each alternative repair method, the system calculates a comprehensive evaluation score for each repair method. This calculation process may be based on a weighted multi-objective optimization algorithm, multiplying the security evaluation value, performance evaluation value, and cost evaluation value by the corresponding weights and summing them to obtain the final evaluation value. The system compares these comprehensive evaluation values and selects the repair method with the highest score as the target repair method, which is then executed by the repair agent. For example, if during a fault detection, the system increases the security weight to 80%, reduces the performance weight to 15%, and maintains the cost weight at 5% based on the current network status of the security incident, then even if a repair method is not optimal in terms of cost and performance, if it can significantly improve network security, it will be selected as the target repair method because the weight of the security evaluation value dominates the total evaluation value.

[0089] This strategy ensures that the network fault self-healing system's decision-making is both flexible and intelligent. It can dynamically optimize the selection of repair methods based on the needs of different scenarios, such as emergency handling of security incidents or performance assurance during peak business periods, thereby finding the optimal balance between security, performance, and cost, and achieving efficient and reliable network fault handling.

[0090] In the technical solution provided in step S206, determining the repair method corresponding to the fault includes: determining the repair method corresponding to the fault through preset expert rules; or determining the repair method corresponding to the fault through a deep learning network, wherein the deep learning network is used to determine the repair instructions corresponding to the fault in the action space based on the state space for indicating the repair method, the state space includes the network topology state of the system, the equipment load condition and the fault type of the fault, and the action space includes a set of executable alternative repair instructions.

[0091] Specifically, the preset expert rules are deterministic rules generated based on network operation and maintenance experience, for example:

[0092] def Generate link switching instructions (fault type, topology status):

[0093] If the fault type == "link down" and a backup link exists:

[0094] return "Switch to backup link"

[0095] elif fault type == "device overload" and adjacent device load < critical value:

[0096] return "Start load balancing"

[0097] That is, the rule is to receive the fault type and topology status, and switch to the backup link if the fault type is link interruption and there is a backup link; if the fault type is device overload and the load of the adjacent device is less than the critical value, load balancing is started.

[0098] The state space of the deep learning network is the network topology, device load, and fault type. The action space is the set of executable repair instructions. The training method is to learn the optimal strategy online based on the Deep Q Network (DQN). The reward function is as follows:

[0099] Reward = α × repair success rate + β × repair speed + γ × resource consumption

[0100] Among them, α, β, γ are preset weights, and α+β+γ=1.

[0101] Step S208: Verify the repair method through the system simulation model corresponding to the system, and use the repair method to repair the fault after passing the verification.

[0102] In the technical solution provided in step S208, verifying the repair method through the system simulation model corresponding to the system includes: performing routing loop detection on the repair method through the system simulation model to determine whether there is a loop path in the system after the repair method is executed; performing bandwidth congestion detection on the repair method through the system simulation model to determine whether there is a link whose utilization exceeds a preset threshold after the repair method is executed; if the detection result is that there is no loop path and no link whose utilization exceeds the preset threshold, it is determined that the repair method has passed the verification.

[0103] Optionally, the system simulation model is built using digital twin technology: a high-fidelity network simulation environment is built using NS3 or Mininet, and configurations (such as routing tables and VLAN settings) are synchronized from physical devices to the digital twin.

[0104] Optionally, routing loop detection checks whether there is a loop path based on a graph traversal algorithm (such as DFS). Bandwidth congestion detection simulates traffic load to detect whether link utilization exceeds a preset threshold (such as >90%).

[0105] Optionally, Figure 3 A network fault recovery system is provided, such as Figure 3 As shown, the system includes the following modules: detection agent 01, repair agent 02, verification agent 03, feedback controller 04. This system completes the end-to-end automated repair process through the collaboration of the three agents: detection agent, repair agent, and verification agent.

[0106] Figure 4 A workflow diagram for detecting an intelligent agent is shown in FIG. Figure 4 As shown, the process includes the following steps:

[0107] S401 Data Collection: Real-time collection of network device status data (such as traffic, delay, packet loss rate) and topological connection relationships.

[0108] S402 Dynamic GNN: Dynamic GNN is mainly divided into three modules: time series embedding module, multi-scale graph attention layer and topology reconstruction module. It is used to continuously update the network topology map, capture the time series changes of device status (such as link load fluctuations, sudden equipment failures), and finally output the abnormal probability value of each node (device) and the health status score of each edge. Among them, the time series embedding module analyzes the time series changes of device status through the LSTM module to generate dynamic node features. The multi-scale graph attention layer is used to calculate the association weights between nodes. The local weight includes the traffic correlation of adjacent switches, and the global weight includes the load balancing status of the entire network. The topology reconstruction module generates a new topology map based on the updated node features and edge weight relationship (such as disconnecting abnormal links and increasing the weight of backup links).

[0109] S403 Adjust threshold: Calculate the adjustment threshold based on the output of dynamic GNN

[0110] Algorithm principle:

[0111] (1). Input: GNN output data + historical fault data.

[0112] (2) Sliding window statistics: Based on the data of the last N minutes (e.g. 5 minutes), calculate the mean (μ) and standard deviation (σ) of the indicator.

[0113] (3) Adaptive threshold calculation:

[0114] Initially

[0115] μ-3σ,μ+3σ

[0116] If anomalies occur M times consecutively (e.g., M=3), a fault alarm is triggered and the k value is reduced (e.g., k=2) to increase sensitivity.

[0117] (4) Feedback mechanism: According to the results of the verification agent (false positive / false negative), the k value and window size are adjusted in reverse.

[0118] S404 Abnormality determination and output: If the node abnormality probability or edge health score output by the dynamic GNN exceeds the threshold range, it is marked as a fault.

[0119] Optionally, the repair agent generates repair instructions based on the fault information output by the detection agent, using an adaptive rule engine and multi-objective optimization strategy to achieve efficient instruction generation and protocol conversion for cross-vendor devices. Its core modules are as follows:

[0120]

[0121]

[0122] Among them, the protocol semantic analysis and adaptation process includes

[0123] (1) General Instruction → Intermediate Representation (IR)

[0124] Semantic parsing: Parsing fixup instructions into device-independent IR

[0125] (2) IR → Manufacturer-specific instructions

[0126] Dynamic template matching: Select a command template based on the target device type (such as manufacturer) and bind dynamic parameters.

[0127] Parameter binding: Extract parameters such as source_port, target_port, and vlan_id from IR and fill them into the template.

[0128] The input of the priority arbitration module is the candidate instruction set generated by multi-objective optimization and the real-time network policy weight. The decision-making process is as follows:

[0129] Dynamic weight adjustment: Automatically adjust target weights based on network status. For example, if there is a security incident (such as a DDoS attack), the security weight increases to 80%; if the service is at peak hours, the performance weight increases to 50%. The utility function calculation is as follows:

[0130] Utility = Σ(target score × target weight)

[0131] Select the best instruction: Select the instruction with the highest utility value and issue it for execution.

[0132] Optionally, the verification agent includes two parts: digital twin construction and conflict detection model. The simulation tool used to build the digital twin includes NS3 or Mininet, and parameter synchronization includes synchronizing configurations (such as routing tables and VLAN settings) from the physical device to the digital twin. The conflict detection model includes routing loop detection (checking for loop paths based on graph traversal algorithms such as DFS) and bandwidth congestion detection (simulating traffic load to detect whether link utilization exceeds a threshold (such as >90%)).

[0133] Optionally, Figure 5 The diagram of the organization of multi-agent association is shown in Figure 5 As shown, closed-loop feedback control and distributed decision optimization are adopted among multiple agents to automate the entire process of "detection-repair-verification", significantly improving the efficiency and security of network fault handling. The functions and technical implementations of each agent are shown in the following table:

[0134]

[0135]

[0136] The multi-agent positive feedback process includes detection → repair → verification:

[0137] The detection agent analyzes the network topology through dynamic GNN, locates the faulty nodes and outputs the fault type (such as link interruption, device overload).

[0138] The repair agent generates repair instructions based on a hybrid rule engine and converts them into the target device-specific protocol through a dynamic protocol adapter.

[0139] The verification agent simulates repair operations in the digital twin and detects potential conflicts (such as routing loops and bandwidth congestion). If the verification passes, the command is sent to the physical network; otherwise, the rollback mechanism is triggered.

[0140] Among them, the multi-agent positive feedback process includes feedback → optimization:

[0141] The feedback controller monitors the repair results (such as repair success rate and repair delay) and evaluates the performance of the agent.

[0142] Dynamically adjust the agent strategy based on the evaluation results:

[0143] Detection agent: adjust the anomaly detection threshold (e.g., lower the k value to increase sensitivity);

[0144] Repair the agent: update the rule base weight (such as increasing the priority of security rules);

[0145] Verification Agent: Optimize the conflict detection model (e.g., increase the accuracy of bandwidth congestion detection).

[0146] Through the above steps, a highly automated, efficient, and flexible network fault repair method can be implemented. The method embodiment of the present application can respond to network changes in real time, quickly locate faults, and comprehensively evaluate repair solutions from multiple perspectives, ensuring that repair costs are effectively controlled while improving network stability and performance. Specifically, the method embodiment of the present application has the following advantages:

[0147] (1) Improve fault detection accuracy and real-time performance

[0148] Through dynamic topology modeling technology, network device status changes are captured in real time, significantly improving the detection accuracy of hidden faults (such as cascading interruptions).

[0149] GNN captures the temporal changes of device status through a time-series embedding module (LSTM / GRU), the multi-scale graph attention layer fuses local connection status with global network load characteristics, and the topology reconstruction module updates the network topology map in real time (e.g., once per second). Compared with traditional detection methods, dynamic GNN supports real-time updates of network topology and detection of hidden faults (e.g., cascading interruptions), solving the pain point that static models cannot adapt to dynamic changes.

[0150] (2) Ensure the safety of repair operations

[0151] The introduction of a pre-verification mechanism avoids secondary failures caused by incorrect operations (such as routing loops and bandwidth congestion), greatly improving the success rate of repair operations.

[0152] (3) Achieve seamless collaboration of cross-vendor devices

[0153] Through lightweight protocol conversion technology, it supports instruction set compatibility of mainstream manufacturers' equipment, significantly reducing the time spent on cross-vendor troubleshooting.

[0154] (4) Reduce operation and maintenance costs and manual intervention

[0155] Through end-to-end automated repair processes, we can reduce dependence on professional operation and maintenance teams and lower labor costs.

[0156] Through a closed-loop feedback control mechanism involving multi-agent collaboration, the detection, repair, and verification agents interact with each other in real time via a message queue (e.g., Kafka). Furthermore, the feedback controller dynamically adjusts policies (e.g., anomaly thresholds, rule weights) based on the repair results to support multi-agent task allocation and load balancing. Compared to existing detection or repair methods, the method embodiments of this application implement a fully closed-loop "detection-repair-verification-feedback" process, breaking through the limitations of the separation of detection and repair.

[0157] The present invention provides a data storage device. Figure 6 is a structural diagram of the device, such as Figure 6As shown, the device includes: a first processing module 60, which is used to collect status data of each network device in the system, and the topological connection relationship between each network device, and update the network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; a second processing module 62, which is used to determine the abnormal probability prediction value of each node in the network topology map, and the health status score of each edge; a third processing module 64, which is used to determine whether a fault occurs based on the abnormal probability prediction value and the health status score, and after determining that a fault occurs, determine the corresponding repair method of the fault; a fourth processing module 66, which is used to verify the repair method through the system simulation model corresponding to the system, and adopt the repair method to repair the fault after the verification is passed.

[0158] In some embodiments of the present application, the second processing module 62 determines the abnormality probability prediction value of each node in the network topology diagram, and the health status score of each edge, including: determining the local characteristics of each node based on the status data, and the global characteristics corresponding to the system, wherein the local characteristics include the device status corresponding to the adjacent nodes of the node, and the global characteristics include the global load distribution status of the system; determining the local feature weights corresponding to the local features, and the global feature weights corresponding to the global features; determining the abnormality probability prediction value of each node, and the health status score of each edge based on the local characteristics and the local feature weights, and the global characteristics and the global feature weights.

[0159] In some embodiments of the present application, the third processing module 64 determines whether a fault has occurred based on the abnormal probability prediction value and the health status score, including: determining historical fault data, wherein the historical fault data includes historical abnormal probability prediction values and historical health status scores; determining a first distribution of historical abnormal probability prediction values, and a second distribution of historical health status scores; determining a first threshold based on the first distribution, and determining a second threshold based on the second distribution; comparing the abnormal probability prediction value of each node with the first threshold, and determining whether there is a faulty node based on the comparison result, and comparing the health status score and the second threshold of each edge, and determining whether there is a faulty edge based on the comparison result.

[0160] In some embodiments of the present application, the third processing module 64 determines the repair method corresponding to the fault, including: determining the fault type of the fault, as well as the status data and link status of each network device in the system, and determining the general fault repair instruction corresponding to the fault; determining the intermediate representation instruction corresponding to the general fault repair instruction, wherein the intermediate representation instruction is a fault repair instruction that is independent of the device type, and the intermediate representation instruction includes a fault repair parameter for indicating the fault repair method; in the case where the fault type is a device fault, determining the device type of the faulty device, and determining the instruction template corresponding to the device type; determining the device fault repair instruction based on the intermediate representation instruction and the instruction template, wherein the fault repair instruction is used to indicate the fault repair method.

[0161] In some embodiments of the present application, the third processing module 64 determines the repair method corresponding to the fault, including: determining multiple alternative repair methods corresponding to the fault, and a set of evaluation values of the alternative repair methods, wherein the evaluation value set includes a safety evaluation value, a performance evaluation value, and a cost evaluation value; determining a weight set of the system, wherein the weight set includes a safety weight, a performance weight, and a cost weight; and determining a target repair method corresponding to the fault from multiple alternative repair methods based on the weight set and the evaluation value set corresponding to each alternative repair method.

[0162] In some embodiments of the present application, the third processing module 64 determines the weight set of the system, including: determining the network status of the system, and adjusting the weights in the weight set based on the network status, wherein the network status includes at least one of the following: encountering a security incident, business load conditions.

[0163] In some embodiments of the present application, the third processing module 64 determines the repair method corresponding to the fault, including: determining the repair method corresponding to the fault through preset expert rules; or determining the repair method corresponding to the fault through a deep learning network, wherein the deep learning network is used to determine the repair instructions corresponding to the fault in the action space based on the state space for indicating the repair method, the state space includes the network topology state of the system, the equipment load condition and the fault type of the fault, and the action space includes a set of executable alternative repair instructions.

[0164] In some embodiments of the present application, the fourth processing module 66 verifies the repair method through the system simulation model corresponding to the system, including: performing routing loop detection on the repair method through the system simulation model to determine whether there is a loop path in the system after the repair method is executed; performing bandwidth congestion detection on the repair method through the system simulation model to determine whether there is a link whose utilization exceeds a preset threshold after the repair method is executed; if the detection result is that there is no loop path and no link whose utilization exceeds the preset threshold, it is determined that the repair method has passed the verification.

[0165] An embodiment of the present application provides a non-volatile storage medium, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the following network fault repair method: collecting status data of each network device in the system, and the topological connection relationship between each network device, and updating the network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; determining the abnormal probability prediction value of each node in the network topology map, and the health status score of each edge; determining whether a fault occurs based on the abnormal probability prediction value and the health status score, and after determining that a fault occurs, determining the repair method corresponding to the fault; verifying the repair method through a system simulation model corresponding to the system, and using the repair method to repair the fault after the verification is passed.

[0166] An embodiment of the present application provides an electronic device, comprising: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the following network fault repair method when running: collecting status data of each network device in the system, and the topological connection relationship between each network device, and updating a network topology map based on the status data and the topological connection relationship, wherein nodes in the network topology map correspond to network devices, and edges correspond to connection links between network devices; determining an abnormality probability prediction value for each node in the network topology map, and a health status score for each edge; determining whether a fault has occurred based on the abnormality probability prediction value and the health status score, and determining a repair method corresponding to the fault after determining that a fault has occurred; verifying the repair method through a system simulation model corresponding to the system, and using the repair method to repair the fault after the verification passes.

[0167] An embodiment of the present application provides a computer program product, including a computer program, which implements the following network fault repair method when executed by a processor: collecting status data of each network device in the system, as well as the topological connection relationship between each network device, and updating a network topology map based on the status data and the topological connection relationship, wherein the nodes in the network topology map correspond to the network devices, and the edges correspond to the connection links between the network devices; determining an abnormality probability prediction value for each node in the network topology map, as well as a health status score for each edge; determining whether a fault has occurred based on the abnormality probability prediction value and the health status score, and after determining that a fault has occurred, determining a repair method corresponding to the fault; verifying the repair method through a system simulation model corresponding to the system, and using the repair method to repair the fault after the verification is passed.

[0168] It should be noted that the various modules in the above-mentioned network fault repair device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0169] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0170] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0171] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0172] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0174] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A network fault repair method, characterized in that: include: Collecting status data of each network device in the system and the topological connection relationship between the network devices, and updating a network topology map based on the status data and the topological connection relationship, wherein nodes in the network topology map correspond to the network devices and edges correspond to connection links between the network devices; Determining an abnormality probability prediction value of each node in the network topology graph and a health status score of each edge; Determining whether a fault has occurred based on the abnormal probability prediction value and the health status score, and after determining that a fault has occurred, determining a repair method corresponding to the fault; The repair method is verified by a system simulation model corresponding to the system, and the fault is repaired by using the repair method after passing the verification.

2. The network fault repair method according to claim 1, characterized in that: Determining the abnormal probability prediction value of each node in the network topology graph and the health status score of each edge includes: Determining local characteristics of each node and global characteristics corresponding to the system based on the status data, wherein the local characteristics include device status corresponding to adjacent nodes of the node, and the global characteristics include global load distribution status of the system; Determining a local feature weight corresponding to the local feature and a global feature weight corresponding to the global feature; The abnormality probability prediction value of each node and the health status score of each edge are determined based on the local features and the local feature weights, and the global features and the global feature weights.

3. The network fault repair method according to claim 1, characterized in that: Determining whether a fault occurs based on the abnormal probability prediction value and the health status score includes: Determining historical fault data, wherein the historical fault data includes historical abnormality probability prediction values and historical health status scores; Determining a first distribution of the historical abnormal probability prediction values and a second distribution of the historical health status scores; determining a first threshold value according to the first distribution, and determining a second threshold value according to the second distribution; Compare the abnormal probability prediction value of each node with the first threshold, and determine whether there is a faulty node based on the comparison result; and compare the health status score of each edge with the second threshold, and determine whether there is a faulty edge based on the comparison result.

4. The network fault repair method according to claim 1, characterized in that: Determining the repair method corresponding to the fault includes: Determining the fault type of the fault, as well as the status data and link status of each network device in the system, and determining a general fault repair instruction corresponding to the fault; Determining an intermediate representation instruction corresponding to the general fault repair instruction, wherein the intermediate representation instruction is a fault repair instruction that is unrelated to the device type, and the intermediate representation instruction includes a fault repair parameter for indicating the fault repair method; In the case where the fault type is a device fault, determining the device type of the faulty device and determining an instruction template corresponding to the device type; A device fault repair instruction is determined according to the intermediate representation instruction and the instruction template, wherein the fault repair instruction is used to indicate the fault repair method.

5. The network fault repair method according to claim 1, characterized in that: Determining the repair method corresponding to the fault includes: Determining multiple alternative repair methods corresponding to the fault and a set of evaluation values for the alternative repair methods, wherein the set of evaluation values includes a safety evaluation value, a performance evaluation value, and a cost evaluation value; Determining a weight set for the system, wherein the weight set includes a security weight, a performance weight, and a cost weight; A target repair method corresponding to the fault is determined from the multiple alternative repair methods based on the weight set and the evaluation value set corresponding to each of the alternative repair methods.

6. The network fault repair method according to claim 5, characterized in that: Determining the set of weights for the system includes: Determine the network status of the system, and adjust the weights in the weight set according to the network status, wherein the network status includes at least one of the following: encountering a security incident, and business load conditions.

7. The network fault repair method according to claim 1, characterized in that: Determining the repair method corresponding to the fault includes: Determine the repair method corresponding to the fault through preset expert rules; or A repair method corresponding to the fault is determined through a deep learning network, wherein the deep learning network is used to determine a repair instruction corresponding to the fault and used to indicate the repair method in an action space based on a state space, the state space includes the network topology state of the system, the equipment load condition and the fault type of the fault, and the action space includes a set of executable alternative repair instructions.

8. The network fault repair method according to claim 1, characterized in that: Verifying the repair method by using a system simulation model corresponding to the system includes: Performing routing loop detection on the repair method through the system simulation model to determine whether there is a loop path in the system after executing the repair method; Performing bandwidth congestion detection on the repair method through the system simulation model to determine whether there is a link whose utilization exceeds a preset threshold after executing the repair method; If the detection result shows that the loop path does not exist and there is no link whose utilization exceeds a preset threshold, it is determined that the repair method has passed the verification.

9. A network fault repair device, characterized in that: include: a first processing module, configured to collect status data of each network device in the system and a topological connection relationship between the network devices, and update a network topology map based on the status data and the topological connection relationship, wherein nodes in the network topology map correspond to the network devices and edges correspond to connection links between the network devices; A second processing module is used to determine an abnormality probability prediction value of each node in the network topology graph and a health status score of each edge; a third processing module, configured to determine whether a fault occurs based on the abnormal probability prediction value and the health status score, and, if a fault occurs, determine a repair method corresponding to the fault; The fourth processing module is used to verify the repair method through a system simulation model corresponding to the system, and adopt the repair method to repair the fault after the verification is passed.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the network fault repair method according to any one of claims 1 to 8.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the network fault repair method according to any one of claims 1 to 8 is executed when the program is run.

12. A computer program product, characterized in that The method comprises a computer program, which implements the network fault repair method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Cited By

  • Graphics processor risk processing method and device, equipment and medium

    CN121256232A

  • Self-checking and self-repairing method and device of electronic equipment, equipment and storage medium

    CN121277743A

  • Self-checking and self-repairing method and device of electronic equipment, equipment and storage medium

    CN121277743B

  • Fault repairing method and system for GPU (Graphic Processing Unit)

    CN122132235A