Intelligent fault detection method and system for network switch

By building a risk prediction model and analyzing the network topology, the problem of difficult to predict and prevent network switch failures in the existing technology is solved, and more accurate and reliable fault prediction and prevention are achieved, reducing enterprise operation risks.

CN120075078AInactive Publication Date: 2025-05-30SHENZHEN MAXTOPIC TECH CO LTD

Patent Information

Application Number
CN202510551843.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively predict and prevent network switch failures, resulting in network service downtime and data loss, increasing the complexity and cost of enterprise operations.

Method used

By obtaining the historical operation characteristic information of the target network device, analyzing the current operating data and risk trends, determining the failure rate, and determining the fault propagation path based on the network topology, thereby predicting the failure rate of the associated network device.

Benefits of technology

The transition from passive detection to active prevention has been achieved, which improves the accuracy and reliability of network switch failure prediction, and reduces the risks and losses caused by potential failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075078A_ABST
    Figure CN120075078A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent fault detection method and system for a network switch, relates to the technical field of computers, and aims to realize conversion from passive detection to active prevention by quantifying fault probability and influence, analyzing risk trend and relevance and visually displaying risk distribution. The method comprises the following steps: acquiring historical operation feature information of target network equipment, and constructing a risk prediction model according to the historical operation feature information; determining the failure rate of the target network equipment based on the current operation data of the target network equipment and a risk prediction model; when the fault rate of the target network equipment is greater than or equal to a preset fault rate, obtaining a network topology corresponding to the target network equipment, and determining a fault propagation path according to the network topology; and according to the fault propagation path and the fault rate of the target network device, determining the fault rate of the associated network device on the fault propagation path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an intelligent fault detection method and system for network switches. Background Art

[0002] With the rapid development of cloud computing and Internet services, the number of network devices has increased sharply, including various switches such as Top-of-Rack (TOP) switches and aggregation switches. Over time, the problem of network device aging has become increasingly prominent, and the probability of device failure will also increase rapidly. Especially for devices that have been in use for more than 3 years, the failure rate will show a sharp increase. The originally occasional network failures have gradually become normal. Among the total network device failures, switch failures dominate and have the most serious impact, causing catastrophic effects such as service outages and even data loss, greatly increasing the complexity and cost of enterprise operations. Therefore, predicting switch failures and improving the reliability of the data center network have become the key to ensuring the high efficiency and stability of the system.

[0003] Current data center network fault tolerance solutions mainly focus on changing protocols and network topologies. In this way, the data center network can automatically recover from network failures. However, the above methods do not cover all switches, and sometimes it is necessary for operation and maintenance personnel to quickly diagnose and locate. These methods either face deployment problems or require a large amount of time to locate and solve switch failures. Therefore, it is necessary to predict switch failures and let operation and maintenance personnel intervene and solve them before potential failures occur. Summary of the Invention

[0004] Based on the above technical problems, this application provides an intelligent fault detection method and system for network switches, which realizes the transformation from passive detection to active prevention by quantifying fault probabilities and impacts, analyzing risk trends and correlations, and intuitively displaying risk distributions.

[0005] In a first aspect, the present application provides an intelligent fault detection method for a network switch. The method includes: obtaining historical operation characteristic information of a target network device and constructing a risk prediction model based on the historical operation characteristic information; the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation state of the target network device; determining the failure rate of the target network device based on the current operation data of the target network device and the risk prediction model; in the case where the failure rate of the target network device is greater than or equal to a preset failure rate, obtaining the network topology corresponding to the target network device and determining the fault propagation path according to the network topology; determining the failure rates of the associated network devices on the fault propagation path according to the fault propagation path and the failure rate of the target network device; the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

[0006] In a possible implementation, the target network device is the central device in the network topology, or the amount of historical operation data of the target network device is less than or equal to a preset threshold.

[0007] In a possible implementation, determining the fault propagation path according to the network topology includes: processing the network topology by using a trained graph neural network to obtain the fault propagation path; the graph neural network reflects the fault relationship between network devices in the network topology through the edges between nodes.

[0008] In a possible implementation, determining the failure rates of the associated network devices on the fault propagation path according to the fault propagation path and the failure rate of the target network device includes: for any associated network device on the fault propagation path, obtaining the fault relationship between the associated network device and the target network device; the fault relationship includes the influence degree of the target network device on the associated network device; determining the failure rate of the associated network device on the fault propagation path based on the influence degree of the target network device on the associated network device and the failure rate of the target network device.

[0009] In a possible implementation, the method further includes: for any associated network device on the fault propagation path, in the case where the influence degree of the target network device on the associated network device is greater than or equal to a preset influence degree, updating the network topology according to the data flow direction between the target network device and the associated network device to adjust the device distance between the target network device and the associated network device, so as to obtain an updated network topology; the influence degree of any associated network device on the fault propagation path in the updated network topology affected by the target network device is less than or equal to the preset influence degree.

[0010] In a possible implementation, the method further includes: clustering associated network devices according to failure rates to obtain a clustering result; dividing risk regions in the network topology based on the clustering result to obtain multiple risk regions; network devices in one risk region correspond to the same cluster, and different risk regions correspond to different risk levels; determining risk boundaries in the network topology according to the multiple risk regions, and visually displaying the risk boundaries and the multiple risk regions.

[0011] In a possible implementation, the method further includes: sorting the multiple risk regions in descending order of risk level to obtain a sorting result, and determining a risk prevention strategy for the network topology according to the sorting result.

[0012] The technical solution provided by this application at least brings the following beneficial effects: (1) This application combines historical data and historical operating status, and uses machine learning algorithms to construct a risk trend prediction model applicable to target network devices. Furthermore, the failure rate of the target network device can be determined based on the risk prediction model, so as to achieve the effect of predicting the failure of the target network device. Further, if the failure rate of the target network device is greater than or equal to a preset failure rate, this application will also obtain the network topology corresponding to the target network device, and determine the fault propagation path according to the network topology. Furthermore, according to the fault propagation path and the failure rate of the target network device, the failure rates of associated network devices on the fault propagation path are determined, so as to achieve the effect of predicting the failures of each network device in the network topology.

[0013] (2) This application can select a key node in a certain network topology as the target network device, such as the central device in the network topology. Since the central device has a greater impact on other network devices in the network topology, selecting the central device as the target network device for fault prediction can strengthen the security guarantee for other network devices and the entire network topology. In addition, the embodiments of this application can also select nodes with small data volume or obvious features as the target network devices. For example, network devices with historical operation data volume less than or equal to a certain preset threshold can be selected as the target network devices. Since the data volume of these network devices is small, it is more computationally efficient and easier to analyze their data characteristics during data processing, and thus a corresponding risk prediction model can be constructed quickly.

[0014] (3) Graph neural network is a type of deep learning model specifically used to process graph-structured data. Compared with traditional neural networks, graph neural network has significant advantages in processing non-Euclidean space data, and thus is more suitable for processing network topologies.

[0015] After determining the failure rate of the target network device, the present application further determines the failure rate of the associated network devices on the fault propagation path by using the influence degree of the target network device on the associated network devices, which can improve the prediction ability of the faults in the entire network topology.

[0016] (5)For any associated network device on the fault propagation path, when the influence degree of the target network device on the associated network device is greater than or equal to the preset influence degree, the present application can also update the network topology according to the data flow direction between the target network device and the associated network device, so as to adjust the device spacing between the target network device and the associated network device, and obtain the updated network topology, so that the influence degree of any associated network device on the fault propagation path in the updated network topology by the target network device is less than or equal to the preset influence degree, thereby achieving the purpose of optimizing the network topology.

[0017] (6)The present application can also divide risk areas in the network topology and visually display the boundaries of the divided areas, thereby facilitating the user to intuitively understand each risk area.

[0018] (7)For risk areas with different risk levels, the present application can formulate different risk prevention strategies respectively, thereby more specifically facing various risk situations.

[0019] In a second aspect, the present application provides a network device fault prediction device, which includes an acquisition unit and a determination unit; the acquisition unit is used to acquire the historical operation characteristic information of the target network device and construct a risk prediction model according to the historical operation characteristic information; the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation state of the target network device; the determination unit is used to determine the failure rate of the target network device based on the current operation data of the target network device and the risk prediction model; the determination unit is also used to, when the failure rate of the target network device is greater than or equal to the preset failure rate, acquire the network topology corresponding to the target network device and determine the fault propagation path according to the network topology; the determination unit is also used to determine the failure rate of the associated network devices on the fault propagation path according to the fault propagation path and the failure rate of the target network device; the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

[0020] In a possible implementation manner, the target network device is the central device in the network topology, or the data volume of the historical operation data of the target network device is less than or equal to the preset threshold.

[0021] In one possible implementation, the determining unit is specifically configured to: process the network topology using a trained graph neural network to obtain a fault propagation path; the graph neural network reflects the fault relationship between network devices in the network topology through the edges between nodes.

[0022] In one possible implementation, the determining unit is specifically configured to: for any associated network device on the fault propagation path, obtain the fault relationship between the associated network device and the target network device; the fault relationship includes the degree of influence of the target network device on the associated network device; based on the degree of influence of the target network device on the associated network device and the failure rate of the target network device, determine the failure rate of the associated network device on the fault propagation path.

[0023] In one possible implementation, the network device fault prediction apparatus further includes a processing unit, and the processing unit is configured to: for any associated network device on the fault propagation path, when the degree of influence of the target network device on the associated network device is greater than or equal to a preset influence degree, update the network topology according to the data flow direction between the target network device and the associated network device to adjust the device distance between the target network device and the associated network device, and obtain an updated network topology; the degree of influence of any associated network device on the fault propagation path in the updated network topology by the target network device is less than or equal to the preset influence degree.

[0024] In one possible implementation, the processing unit is further configured to: cluster the associated network devices according to the failure rate to obtain a clustering result; divide risk regions in the network topology based on the clustering result to obtain multiple risk regions; the network devices in one risk region correspond to the same cluster, and different risk regions correspond to different risk levels; according to the multiple risk regions, determine the risk boundary in the network topology and visually display the risk boundary and the multiple risk regions.

[0025] In one possible implementation, the determining unit is further configured to: sort the multiple risk regions in descending order of risk level to obtain a sorting result, and determine a risk prevention strategy for the network topology according to the sorting result.

[0026] In a third aspect, the present application provides an electronic device, including: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device implements the method of the first aspect as described above.

[0027] In a fourth aspect, the present application provides a computer program product, when the computer program product runs in an electronic device, enabling the electronic device to execute the method related to the first aspect as described above to implement the method of the first aspect.

[0028] Fifth aspect, the present application provides a computer-readable storage medium, which includes: software instructions; when the software instructions run on an electronic device, the electronic device is enabled to implement the method of the first aspect above.

[0029] The beneficial effects of the second to fifth aspects above can be referred to the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 It is a schematic structural diagram of a network device fault prediction system provided by an embodiment of the present application; Figure 2 It is a schematic composition diagram of an electronic device provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of an intelligent fault detection method for a network switch provided by an embodiment of the present application; Figure 4 It is a schematic composition diagram of a network device fault prediction device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0034] In addition, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. "And / or" herein is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0035] Before explaining the embodiments of the present application in detail, some related terms and related technologies involved in the embodiments of the present application will be introduced first.

[0036] As a core component of modern communication systems, the stable operation of network switches is crucial for ensuring the reliability and efficiency of data transmission. With the expansion of network scale and the increase in complexity, fault detection and risk management have become key links to ensure system performance. Traditional methods usually rely on manual experience or simple threshold monitoring. Although effective in some scenarios, they tend to be inadequate in the face of a dynamically changing network environment and potential complex faults. These methods mainly focus on the detection and repair of occurred faults and lack the ability to predict potential risks, resulting in the system being prone to unforeseen downtime under high load or abnormal conditions.

[0037] The limitations of existing solutions are that they can usually only respond passively to exposed problems and cannot actively identify and manage hidden dangers that have not yet emerged. Especially with the rapid increase in device interconnectivity and data traffic in the network, a single fault may trigger a chain reaction, and existing methods have obvious deficiencies in quantifying the impact of risks, predicting development trends, and revealing the correlations between multiple risks. This shortcoming of passive detection makes it difficult for network managers to take effective preventive measures before faults occur.

[0038] The core challenges in the research field focus on how to accurately quantify the occurrence probability and impact degree of different fault scenarios, and how to extract the development trends and correlation features of risks from complex network data. First, risk assessment requires establishing a scientific model to cope with the changing operating conditions and technical parameters in the network; second, the prediction of risk trends depends on in-depth analysis of historical data and real-time status, which poses higher requirements for the computing power and algorithm design of the system; finally, the correlation analysis of multiple risks involves complex mappings of network topologies and fault propagation paths, and existing technologies still struggle to handle these dynamic relationships. These unsolved technical problems have led to the identification and prevention of potential risks becoming a weak link in network optimization.

[0039] In view of the above problems, an embodiment of the present application provides an intelligent fault detection method for a network switch, aiming to construct a comprehensive risk assessment and prediction system, by quantifying the fault probability and impact, analyzing the risk trend and correlation, and intuitively displaying the risk distribution, so as to realize the transformation from passive detection to active prevention, and further improve the intelligent fault detection ability of the network switch.

[0040] The following will describe in detail the intelligent fault detection method for a network switch provided by the embodiment of the present application with reference to the accompanying drawings.

[0041] The intelligent fault detection method for a network switch provided by the embodiment of the present application can be applied to a network device fault prediction system. Figure 1 A schematic structural diagram of the network device fault prediction system is shown. As Figure 1 shown, the network device fault prediction system 10 includes a network device fault prediction device 11 and a plurality of network devices 12. Among them, the network devices 12 are connected in a wired or wireless manner to form a network topology; the network device fault prediction device 11 is connected to the network topology formed by the plurality of network devices 12 in a wired or wireless manner. Specifically, the network device fault prediction device 11 can be connected to one or more network devices in the network topology in a wired or wireless manner.

[0042] In the embodiment of the present application, the network device connected to the network device fault prediction device 11 can also be referred to as a target network device.

[0043] In the embodiment of the present application, the network device 12 can be, for example, a switch, an access router, a backbone router, etc., but is not limited thereto.

[0044] The network device fault prediction device 11 can be any electronic device with data processing functions. For example, the network device fault prediction device 11 can be a server, a computer, or can also be a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. Optionally, the server can be a central server, and the server can also be implemented on a cloud platform. For example, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, and a multi-cloud, etc., or any combination thereof. The embodiment of the present application does not limit this.

[0045] The execution subject of the intelligent fault detection method for a network switch provided by an embodiment of this application may be the aforementioned network device fault prediction apparatus 11. As described above, the network device fault prediction apparatus 11 may be an electronic device with data processing capabilities such as a computer or a server. Optionally, the network device fault prediction apparatus 11 may also be a processor (such as a central processing unit (CPU)) in the aforementioned electronic device; or, the network device fault prediction apparatus 11 may also be an application (APP) installed in the aforementioned electronic device with model training capabilities; or, the network device fault prediction apparatus 11 may also be a functional module with model training capabilities in the aforementioned electronic device, etc. The embodiments of this application do not limit this.

[0046] For simplicity of description, the following will uniformly introduce the network device fault prediction apparatus 11 as an electronic device as an example.

[0047] Figure 2 It is a schematic diagram of the composition of the electronic device provided by an embodiment of this application. As Figure 2 shown, the electronic device may include: a processor 20, a memory 21, a communication line 22, a communication interface 23, and an input / output interface 24.

[0048] Among them, the processor 20, the memory 21, the communication interface 23, and the input / output interface 24 may be connected through the communication line 22.

[0049] The processor 20 is configured to execute instructions stored in the memory 21 to implement the fault analysis method provided by the following embodiments of this application. The processor 20 may be a CPU, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller (MCU), a programmable logic device (PLD), or any combination thereof. The processor 20 may also be any other device with processing capabilities, such as a circuit, a device, or a software module. The embodiments of this application do not limit this. In one example, the processor 20 may include one or more CPUs, such as Figure 2 the CPU0 and CPU1 in Figure 2 shown as a dotted line). As an optional implementation manner, the electronic device may include multiple processors. For example, in addition to the processor 20, it may also include a processor 25 (

[0050] A memory 21 for storing instructions. For example, the instructions can be a computer program. Optionally, the memory 21 can be a read-only memory (ROM) or other types of static storage devices that can store static information and / or instructions, or it can be a random access memory (RAM) or other types of dynamic storage devices that can store information and / or instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc. The embodiments of the present application do not limit this.

[0051] It should be noted that the memory 21 can exist independently of the processor 20 or can be integrated with the processor 20. The memory 21 can be located inside the electronic device or outside the electronic device. The embodiments of the present application do not limit this.

[0052] A communication line 22 for transmitting information between the various components included in the electronic device.

[0053] A communication interface 23 for communicating with other devices or other communication networks. The other communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 23 can be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0054] An input / output interface 24 for implementing the human-computer interaction between the user and the electronic device. For example, it realizes the action interaction or information interaction between the user and the electronic device.

[0055] Exemplarily, the input / output interface 24 can be a mouse, a keyboard, a display screen, or a touch display screen, etc. Through a mouse, a keyboard, a display screen, or a touch display screen, etc., the action interaction or information interaction between the user and the electronic device can be realized.

[0056] It should be noted that Figure 2 the structure shown in Figure 2 does not constitute a limitation on the electronic device. In addition to

[0057] The following introduces the intelligent fault detection method for network switches provided in the embodiments of this application.

[0058] Figure 3 It is a schematic flowchart of the intelligent fault detection method for network switches provided in the embodiments of this application. Optionally, this method can be executed by an electronic device with the above Figure 2 shown hardware structure, such as Figure 3 shown, this method includes S301 to S304.

[0059] S301. Obtain the historical operation characteristic information of the target network device, and construct a risk prediction model according to the historical operation characteristic information.

[0060] Among them, the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation state of the target network device.

[0061] It should be noted that the selection of the target network device can be arbitrary. In the embodiments of this application, a key node in a certain network topology can be selected as the target network device, such as the central device in the network topology. Since the central device has a greater impact on other network devices in the network topology, selecting the central device as the target network device for fault prediction can strengthen the security guarantee for other network devices and the entire network topology. In addition, in the embodiments of this application, a node with a small amount of data or obvious features can also be selected as the target network device. For example, a network device with a historical operation data volume less than or equal to a certain preset threshold can be selected as the target network device. Since the data volume of these network devices is small, it is more computationally efficient and easier to analyze their data characteristics during data processing, and thus a corresponding risk prediction model can be constructed quickly.

[0062] For the acquisition of the historical operation characteristic information of the target network device, in the embodiments of this application, device logs and traffic records can be obtained from the target network device to obtain a preliminary set of historical data and historical operation states. Further, a time series analysis method is used to process the historical data to extract the change trend of the failure probability and obtain an initial description of the distribution characteristics.

[0063] Among them, obtaining device logs and traffic records from the operation of the switch through a distributed system is the basis for realizing network status monitoring.

[0064] For example, in an enterprise network, the distributed system can be deployed on multiple switch nodes to collect log files and traffic data in real time, such as the number of packets transmitted per second or the packet loss rate, and initially form a set of historical data and real-time status.

[0065] For example, historical data may include traffic peak records for the past 30 days, while real-time status reflects traffic fluctuations within the current hour. This approach can efficiently aggregate large-scale network data and provide support for subsequent analysis. Using time series analysis methods to process historical data can reveal the changing trend of failure probability.

[0066] Specifically, by analyzing the daily traffic peaks and downtime records of the switches, the correlation between traffic surges and failures can be extracted.

[0067] For example, if the downtime probability of a switch increases from 5% to 20% when the traffic exceeds 1000Mbps, the distribution characteristics of the failure probability can be preliminarily described. This trend analysis helps predict potential risk points and improve the pertinence of network maintenance. Judging abnormal fluctuations in real-time traffic records is a key step.

[0068] In one possible implementation, the preset threshold is 1.5 times the normal traffic. For example, if the normal traffic is 500Mbps, the threshold is set to 750Mbps. If the traffic suddenly increases to 800Mbps at a certain moment, it indicates an abnormality and the probability of potential failure increases. This real-time monitoring can quickly respond to network abnormalities and reduce the possibility of fault propagation. Based on the correlation between distribution characteristics and failure probability, it is particularly important to obtain key nodes in topology information.

[0069] For example, in a network containing 10 switches, analysis found that the core switch carries 80% of the traffic. If its failure probability rises to 30%, the risk is much higher than that of the edge node.

[0070] Preferably, this risk quantification distribution can help operation and maintenance personnel prioritize high-risk nodes and improve resource allocation efficiency. Verifying the integrity of data collection through a distributed system can ensure the reliability of analysis.

[0071] It should be noted that if the topology information shows that a switch should have 3 connection ports, but the device log only records the traffic of 2 ports, the data needs to be retrieved again. This consistency check can avoid risk misjudgment caused by missing data and ensure the accuracy of the results. Using the random forest algorithm to integrate historical data and real-time status change trends is an effective way to improve the dynamic nature of risk quantification.

[0072] In one embodiment, the algorithm can combine the downtime frequency in the past week and the current traffic fluctuations to predict the probability of failure in the next hour.

[0073] For example, historical data shows that traffic peaks on Mondays are likely to trigger failures. Combining with the current abnormal traffic, the algorithm may adjust the failure probability from 10% to 25%. This dynamic adjustment can more accurately reflect the network status. Analyzing the propagation paths of potential failures based on key nodes and the updated failure probability distribution is the ultimate goal.

[0074] It can be understood that if the failure probability of the core switch is 30%, the three downstream edge switches may be affected due to traffic overload. Through topology analysis, the path of failure spreading from the core node to the edge node can be predicted, and the overall risk can be quantified.

[0075] For example, a core node failure may cause 50% of the downstream services to be interrupted. This path analysis can provide a basis for emergency response plans and significantly reduce losses.

[0076] Therefore, the embodiments of the present application can use historical operation data as input and historical operation status as labels, and construct a risk prediction model (also called a risk trend prediction model) through machine learning algorithms, and then determine the change trend of risk over time based on this model.

[0077] In some embodiments, the present application can also extract key metrics through a set of operation modes, and combine with the risk quantification results to judge the initial distribution of abnormal signals.

[0078] Specifically, the electronic device can obtain the long-term operation mode from the historical data of the switch.

[0079] It can be understood that usually, based on the device logs and traffic records accumulated during the operation of the switch, the regularity in the time dimension is analyzed.

[0080] For example, by counting the time periods when traffic peaks and troughs occurred in the past year, a preliminary set of operation modes is formed.

[0081] Exemplarily, assume that the traffic of a certain switch surges from 8 am to 9 am every day and gradually stabilizes after 10 pm. This periodic feature can be summarized as a mode. When using statistical methods to analyze data features.

[0082] Preferably, the mean and variance of the traffic data can be calculated, and characteristic values such as the daily average traffic of 500 Mbps and the peak fluctuation range of 200 - 800 Mbps can be extracted to obtain the set of operation modes. When extracting key metrics through the set of operation modes.

[0083] Specifically, metrics such as the duration of traffic peaks and the frequency of abnormal disconnections can be concerned. Combining with the risk quantification results, for example, the historical failure probability is 5%, to judge the initial distribution of abnormal signals.

[0084] In a possible implementation, if the traffic peak continues to exceed 30 minutes and the disconnection count increases within a certain period of time, it may indicate that the probability distribution of abnormal signals has increased from 5% to 10%. The beneficial effect of this is that it can quickly locate the starting point of potential problems. When obtaining real-time status data and analyzing traffic fluctuation characteristics.

[0085] For example, the current traffic suddenly surges from 300 Mbps to 700 Mbps, and the fluctuation amplitude exceeds the normal range by 50%.

[0086] It should be noted that by comparing with historical patterns, the correlation between fluctuations and abnormal signals is determined. If historical data shows that the failure probability increases to 15% after similar fluctuations, then the risk trend of the current state can be initially judged. Such analysis can effectively improve the accuracy of real-time monitoring. When constructing a feature set using the random forest algorithm.

[0087] In one embodiment, the 10% probability of the initial distribution of abnormal signals, the 50% fluctuation amplitude, and the historical disconnection frequency, etc. can be used as input features to generate a basic model for risk trend prediction. When processing time series data, the basic model may show that the risk probability increases from 10% to 20% within the next 2 hours. If the change trend exceeds the preset threshold of 15%, the model parameters are adjusted by the support vector machine.

[0088] Exemplarily, the support vector machine will re-divide the feature boundary to make the prediction result closer to the actual risk. For example, the probability stabilizes at 18% after adjustment. This optimization can enhance the robustness of the prediction. When analyzing the long-term risk distribution based on the optimized prediction result.

[0089] It can be understood that by combining historical patterns and real-time fluctuations, the distribution curve of the risk probability within the next week is deduced.

[0090] For example, the prediction shows that the risk reaches 20% on Monday morning and remains below 10% at other times.

[0091] Specifically, by analyzing the load of key nodes through topological information, if a certain node has long-term traffic overload, it may become the starting point of fault propagation. This analysis method helps to plan maintenance strategies in advance and reduce the possibility of system interruption.

[0092] In one embodiment, for the traffic overload situation of key nodes, the impact on neighboring nodes can be further deduced by combining historical data.

[0093] For example, if the traffic of the core switch exceeds 80% of its capacity, the neighboring nodes may have a 30% increase in the failure probability within 2 hours.

[0094] Preferably, by adjusting traffic distribution or adding redundant paths, the risk propagation can be effectively alleviated. This multi-faceted supported analysis not only improves the comprehensiveness of risk prediction but also provides diverse coping solutions for actual operation and maintenance.

[0095] S302. Determine the failure rate of the target network device based on the current operating data of the target network device and the risk prediction model.

[0096] As a possible implementation, the electronic device can input the current operating data of the target network device into the trained risk prediction model to obtain the failure rate of the target network device.

[0097] S303. When the failure rate of the target network device is greater than or equal to the preset failure rate, obtain the network topology corresponding to the target network device, and determine the fault propagation path according to the network topology.

[0098] Among them, the preset failure rate can be set in advance by the operation and maintenance personnel, and the specific data value of the preset failure rate is not limited in the embodiments of the present application. For example, the preset failure rate can be 50%. If the failure rate of the target network device is greater than or equal to 50%, it indicates that the probability of the target network device failing is relatively high. Therefore, it is necessary to further perform fault troubleshooting in combination with the network topology of the target network device.

[0099] As a possible implementation, the electronic device can use the trained graph neural network to process the network topology to obtain the fault propagation path; among them, the graph neural network reflects the fault relationship between network devices in the network topology through the edges between nodes.

[0100] It can be understood that in actual applications, the actual impact of each node on the data flow in the topology is usually analyzed based on the connection structure of switches and routers and traffic logs.

[0101] For example, a network includes a core switch A and edge switches B and C. Historical data shows that the traffic of A accounts for 70% of the total, while B and C share 20% and 10% respectively. This distribution reflects the hierarchical constraints of the topology.

[0102] Exemplarily, if the bandwidth capacity of A is 1000 Mbps, and those of B and C are 500 Mbps each, the mapping relationship indicates that A is a traffic bottleneck point. When extracting data traffic characteristics from real-time data.

[0103] Specifically, it can be observed whether the current traffic deviates from the normal range.

[0104] In a possible implementation, assume that the normal traffic mean is 400 Mbps and the preset threshold is 600 Mbps. If the real-time traffic suddenly increases to 700 Mbps, it exceeds the threshold by 100 Mbps, triggering an anomaly judgment.

[0105] It should be noted that this feature extraction needs to be combined with a time window, such as the average value within 5 minutes, to avoid misjudgment due to instantaneous fluctuations. If the traffic surges beyond the preset threshold, when using a graph neural network to analyze the network topology.

[0106] Preferably, the connection weights and traffic loads between nodes will be used as inputs.

[0107] For example, if the core switch A is connected to B and C, and the traffic of A is overloaded, the graph neural network can calculate the possibility of the fault spreading from A to B and C.

[0108] In one embodiment, after the analysis shows that A is overloaded, the failure probability of B increases to 25%, which helps to locate the propagation path. This kind of analysis can quickly identify the vulnerability of key nodes. When extracting the propagation range feature according to the fault propagation path.

[0109] For example, if the overload of A affects B, the propagation range may cover the downstream node D of B. The initial value of the influence degree can be set as the proportion of the traffic of the affected node. Assume that D accounts for 15% of the total traffic, then the initial value is 15%.

[0110] Specifically, this feature extraction can also reflect the availability of redundant paths in the topology, improving the comprehensiveness of the evaluation. When processing the initial value of the influence degree through time series data.

[0111] In a possible implementation, a sliding window can be used to analyze the change trend of 15%.

[0112] For example, if the initial value has increased from 15% to 20% in the past 1 hour, it indicates that the dynamic evaluation value shows an upward trend. This trend analysis can capture the evolution process of risks. When obtaining the historical distribution of the dynamic evaluation value.

[0113] It can be understood that by comparing the historical anomaly fluctuation records, the correlation with the fault propagation is judged.

[0114] Exemplarily, if the historical data shows that in 80% of the cases, node failures occur after the traffic surges, then the current 20% dynamic evaluation value is highly correlated with the failure. This correlation analysis can optimize the accuracy of subsequent decisions. When adjusting the mapping relationship parameters according to the correlation.

[0115] Preferably, the weight coefficient of A can be increased, from 70% to 75%, to reflect its greater constraint on the traffic.

[0116] For example, the adjusted mapping result shows that the traffic sharing of B has decreased from 20% to 18%. Such optimization can more realistically match the real-time status and topology, effectively improving the adaptability of traffic management.

[0117] In another possible way, the electronic device can determine the directly connected network devices and non-directly connected network devices of the target network device in the network topology, determine the path between the target network device and the directly connected network device as the first fault propagation path, and determine the path between the target network device and the non-directly connected network device as the second fault propagation path. Further, the electronic device can obtain the fault propagation path based on the first fault propagation path and the second fault propagation path.

[0118] For example, the electronic device can regard both the first fault propagation path and the second fault propagation path as the fault propagation path. Or, in order to save computing resources, the electronic device can also regard only the first fault propagation path as the fault propagation path and temporarily not regard the second fault propagation path as the fault propagation path.

[0119] S304. Determine the failure rate of the associated network devices on the fault propagation path according to the fault propagation path and the failure rate of the target network device.

[0120] Among them, the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

[0121] As a possible implementation, for any associated network device on the fault propagation path, the electronic device can obtain the fault relationship between the associated network device and the target network device; among them, the fault relationship includes the influence degree of the target network device on the associated network device. Further, the electronic device can determine the failure rate of the associated network devices on the fault propagation path based on the influence degree of the target network device on the associated network device and the failure rate of the target network device.

[0122] In a design, for any associated network device on the fault propagation path, when the influence degree of the target network device on the associated network device is greater than or equal to the preset influence degree, the electronic device can update the network topology according to the data flow direction between the target network device and the associated network device to adjust the device spacing between the target network device and the associated network device, and obtain the updated network topology; among them, the influence degree of any associated network device on the fault propagation path in the updated network topology by the target network device is less than or equal to the preset influence degree.

[0123] In addition, the electronic device can obtain the interaction data between multiple devices through the real-time status.

[0124] It can be understood that network devices usually rely on communication logs and status update records between devices.

[0125] For example, in an enterprise network that includes server S1, switch S2, and terminal device S3, the interaction data may show that the packet rate sent from S1 to S2 suddenly increases, and the forwarding delay from S2 to S3 becomes longer.

[0126] In one possible implementation, the association rule mining technology analyzes this data and discovers the association feature of "when the traffic of S1 surges, the delay of S2 increases".

[0127] Specifically, assume that under normal circumstances, the sending rate of S1 is 200Mbps and the forwarding delay of S2 is 5ms. If the real-time data changes to 400Mbps and 10ms, the association feature is manifested as "doubling of traffic is related to doubling of delay". This feature can reflect the potential connection of faults between devices. When extracting data flow direction features from the interaction data.

[0128] Preferably, the embodiments of the present application pay attention to the direction and intensity of the data.

[0129] For example, the traffic from S1 to S2 accounts for 60% of the total, and the traffic from S2 to S3 accounts for 30%.

[0130] It can be understood that the distance can be defined as the physical distance or the logical hop count.

[0131] Exemplarily, if the hop from S1 to S2 is one hop and the hop from S2 to S3 is two hops, the state change may be concentrated at S2, such as a greater increase in delay. The initial range can be set to directly connected devices, that is, S1 and S2. When analyzing the fault relationship for the state change, the time series processing technology can come in handy.

[0132] Specifically, record the delay of S2 with a 5-minute sliding window and find that it gradually rises from 5ms to 10ms, and the trend indicates that the fault relationship intensifies.

[0133] In one embodiment, if the traffic of S1 remains at a high level, S2 may lose packets due to buffer overflow, and the dynamic relationship is manifested as "traffic overload causes packet loss". This trend analysis can reveal the process of fault evolution. When extracting the influence range feature from the dynamic relationship.

[0134] For example, the packet loss of S2 will affect the download speed of S3, and the range covers the path from S2 to S3. When judging the propagation path by the association feature.

[0135] Preferably, the analysis shows that the traffic overload of S1 is the starting point, S2 is the key propagation node, and finally affects S3. The distribution value can be set as the traffic proportion of the affected devices, such as S3 accounting for 30%. If the preset threshold is 20%, then 30% exceeds the threshold and needs to be adjusted. When adjusting the device distance through the data flow direction.

[0136] In a possible implementation, the bandwidth from S1 to S2 can be increased, or the number of hops from S2 to S3 can be reduced.

[0137] For example, optimizing two hops to one hop reduces the latency from 10 ms to 7 ms. After updating the dynamic evaluation value with the optimized state change, the association rule mining shows that the feature performance range shrinks, the influence of S3 decreases from 30% to 15%, and finally the association distribution tends to be stable. This adjustment can improve the network robustness.

[0138] It can be understood that if the influence of S3 decreases to 15%, the traffic distribution from S1 to S2 is optimized.

[0139] For example, adding a backup path to share the S1 traffic stabilizes the fault relationship step by step. This method can effectively reduce the risk of cascading failures between multiple devices.

[0140] In a design, the electronic device can also cluster the associated network devices according to the failure rate to obtain a clustering result. The electronic device divides risk regions in the network topology based on the clustering result to obtain multiple risk regions; among them, the network devices in one risk region correspond to the same clustering, and different risk regions correspond to different risk levels. Further, the electronic device can determine the risk boundary in the network topology according to the multiple risk regions, and visually display the risk boundary and the multiple risk regions.

[0141] It can be understood that the topological structure reflects the connection relationship and data flow direction between devices.

[0142] For example, in an enterprise network, it includes a core router R1, an edge switch R2, and a terminal server R3. The historical fault records may show that R1 has crashed 3 times, the packet loss rate of R2 reaches 10%, and R3 has timed out 5 times. When using the clustering analysis method to divide the risk distribution area.

[0143] Preferably, classification is based on the failure frequency and influence range.

[0144] Specifically, the crash of R1 affects the entire network and is classified into the high-risk area; the packet loss of R2 affects some subnets and is classified into the medium-risk area; the timeout of R3 only affects itself and is classified into the low-risk area. When extracting the manifestation form of the associated features from the classification boundary.

[0145] For example, the feature of the high-risk area may be "a single point of failure causes global paralysis", and the medium-risk area is "local packet loss causes delay".

[0146] In a possible implementation, by judging the regional division of the risk distribution through the topological structure, the dependence between devices can be concerned.

[0147] Exemplarily, if R1 is an upstream node of R2 and the probability that R2's downtime is affected by R1 reaches 70%, then the high-risk area is concentrated in the upstream core devices. When analyzing the relationship between historical faults and risk distribution based on the distribution characteristics, time series processing technology can reveal trends.

[0148] Specifically, by analyzing R1's downtime records with a 30-day window, it is found that the fault interval per month shortens from 10 days to 5 days, indicating an intensification of risks. When judging the propagation path of the manifestation form through the change trend.

[0149] Preferably, the analysis shows that R1's faults spread downstream, and R2's packet loss rate increases by 10% accordingly. The propagation path is from R1 to R2. When extracting the boundary range of multiple risks from the regional division.

[0150] For example, the boundary of the medium-risk area may be set as the subnet with a packet loss rate exceeding 5%, and the range covers R2 and its downstream nodes. When determining the adjustment direction of the topological structure.

[0151] In one embodiment, if R1's failure rate exceeds the preset threshold of 2 times per month, a standby router can be added to share the traffic. If the adjustment direction exceeds the threshold, the risk distribution is updated through the fault records.

[0152] For example, the new fault shows that R2's packet loss rate rises to 15%, and the range of the medium-risk area expands after the update. When obtaining the new classification boundary from the updated distribution, cluster analysis may upgrade R2 to a high-risk area. When using the cluster analysis method to analyze the distribution range of the manifestation form.

[0153] Specifically, the new boundary may show that the high-risk area expands from a single R1 to R1 and R2, and the low-risk area shrinks to R3. The optimized risk distribution can more accurately locate the problem area.

[0154] For example, after the adjustment, R1's failure rate drops to 1 time per month, and R2's packet loss rate stabilizes at 5%, indicating a more balanced risk distribution. This way can enhance the network's fault tolerance.

[0155] It should be noted that cluster analysis depends on the integrity of historical data. If records are missing, real-time monitoring data may need to be supplemented to ensure the accuracy of the boundary.

[0156] In one embodiment, adding monitoring points to cover the downstream of R2 can further refine the risk area and enhance the pertinence of the adjustment.

[0157] It can be understood that the traffic distribution reflects the real-time load of data transmission in the network, while the topological change shows the dynamic adjustment of device connections.

[0158] For example, in an enterprise network, the traffic load of the core router R1 surges from the normal 50% to 90%, and the downstream connection of the edge switch R2 is temporarily disconnected due to maintenance. Based on the classification boundary judgment, the failure probability in the R1 area may increase to 20% due to traffic overload, exceeding the preset threshold of 15%. When the failure probability exceeds the preset threshold, the Monte Carlo simulation is used to quantify the potential impact.

[0159] Preferably, the failure scenarios after the overload of R1 are simulated through multiple random samplings.

[0160] Specifically, the simulation shows that the network-wide latency increases by 30% in 80% of the scenarios, and the R2 subnet is paralyzed in 50% of the scenarios. The potential impact distribution ranges from the whole network to local subnets.

[0161] For example, the boundary of the high-risk area is set as the nodes with traffic overload exceeding 80% and their direct downstream, including the upstream parts of R1 and R2.

[0162] In a possible implementation, the traffic pressure on R1 is preferably relieved first.

[0163] Exemplarily, the adjustment direction can be to divert 50% of the traffic to the standby router, and the initial priority order is R1 prior to R2.

[0164] Specifically, the sudden increase in the traffic of R1 and the disconnection of the connection to R2 are the core features. After optimization, the sorting is adjusted to give equal importance to R1 and R2 because the disconnection of R2 exacerbates the downstream risk. When using the optimized sorting result to adjust the boundary of the area division.

[0165] For example, both R1 and R2 are included in the high-risk area, and the new risk distribution range extends to the downstream subnet of R2.

[0166] In one embodiment, the definition of the classification boundary is adjusted to the area with traffic overload exceeding 75% or downstream disconnection exceeding 2 times, and the adjusted priority sequence is R1, R2, R3.

[0167] It should be noted that the adjusted sequence pays more attention to the cascading impact of topological changes on the downstream. When considering from multiple aspects, for the Monte Carlo simulation, the simulation is not limited to traffic overload, and the hardware aging of R1 can also be considered, with the failure rate increasing from 5% to 15%, and the impact distribution being wider. For the extraction of key features, if the downstream nodes of R2 feedback that the latency exceeds 200 milliseconds, the feature weight needs to be dynamically adjusted, and the sorting is more accurate.

[0168] Preferably, this method can quickly locate high-risk nodes and optimize resource allocation. From the perspective of the adjustment direction, if the diversion effect of the standby router is not good, the bandwidth can be further increased to ensure network stability. This multi-faceted analysis ensures the comprehensiveness of risk assessment and the feasibility of adjustment.

[0169] In one design, the electronic device can also sort multiple risk areas in descending order of risk level to obtain a sorting result, and determine the risk prevention strategy of the network topology according to the sorting result.

[0170] In an enterprise network, the historical data of the core router R1 may show that traffic peaks often occur at 10 am, and the downstream connection of the edge switch R2 is disconnected about 3 times a week. The feature set may include traffic load peaks, connection stability, etc. When analyzing the risk trend through the feature set.

[0171] Specifically, the trend may show that the risk of traffic overload of R1 is gradually increasing, while the downstream interruption of R2 shows periodic fluctuations.

[0172] In a possible implementation, the reinforcement learning algorithm adjusts the prediction model parameters through trial and error.

[0173] For example, increase the weight of the traffic peak on the risk. The adjusted risk trend prediction result may show that the risk of R1 rises to 25% within the next 24 hours, and that of R2 is 10%.

[0174] Exemplarily, the high-risk area may cover R1 and its downstream subnets, while the influence range of R2 is more local. The new area sorting sequence is adjusted to give priority to R1 and then R2. When obtaining the update direction of feature extraction from historical data.

[0175] Preferably, pay attention to the duration of traffic overload of R1 and the recovery speed after the disconnection of R2. The optimized feature set may add an index of "interruption recovery duration". When updating the input conditions of the prediction model with the optimized feature set.

[0176] In one embodiment, the duration of traffic overload of R1 is set as a key variable, and the improved risk trend distribution shows that the risk peak of R1 advances to 8 hours later. When adjusting the execution order of the prevention and control strategy.

[0177] For example, first perform traffic diversion on R1, and then repair the connection of R2. The final strategy deployment plan may be to divert 30% of the traffic to the standby router and upgrade the R2 line. If the risk trend of a certain area exceeds the preset threshold, for example, the risk of R1 reaches 30% exceeding the threshold of 25%, the impact range is quantified through Monte Carlo simulation.

[0178] Specifically, the simulation shows that in 60% of the scenarios, the network-wide delay increases by 20%, and in 30% of the scenarios, the R1 subnet is paralyzed. The adjusted prevention and control priority sequence may become equal emphasis on R1 and R2, and give priority to ensuring the stability of the core nodes.

[0179] It should be noted that when the reinforcement learning adjusts the model parameters, it can learn from historical false alarms to avoid overestimating the impact of low-risk areas.

[0180] In a possible implementation, simulating the hardware aging of R1 causes the failure rate to rise to 10%, and combined with traffic characteristic analysis, it ensures that the prediction is closer to the actual situation.

[0181] Preferably, multi-faceted analysis such as combining interruption frequency and delay feedback can improve the accuracy of trend prediction and optimize the efficiency of resource allocation.

[0182] For example, if the downstream delay of R2 exceeds 150 milliseconds, it is preferred to repair the connection rather than simply divert traffic. This method ensures that the prevention and control strategy is more targeted and effectively reduces network risks by dynamically adjusting the feature weights.

[0183] In one embodiment, if the diversion effect of the prevention and control strategy adjustment direction is insufficient, bandwidth support can be increased to ensure long-term stability.

[0184] It can be understood that this multi-level analysis and optimization can quickly locate high-risk nodes and improve network robustness, providing a reliable basis for subsequent operation and maintenance.

[0185] Specifically, by the interaction data of the real-time state and the network topology, the update frequency of the dynamic relationship can be obtained, which can be understood as the speed of state change between nodes in the network.

[0186] For example, in an enterprise network, the data exchange between the core server S4 and the database server S5 may be updated once a minute, while the connection status between the edge node N1 and the external gateway changes once every 5 minutes.

[0187] Specifically, regions with a high update frequency often mean active data streams and may hide risk points.

[0188] Exemplarily, the high-frequency interaction between S4 and S5 may cause delays due to a sharp increase in load. The initial risk distribution map may show that the risk value in the S4 region is 20% and that in the N1 region is 5%. When generating a risk heat map according to the update frequency of the dynamic relationship.

[0189] In a possible implementation, the risk level can be represented by the color depth.

[0190] For example, the S4 region shows dark red, indicating high risk, while the N1 region is light yellow, indicating low risk.

[0191] It should be noted that the heat map can also visually demarcate the boundaries of high-risk areas. For example, the subnet around S4 is framed as the key monitoring scope, while the impact of N1 is limited to a single link.

[0192] Preferably, the bandwidth occupancy and dependency degree between nodes can be concerned.

[0193] For example, the connection bandwidth occupancy between S4 and its downstream switch E1 reaches 80%, far exceeding the threshold of 50%, while the connection strength between N1 and the gateway is only 30%.

[0194] Specifically, the high connection strength of S4 indicates that its paralysis may affect the entire subnet, and the priority of the prevention and control strategy naturally tilts towards S4. When the connection strength in the high-risk area exceeds the preset threshold and the coverage range is quantified through Monte Carlo simulation.

[0195] In one embodiment, after simulating the failure of S4, 70% of the scenarios show that the downstream delay increases by 15 milliseconds, and 20% of the scenarios lead to the paralysis of the subnet. The adjusted execution order may become to process S4 first and then focus on N1.

[0196] It should be noted that the simulation can also reveal hidden risks, such as the chain reaction of E1 after S4 goes down, so as to optimize the priority distribution. After adopting the adjusted execution order and updating the real-time state of the risk distribution, the optimized heat map display may show that the risk of S4 drops to 10% and N1 remains stable.

[0197] Exemplarily, this kind of dynamic update can help the operation and maintenance personnel quickly lock the changing high-risk points and improve the response efficiency.

[0198] In one possible implementation, bandwidth expansion is deployed for S4, while N1 strengthens link redundancy, and the final sequence may be S4 first and N1 second.

[0199] It can be understood that this kind of deployment can effectively balance resource allocation and ensure the stability of the core nodes.

[0200] For example, the "bandwidth saturation duration" of S4 and the "link switching success rate" of N1 are added.

[0201] In one embodiment, if the saturation duration of S4 exceeds 10 minutes, priority is given to capacity expansion.

[0202] Preferably, these features can make the prevention and control more accurate. When obtaining the iterative scheme of the prevention and control strategy.

[0203] Specifically, if the risk of S4 is still high after capacity expansion, a standby server can be introduced to share the traffic.

[0204] Exemplarily, after verifying from multiple aspects such as delay feedback and connection stability, the iterative scheme can significantly improve the network robustness and provide a flexible basis for subsequent operation and maintenance.

[0205] The technical solutions provided by the embodiments of this application at least bring the following beneficial effects: The embodiments of this application collect device logs, traffic records, and topology information through a distributed storage system, extract the characteristics of the failure probability distribution using time series analysis, and obtain the initial risk quantification result. Combining historical data and real-time status, a risk trend prediction model is constructed using machine learning algorithms. When the traffic surges beyond the threshold, the fault propagation path is calculated through a graph neural network to evaluate the impact degree. Further, association rule mining and clustering analysis are used to divide the risk distribution area, and the potential impact degree is quantified through Monte Carlo simulation. Finally, a reinforcement learning algorithm is used to optimize the risk prevention and control strategy, and a heat map of the risk distribution is generated. The present invention realizes the accurate prediction and dynamic prevention and control of the operation risks of network switches, improving network security and reliability.

[0206] The above mainly introduces the solutions provided by the embodiments of this application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0207] In an exemplary embodiment, the embodiments of this application also provide a network device fault prediction device. Figure 4 It is a schematic diagram of the composition of the network device fault prediction device provided by the embodiments of this application. As Figure 4 shown, the network device fault prediction device includes: an acquisition unit 401, a determination unit 402, and a processing unit 403.

[0208] An acquisition unit 401, configured to acquire historical operation characteristic information of a target network device, and construct a risk prediction model according to the historical operation characteristic information; the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation state of the target network device; a determination unit 402, configured to determine the failure rate of the target network device based on the current operation data of the target network device and the risk prediction model; the determination unit 402 is further configured to, when the failure rate of the target network device is greater than or equal to a preset failure rate, acquire the network topology corresponding to the target network device, and determine the fault propagation path according to the network topology; the determination unit 402 is further configured to determine the failure rate of the associated network devices on the fault propagation path according to the fault propagation path and the failure rate of the target network device; the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

[0209] In a possible implementation, the target network device is a central device in the network topology, or the data volume of the historical operation data of the target network device is less than or equal to a preset threshold.

[0210] In a possible implementation, the determination unit 402 is specifically configured to: process the network topology by using a trained graph neural network to obtain a fault propagation path; the graph neural network reflects the fault relationship between network devices in the network topology through the edges between nodes.

[0211] In a possible implementation, the determination unit 402 is specifically configured to: for any associated network device on the fault propagation path, acquire the fault relationship between the associated network device and the target network device; the fault relationship includes the influence degree of the target network device on the associated network device; determine the failure rate of the associated network device on the fault propagation path based on the influence degree of the target network device on the associated network device and the failure rate of the target network device.

[0212] In a possible implementation, the network device fault prediction device further includes a processing unit 403, and the processing unit 403 is configured to: for any associated network device on the fault propagation path, when the influence degree of the target network device on the associated network device is greater than or equal to a preset influence degree, update the network topology according to the data flow direction between the target network device and the associated network device to adjust the device spacing between the target network device and the associated network device, so as to obtain an updated network topology; the influence degree of any associated network device on the fault propagation path in the updated network topology affected by the target network device is less than or equal to the preset influence degree.

[0213] In a possible implementation, the processing unit 403 is further configured to: cluster the associated network devices according to the failure rate to obtain a clustering result; divide risk regions in the network topology based on the clustering result to obtain a plurality of risk regions; the network devices in one risk region correspond to the same clustering, and different risk regions correspond to different risk levels; determine the risk boundaries in the network topology according to the plurality of risk regions, and visually display the risk boundaries and the plurality of risk regions.

[0214] In a possible implementation, the determining unit 402 is further configured to: sort the plurality of risk regions in descending order of risk level to obtain a sorting result, and determine a risk prevention strategy for the network topology according to the sorting result.

[0215] It should be noted that Figure 4 the division of modules herein is illustrative, merely a logical function division, and there may be other division methods in actual implementation. For example, two or more functions may also be integrated into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional unit.

[0216] In an exemplary embodiment, the embodiment of the present application further provides a computer-readable storage medium, including software instructions, which when running on an electronic device, cause the electronic device to execute any one of the methods provided in the above embodiments.

[0217] In an exemplary embodiment, the embodiment of the present application further provides a computer program product containing computer-executable instructions, which when running on an electronic device, cause the electronic device to execute any one of the methods provided in the above embodiments.

[0218] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer-executable instructions. When the computer-executable instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer-executable instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer-executable instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0219] Although the present application has been described in conjunction with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0220] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present application defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

[0221] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An intelligent fault detection method for a network switch, characterized in that: The method comprises: Acquire historical operation characteristic information of the target network device, and construct a risk prediction model based on the historical operation characteristic information; the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation status of the target network device; Determining a failure rate of the target network device based on current operating data of the target network device and the risk prediction model; When the failure rate of the target network device is greater than or equal to a preset failure rate, obtaining a network topology corresponding to the target network device, and determining a fault propagation path according to the network topology; According to the fault propagation path and the failure rate of the target network device, the failure rate of the associated network device on the fault propagation path is determined; the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

2. The method according to claim 1, characterized in that The target network device is a central device in the network topology, or the amount of historical operation data of the target network device is less than or equal to a preset threshold.

3. The method according to claim 1, characterized in that The determining of the fault propagation path according to the network topology includes: The network topology is processed using a trained graph neural network to obtain the fault propagation path; the graph neural network reflects the fault relationship between each network device in the network topology through the edges between each node.

4. The method according to claim 3, characterized in that The determining, according to the fault propagation path and the failure rate of the target network device, the failure rate of the associated network device on the fault propagation path includes: For any associated network device on the fault propagation path, obtaining a fault relationship between the associated network device and the target network device; the fault relationship includes the degree of influence of the target network device on the associated network device; Based on the degree of influence of the target network device on the associated network device and the failure rate of the target network device, the failure rate of the associated network device on the fault propagation path is determined.

5. The method according to claim 4, characterized in that The method further comprises: For any associated network device on the fault propagation path, when the degree of influence of the target network device on the associated network device is greater than or equal to the preset degree of influence, the network topology is updated according to the data flow direction between the target network device and the associated network device to adjust the device distance between the target network device and the associated network device to obtain an updated network topology; in the updated network topology, the degree of influence of the target network device on any associated network device on the fault propagation path is less than or equal to the preset degree of influence.

6. The method according to claim 1, characterized in that The method further comprises: Clustering the associated network devices according to the failure rate to obtain clustering results; Dividing the risk area in the network topology based on the clustering result to obtain multiple risk areas; network devices in a risk area correspond to the same cluster, and different risk areas correspond to different risk levels; According to the multiple risk areas, a risk boundary in the network topology is determined, and the risk boundary and the multiple risk areas are visualized.

7. The method according to claim 6, characterized in that The method further comprises: The multiple risk areas are sorted in descending order of risk level to obtain a sorting result, and a risk prevention strategy for the network topology is determined according to the sorting result.

8. An intelligent fault detection system for a network switch, characterized in that: The intelligent fault detection system includes a network equipment fault prediction device, which includes an acquisition unit and a determination unit; The acquisition unit is used to acquire historical operation characteristic information of the target network device and construct a risk prediction model according to the historical operation characteristic information; the historical operation characteristic information is used to reflect the correlation between the historical operation data and the historical operation status of the target network device; The determining unit is used to determine the failure rate of the target network device based on the current operation data of the target network device and the risk prediction model; The determining unit is further configured to obtain a network topology corresponding to the target network device when the failure rate of the target network device is greater than or equal to a preset failure rate, and determine a fault propagation path according to the network topology; The determination unit is also used to determine the failure rate of associated network devices on the fault propagation path based on the fault propagation path and the failure rate of the target network device; the failure rate of an associated network device is related to the failure rate of the target network device and the position of the associated network device in the fault propagation path.

9. An electronic device, characterized in that: include: Processor and memory; The memory stores instructions executable by the processor; When the processor is configured to execute the instructions, the electronic device implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The readable storage medium includes: software instructions; When the software instructions are executed in an electronic device, the electronic device implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fault processing method and device for smart power grid

    CN118868079A

  • High-voltage circuit breaker and heating fault prediction method and system for isolating switch of high-voltage circuit breaker

    CN119441763A

  • Distributed system fault positioning diagnosis method and system based on log analysis

    CN119668990A

Cited By

  • Routing loop early warning and suppression method and device, equipment and storage medium

    CN121396747A