Fault analysis method and device, storage medium and electronic equipment

Through adaptive monitoring and pre-training model analysis of switch port status information, the problem of delay and misjudgment of switch port status monitoring is solved, accurate identification and automated processing of network failures is realized, and the stability and operation and maintenance efficiency of the network system are improved.

CN120358131AInactive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510828061.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, switch port status monitoring has problems such as high monitoring delay, high misjudgment rate and inability to accurately identify the root cause of complex faults.

Method used

By adaptively monitoring the port status information of the switch port, the pre-trained fault analysis model is used to analyze the root cause of state change events and multi-dimensional performance data, and combined with event-driven models and dynamic threshold adjustments, the precise identification of network fault types is achieved.

Benefits of technology

It significantly reduces monitoring delays and false alarms, improves the comprehensiveness and accuracy of fault identification, can quickly locate and automate the handling of multiple network failures, and improves the stability and operation and maintenance efficiency of the network system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358131A_ABST
    Figure CN120358131A_ABST
Patent Text Reader

Abstract

The invention discloses a fault analysis method and device, a storage medium and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: firstly, carrying out the adaptive monitoring of the port state information of a switch port, and judging whether the port state information is changed or not; when it is monitored that the port state information changes, receiving a state change event sent by the switch port; then acquiring multi-dimensional performance data of a network connected with the switch port; and performing root cause analysis on the state change event and the multi-dimensional performance data by using a pre-trained fault analysis model to determine a network fault type. Compared with the prior art, the method effectively reduces the problems of monitoring delay and false alarm and missing alarm, improves the comprehensiveness and accuracy of fault recognition, can accurately distinguish a plurality of network fault types, further can automatically repair common faults through the self-healing function of the system, and improves the fault recognition efficiency. And the levels of rapid fault positioning and automatic processing are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a fault analysis method, apparatus, storage medium, and electronic device. Background Art

[0002] With the continuous expansion of the network scale, the real-time monitoring of the switch port status is crucial for ensuring network stability and service quality. By monitoring and analyzing the switch port status, potential fault hazards can be detected in a timely manner, and the network operation and maintenance efficiency can be improved.

[0003] Currently, in the related art, the port status of the switch is mainly monitored by relying on the traditional polling mechanism, and the fault is judged by regularly checking the port status. However, this method has problems such as high monitoring latency, high misjudgment rate, and inability to accurately identify the root cause of complex faults. Summary of the Invention

[0004] The present disclosure provides a fault analysis method, apparatus, storage medium, and electronic device. Its main purpose is to solve the problems of high monitoring latency, high misjudgment rate, and inability to accurately identify the root cause of complex faults in the related art.

[0005] In a first aspect, the present application provides a fault analysis method, including: Judging whether the port status information has changed by adaptively monitoring the port status information of the switch port; When it is monitored that the port status information has changed, receiving a status change event sent by the switch port; Obtaining multi-dimensional performance data of the network connected to the switch port; Performing root cause analysis on the status change event and the multi-dimensional performance data by using a pre-trained fault analysis model to determine the network fault type.

[0006] In a second aspect, the present application provides a fault analysis apparatus, including: A monitoring module configured to judge whether the port status information has changed by adaptively monitoring the port status information of the switch port; A receiving module configured to receive a status change event sent by the switch port when it is monitored that the port status information has changed; An obtaining module configured to obtain multi-dimensional performance data of the network connected to the switch port; A determining module configured to perform root cause analysis on the status change event and the multi-dimensional performance data by using a pre-trained fault analysis model to determine the network fault type.

[0007] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the method of the first aspect is implemented.

[0009] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0010] The fault analysis method, device, storage medium, and electronic device provided by the present disclosure, wherein the method includes: first, adaptively monitoring the port status information of the switch port to determine whether the port status information has changed; when it is monitored that the port status information has changed, receiving the status change event sent by the switch port; then obtaining the multi-dimensional performance data of the network connected to the switch port; and then using the pre-trained fault analysis model to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type. Compared with the current related technologies, the present application effectively reduces the problems of monitoring delay and false alarms / missing reports through the adaptive monitoring of port status changes, and combines the acquisition and fusion analysis of multi-dimensional performance data to improve the comprehensiveness and accuracy of fault identification. At the same time, using the pre-trained fault analysis model to infer the root cause of the status change event and its related features can accurately distinguish multiple network fault types. Furthermore, common faults can be automatically repaired through the self-healing function of the system, significantly improving the level of rapid fault location and automated processing.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 Shows a schematic flow chart of a fault analysis method provided by an embodiment of the present application; Figure 2 Shows a schematic flow chart of another fault analysis method provided by an embodiment of the present application; Figure 3 It shows a schematic flowchart of an example provided by an embodiment of the present application; Figure 4 It shows a schematic flowchart of an example provided by an embodiment of the present application; Figure 5 It shows a schematic flowchart of an example provided by an embodiment of the present application; Figure 6 It shows a schematic structural diagram of a fault analysis device provided by an embodiment of the present application. Detailed implementation manners

[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0016] Based on the problems existing in the above server configuration, in order to improve the technical problems of high monitoring latency, high misjudgment rate and inability to accurately identify the root cause of complex faults in the related art. This embodiment provides a fault analysis method, as Figure 1 shown, this method includes the following steps: Step 101: Determine whether the port status information has changed by adaptively monitoring the port status information of the switch port.

[0017] Exemplarily, an event-driven model can be adopted to establish an asynchronous notification channel between the switch and the controller based on the gNMI protocol, where the gNMI protocol is a network management interface that provides an open and standard way to retrieve network status information from network devices and perform configurations.

[0018] In some examples, the port status change is monitored in real time through hardware interrupts or message queues. When the physical layer status of the port changes, the switch actively pushes the status change event to the controller, and the push latency is less than 50 ms, realizing efficient and real-time perception of the port status.

[0019] Step 102: When it is detected that the port status information has changed, receive the status change event sent by the switch port.

[0020] Exemplarily, when the status change of the switch port occurs, the system can immediately trigger the notification mechanism and send the updated status information to the monitoring platform to ensure that fault events and operating status can be captured and responded to in a timely manner.

[0021] In some examples, the status change event is any state migration or attribute change that occurs during the operation of the network device (such as a switch) port. For example, such events include, but are not limited to, the physical connection status of the port changing from "up" to "down" or vice versa, a significant increase in the CRC error count, abnormal fluctuations in the optical module power, changes in the BGP adjacency relationship, etc. Whenever these key performance indicators change, the system records a status change event.

[0022] Step 103: Obtain multi-dimensional performance data of the network connected to the switch port.

[0023] For example, the multi-dimensional performance data may include: indicators such as port traffic, bit error rate, Cyclic Redundancy Check (CRC) error, link quality, port utilization rate, queue congestion status, protocol layer statistics, etc. These data reflect the operating status of the network device and the link health status from different levels, providing comprehensive data support for fault prediction, anomaly detection, and network optimization.

[0024] Step 104: Use the pre-trained fault analysis model to perform root cause analysis on the status change event and multi-dimensional performance data to determine the network fault type.

[0025] For example, the pre-trained fault analysis model can be an Artificial Intelligence (AI) model. The fault analysis model can be used to perform root cause analysis on the status change event of the switch port and its related multi-dimensional performance data. The system can automatically identify the fault type and output the corresponding confidence level, achieving efficient and accurate fault location. Among them, the fault types can include hardware faults, configuration errors, and link interference, etc.

[0026] In some examples, a centralized intelligent management platform can also be built to support the aggregated analysis and policy coordination of the status information of multiple switches, improve the overall intelligent operation and maintenance ability of the network, enhance the unity and timeliness of fault response, and provide strong guarantee for the stable operation of the network.

[0027] Compared with the related technologies, in this embodiment, first, the port status information of the switch port is adaptively monitored to determine whether the port status information has changed; when it is detected that the port status information has changed, a status change event sent by the switch port is received; then, multi-dimensional performance data of the network connected to the switch port is obtained; and then, a pre-trained fault analysis model is used to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type. Compared with the current related technologies, through the adaptive monitoring of port status changes, the problems of monitoring delay, false alarms, and missed alarms are effectively reduced. By combining the collection and fusion analysis of multi-dimensional performance data, the comprehensiveness and accuracy of fault identification are improved. At the same time, a pre-trained fault analysis model is used to infer the root cause of the status change event and its related features, which can accurately distinguish multiple network fault types. Furthermore, common faults can be automatically repaired through the self-healing function of the system, significantly improving the level of rapid fault location and automated processing.

[0028] As a refinement of this embodiment, the hard disk installation position can be determined in the following ways, but not limited to: Figure 2 as shown Figure 2 FIG. is a schematic flowchart of a fault analysis method provided by an embodiment of the present disclosure, including: Step 201: Dynamically determine the status oscillation threshold of the switch port according to the historical status change frequency and network load status of the switch port within a preset time period.

[0029] Exemplarily, the status oscillation of the switch port refers to the situation where the port of the network device repeatedly switches between different states in a short period of time. The status oscillation threshold is mainly used to identify and suppress the frequent changes of the port status. By introducing a dynamic threshold adjustment algorithm, the event trigger condition can be automatically optimized according to the historical status change frequency and the current network load situation. Through the intelligent analysis of the port status change trend and the network operation environment, the monitoring sensitivity is dynamically adjusted to avoid false alarms or missed alarms caused by a fixed status oscillation threshold, thereby improving the accuracy and adaptability of fault detection and realizing more refined and intelligent network monitoring.

[0030] In some examples, the port status oscillation threshold can be dynamically calculated. The preset time period can be set to T. Based on the moving average of the number of port status changes within the past T time window, an initial threshold baseline is determined; at the same time, the network traffic volatility is introduced as a weight factor to adjust Formula 1: (Formula 1) where α is a configurable sensitivity coefficient, Baseline represents the basic expected value of the status oscillation threshold under normal conditions, and ΔTraffic represents the deviation between the current traffic and the historical average value, thereby realizing the dynamic adjustment of the threshold.

[0031] Exemplarily, such as Figure 3 As shown, the system obtains the number of state changes in the past T time from historical data and calculates its moving average as the baseline to reflect the change frequency under normal conditions. Then, the system obtains the volatility of the current network traffic in real time, combines the historical baseline with the current network load situation, calculates the weight factor, and uses it to evaluate the rationality of the current state change. Then, the initial threshold is dynamically adjusted according to Formula 1 to ensure that it can adapt to the normal fluctuation range in different network environments. The adjusted threshold can be output for subsequent state change monitoring and fault detection.

[0032] Exemplarily, this embodiment proposes a dynamic threshold algorithm based on exponential weighting, enabling the alarm sensitivity to automatically adapt to changes in network load, improving detection accuracy and robustness. A lightweight AI inference engine can also be deployed locally on the switch to achieve sub-second real-time decision-making at the edge side, effectively supporting high-precision and low-latency network fault monitoring and response.

[0033] Step 202: Monitor the port status information through the asynchronous notification channel corresponding to the switch port based on the state oscillation threshold.

[0034] Exemplarily, by setting the state oscillation threshold, the maximum number of allowed state changes within a specific time window can be defined. If the state change frequency of the port exceeds this threshold, the system will trigger corresponding processing mechanisms, such as alarm notifications, automatic adjustment of port parameters, or temporary isolation of the port to prevent greater impact on the entire network. This mechanism helps maintain network stability, reduces unnecessary traffic fluctuations and service interruptions, and also helps quickly locate and solve potential problem sources.

[0035] Step 203: When it is monitored that the port status information exceeds the range of the state oscillation threshold, it is determined that the port status information has changed.

[0036] For example, when the system monitors that the port status information of the switch (such as state change frequency, CRC error rate, link stability, etc.) exceeds the range of the dynamically calculated state oscillation threshold, it is determined that an abnormal state change has occurred in the port. The state oscillation threshold is dynamically adjusted based on the baseline data of historical state changes and the current network load situation, and can adapt to the normal fluctuation range in different network environments.

[0037] Step 204: Receive the state change event sent by the switch port.

[0038] Exemplarily, when the actual monitored value continuously exceeds the status oscillation threshold, a status change event is triggered, indicating that the port may face potential failure risks, thereby initiating subsequent alarm, diagnosis, or self-healing mechanisms to achieve rapid response and precise handling of network anomalies.

[0039] For example, when the link state of a port changes from normal operation (up) to unavailable (down) due to reasons such as unstable connection, hardware failure, or configuration error, or vice versa, a status change event will be triggered. These events are crucial for monitoring the network health, quickly locating the root cause of problems, and implementing corresponding maintenance measures. By real-time monitoring and analyzing these events, actions can be taken in a timely manner to prevent potential problems from affecting the stability and service quality of the entire network.

[0040] Step 205: Obtain multi-dimensional performance data of the network connected to the switch port.

[0041] In some examples, by real-time monitoring of key metrics such as port traffic statistics, bit error rate, CRC error count, optical module transmit and receive power, BGP adjacency status, and link quality, not only can potential fault hazards be detected in a timely manner, such as port oscillation caused by configuration error or service interruption caused by hardware failure, but also the overall operation quality of the network can be quantitatively evaluated.

[0042] In addition, based on trend analysis of historical data, potential future problems can be predicted and preventive measures can be taken. Using these detailed performance data and combined with intelligent algorithm models, the system can achieve precise identification of fault types, root cause analysis, and automated response strategies, thereby ensuring the high availability and stability of network services.

[0043] Step 206: Use a pre-trained fault analysis model to perform root cause analysis on the status change event and multi-dimensional performance data to determine the network fault type.

[0044] Exemplarily, by using a pre-trained fault analysis model to conduct in-depth root cause analysis on the status change event and multi-dimensional performance data (such as port traffic, bit error rate, CRC error, optical module power, etc.), the network fault type can be accurately determined, including problems such as hardware failure, configuration error, and link interference.

[0045] In some examples, the LSTM network can be used to capture the status time series pattern and combined with a random forest classifier to achieve multi-label prediction, so as to provide detailed fault diagnosis results and their confidence levels, assisting the operation and maintenance personnel to quickly locate the root cause of problems and improve the fault handling efficiency and network reliability.

[0046] Optionally, the training process of the above-mentioned pre-trained fault analysis model may specifically include: obtaining historical status log data of switch ports; performing cleaning processing on the historical status log data; constructing a fault classification system corresponding to network fault types based on the processed historical status log data; and performing model training based on the fault classification system and the time-series sampling data of switch ports to obtain a fault analysis model.

[0047] In some examples, the data preparation and preprocessing stage of the fault analysis model includes comprehensive collection of historical logs of switch ports, which cover key metrics such as port status change events (such as UP / DOWN), traffic statistics, bit error rate, optical module power, etc., and ensure that more than a thousand fault events and normal state samples are included. The data cleaning step filters out invalid data, corrects timestamp problems and removes duplicate records, and then assigns clear labels to each fault type through a manual annotation session, establishing a classification system covering 87 different fault types. And to enhance the robustness of the model, methods such as window sliding sampling and adding Gaussian noise are also used to simulate network fluctuations.

[0048] Exemplarily, during the training process, the Xavier initialization method can be used to set the weights of the Long Short-Term Memory (LSTM) network, and the depth of the random forest tree is set to 15-20 layers. Among them, the loss function uses weighted cross-entropy loss, and the weight is specifically increased for rare fault types to ensure that the model can effectively learn the features of all fault types. Configure the Adam optimizer with an initial learning rate of 0.001, which gradually decreases with the training cycle. An early stopping mechanism can also be used to prevent overfitting, and the training is automatically terminated when the validation set loss has not improved for 10 consecutive epochs.

[0049] Exemplarily, in the evaluation and tuning stage of model training, the data can be divided into a training set, a validation set, and a test set in a ratio of 7:2:1 to ensure a uniform distribution of each fault type. The performance metric requirements are that the precision is higher than 92%, the recall rate exceeds 85%, and the F1-score is greater than 88%. Hyperparameter optimization uses Bayesian optimization technology to adjust the dimension of the LSTM hidden layer and the number of random forest trees to further improve the generalization ability and accuracy of the model.

[0050] Optionally, step 206 may specifically include: using the pre-trained fault analysis model to perform root cause analysis on state change events and multi-dimensional performance data to determine the network fault type, including: based on the state change events, using the long short-term memory network to capture the state time-series pattern of the switch port; combining the state time-series pattern and the random forest classifier to perform multi-label prediction on the multi-dimensional performance data to determine the network fault type.

[0051] Exemplarily, in terms of the model architecture design of the fault analysis model, key features including the state change frequency, CRC error count, optical power attenuation gradient, and the number of changes in the Border Gateway Protocol (BGP) adjacency state are used as inputs. The fault analysis model can adopt a hybrid structure, where the LSTM layer is responsible for capturing the time series dependencies of port states, while the random forest classifier is used to process static features to achieve multi-label prediction. In addition, an attention mechanism is introduced to dynamically weight the importance of features in different dimensions, such as significantly increasing the weight of the optical power mutation feature, so as to more accurately identify potential problems.

[0052] In some examples, as Figure 4 shown, the AI-based network fault diagnosis system is divided into a training phase and an inference phase. In the training phase, historical fault data is first collected, and data cleaning and feature engineering are performed to ensure data quality and the effectiveness of model inputs. Then, the LSTM and random forest algorithms are used to jointly train the data to build a hybrid model that can capture the time series features and the correlations of multi-dimensional performance metrics. After model evaluation, the number of model parameters is reduced through lightweight compression technology to adapt to the resource limitations of edge computing, and finally, it is deployed to the actual environment. In the inference phase, the system collects switch port data in real time and performs preprocessing to extract key features. Subsequently, the processed data is input into the deployed AI model for inference analysis, and the fault type and its confidence level are output. If the confidence level is higher than 90%, the self-healing mechanism is triggered to automatically execute corresponding repair measures; otherwise, the system reports the fault information to the manual processing link, and the operation and maintenance personnel further investigate and solve it.

[0053] Optionally, the method of this embodiment may specifically further include: if the network fault type is port oscillation caused by network configuration errors, the network configuration is restored to the previous stable configuration version; if the network fault type is a hardware-related fault, the switch port is marked as the disabled state and the standby link is triggered to switch.

[0054] In some examples, for the port oscillation problem caused by configuration errors, it can automatically identify the anomaly and roll back to the nearest stable configuration version, thus quickly restoring the normal operation state of the port; for hardware-related faults, the port can be marked as the disabled state and the standby link switch can be triggered to maximize network stability and service continuity. It effectively reduces manual intervention and fault recovery time, improves network stability and operation and maintenance efficiency, and reduces the risk of persistent network interruptions caused by configuration mistakes.

[0055] Optionally, the method of this embodiment may specifically further include: judging the severity of the fault of the network connected to the switch port based on the network fault type; finding and executing the corresponding alarm policy in the hierarchical alarm system according to the severity of the fault.

[0056] Exemplarily, the system can automatically divide the alarm levels according to the fault type, influence range and severity, judge the severity of the fault of the link connected to the switch port, divide the fault into different levels according to the preset hierarchical standard, and then push the alarm information to the corresponding maintenance personnel in a timely manner through text messages, emails or work order systems. Thus, it ensures that different levels of faults can receive differential responses and processing, improves the operation and maintenance efficiency and the timeliness of fault handling, and effectively guarantees the stable operation of the network system.

[0057] Optionally, the method of this embodiment may specifically further include: analyzing the historical port status information of the switch port to generate a port status report; generating optimization suggestions for the network topology structure according to the port status report.

[0058] In some examples, the system supports lightweight deployment and continuous optimization. By using model pruning technology, the number of parameters of the AI model is compressed to 30% of the original model, significantly reducing the computing and storage requirements and adapting to the local edge computing resources of the switch. At the same time, it has the ability of online learning. After deployment, it can continuously collect new fault cases and perform incremental training every month to update the model weights, continuously improving the diagnostic accuracy and adaptability. With the support of the centralized intelligent management platform, the system can be compatible with multi-vendor switch devices through standard protocols such as SNMP and gNMI to achieve unified monitoring and policy coordination.

[0059] Exemplarily, as Figure 5 shown in the intelligent switch port monitoring and fault handling system architecture, it realizes the real-time monitoring of the network status, rapid fault location and automatic processing through the event-driven model, dynamic threshold adjustment, AI root cause analysis and centralized management platform. The switch transmits the port status information to the controller in real time through the gNMI protocol, and then the controller forwards it to the centralized management platform for unified management and policy coordination. At the same time, the switch also directly sends control flow information such as configuration rollback and interface disabling to the centralized management platform through the SNMP or gNMI protocol. On the monitoring platform, the message middleware is responsible for receiving and serializing the event stream, and the event processor makes a preliminary analysis according to the multi-dimensional data and the historical records in the time series database, and triggers the dynamic threshold calculation module to optimize the detection sensitivity and reduce false alarms and missed alarms. The alarm engine generates severe alarms or general alarms according to the analysis results and notifies the maintenance personnel by means of text messages, emails, etc. The AI inference engine deeply analyzes the root cause of the fault, outputs a confidence evaluation, and automatically repairs common faults in combination with the self-healing function. Finally, all policy coordination information is transmitted to the highly available cluster to ensure the stable operation and efficient operation and maintenance of the system.

[0060] For example, in a production environment, real-time monitoring of switch port status is crucial for ensuring network stability and service quality. The method of this embodiment realizes efficient monitoring and automated processing of port status through an event-driven model, dynamic threshold adjustment, AI root cause analysis, and a centralized management platform. First, an asynchronous notification channel based on the gNMI protocol is enabled on the switch to ensure that when the physical layer status of the port changes, the switch can actively push status change events to the controller within 50 ms. The controller receives the events through a message queue and transfers them to the monitoring platform in real time, thus achieving a rapid response to port status changes. Second, a dynamic threshold adjustment mechanism is introduced. The threshold baseline is calculated based on the number of port status changes in the past hour, and the network traffic volatility is combined as a weighting factor to dynamically adjust the threshold. This adjustment mechanism can optimize the triggering conditions according to the actual operating conditions of the network, reduce false alarms and missed alarms, and improve the accuracy of monitoring. At the same time, multi-dimensional data such as port traffic, bit error rate, CRC errors, and link quality are collected and uploaded to the monitoring platform in real time through the gNMI protocol. These data provide comprehensive information support for subsequent fault analysis. In terms of fault analysis, a pre-trained LSTM network is used to capture the state time series pattern, and a random forest classifier is combined for multi-label prediction to output the fault type (such as hardware failure, configuration error, link interference) and confidence level. The training data of the AI model comes from historical fault data in the past six months, ensuring the accuracy and reliability of the analysis results. In addition, a centralized management platform that supports protocols such as SNMP and gNMI is built to realize multi-switch status aggregation and policy coordination. This platform is compatible with switches from different manufacturers, including Ruijie, H3C, etc., and provides a unified management interface for network operation and maintenance. Finally, according to the severity of the fault, maintenance personnel are notified via text message, email, or a work order system. For port oscillations caused by configuration errors, the system automatically rolls back to the most recent stable configuration version; for hardware-related faults, the port is marked as disabled and the standby link is triggered to switch. These self-healing functions significantly reduce the workload of operation and maintenance personnel and improve the efficiency of network operation and maintenance. Through the implementation of the method of this embodiment in network monitoring, real-time monitoring of port status changes, rapid analysis of fault root causes, and automated processing are achieved, significantly reducing network interruption time, optimizing the use of monitoring resources, and providing strong guarantee for the stable operation of the production line network.

[0061] Compared with the current related technologies, in this embodiment, by introducing a dynamic threshold adjustment mechanism, the oscillation detection sensitivity is automatically optimized by combining the historical state change frequency and the real-time network load, avoiding the misjudgment problem caused by a fixed threshold, and a feature input system integrating multi-dimensional performance data such as port state change events, CRC errors, optical module power, and traffic statistics is constructed to improve the characterization ability of fault scenarios. On this basis, a hybrid AI model based on LSTM and random forest is used for root cause analysis, which can accurately identify various fault types such as hardware faults, configuration errors, and link interference, and output a confidence evaluation. It also combines the built-in self-healing strategy of the system. After identifying a fault that can be automatically repaired, it can immediately trigger configuration rollback or link switching operations, significantly improving the adaptability and recovery ability of the network system.

[0062] An embodiment of the present application also provides a fault analysis device, as Figure 1 a specific implementation of the method shown in Figure 6 shown, the device includes: a monitoring module 31, a receiving module 32, an obtaining module 33, and a determining module 34.

[0063] The monitoring module 31 is configured to adaptively monitor the port state information of the switch port to determine whether the port state information has changed; The receiving module 32 is configured to receive the status change event sent by the switch port when it is detected that the port state information has changed; The obtaining module 33 is configured to obtain multi-dimensional performance data of the network connected to the switch port; The determining module 34 is configured to perform root cause analysis on the status change event and the multi-dimensional performance data by using a pre-trained fault analysis model to determine the network fault type.

[0064] In some examples of this embodiment, the monitoring module 31 is specifically configured to dynamically determine the status oscillation threshold of the switch port according to the historical state change frequency and the network load status of the switch port within a preset time period; based on the status oscillation threshold, monitor the port state information through the asynchronous notification channel corresponding to the switch port; when it is detected that the port state information exceeds the range of the status oscillation threshold, it is determined that the port state information has changed.

[0065] In some examples of this embodiment, the determining module 34 is specifically configured to obtain the historical state log data of the switch port; perform cleaning processing on the historical state log data; construct a fault classification system corresponding to the network fault type based on the processed historical state log data; perform model training based on the fault classification system and the time-series sampling data of the switch port to obtain a fault analysis model.

[0066] In some examples of this embodiment, the determination module 34 is further specifically configured to capture the state time series pattern of the switch port based on the state change event by using a long short-term memory network; Combine the state time series pattern and a random forest classifier to perform multi-label prediction on the multi-dimensional performance data to determine the network fault type.

[0067] In some examples of this embodiment, the determination module 34 is further specifically configured to, if the network fault type is port oscillation caused by network configuration errors, restore the network configuration to the previous stable configuration version; if the network fault type is a hardware-related fault, mark the switch port as a disabled state and trigger the switching of the standby link.

[0068] In some examples of this embodiment, the determination module 34 is further specifically configured to judge the severity of the fault of the network connected to the switch port based on the network fault type; and find and execute the corresponding alarm policy in the hierarchical alarm system according to the severity of the fault.

[0069] In some examples of this embodiment, the determination module 34 is further specifically configured to analyze the historical port state information of the switch port to generate a port state report; and generate optimization suggestions for the network topology structure according to the port state report.

[0070] It should be noted that for other corresponding descriptions of each functional unit involved in the fault analysis device provided in this embodiment, reference can be made to the corresponding description in Figure 1 and will not be elaborated here.

[0071] Based on the methods as shown in Figure 1 and Figure 2 correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the methods as shown in Figure 1 and Figure 2 above.

[0072] Based on the methods as shown in Figure 1 and Figure 2 correspondingly, this embodiment further provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, it implements the methods as shown in Figure 1 and Figure 2 above.

[0073] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.

[0074] Based on the above-mentioned methods as Figure 1 and Figure 2 shown, and Figure 6 the virtual device embodiments shown, in order to achieve the above object, the embodiments of the present application further provide an electronic device, such as a personal computer, a server, and the device includes a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to implement the methods as Figure 1 and Figure 2 shown.

[0075] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen, an input unit such as a keyboard, etc., and optionally the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0076] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0077] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement the communication between the components inside the storage medium, and the communication between other hardware and software in the information processing physical device.

[0078] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current related technologies, this embodiment constructs an intelligent switch port monitoring and fault handling system, which realizes real-time monitoring and rapid response to port status changes through an event-driven model, ensuring that status change events are captured in time and the notification mechanism is triggered. The system introduces a dynamic threshold adjustment algorithm, which automatically optimizes the triggering conditions according to the historical change frequency and network load, effectively reducing false alarms and missed alarms. Combining with a pre-trained AI fault analysis model, it conducts root cause analysis on status change events and multi-dimensional performance data, accurately identifies fault types such as hardware faults, configuration errors, or link interference, and outputs corresponding confidence evaluations. At the same time, a self-healing function is integrated. When common faults (such as port oscillation caused by configuration errors) are detected, it can automatically roll back the configuration or switch to an alternative link to improve network stability. In addition, a centralized intelligent management platform is also constructed, which supports the aggregated analysis and policy coordination of the status information of multiple switches, realizes unified monitoring, intelligent alarm, and efficient operation and maintenance, and comprehensively improves the response speed and automation management level of network faults.

[0079] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0080] The above are only the specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments herein, but will conform to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A fault analysis method, characterized in that Including: Adaptively monitor the port status information of the switch port to determine whether the port status information has changed; When it is detected that the port status information has changed, receive the status change event sent by the switch port; Obtain multi-dimensional performance data of the network connected to the switch port; Use a pre-trained fault analysis model to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type.

2. The method according to claim 1, characterized in that The monitoring of the port status information of the switch port to determine in real time whether the port status information has changed includes: Dynamically determine the status oscillation threshold of the switch port according to the historical status change frequency and network load status of the switch port within a preset time period; Based on the status oscillation threshold, monitor the port status information through the asynchronous notification channel corresponding to the switch port; When it is detected that the port status information exceeds the range of the status oscillation threshold, it is determined that the port status information has changed.

3. The method according to claim 1, wherein The training process of the pre-trained fault analysis model includes: Obtain the historical status log data of the switch port; Clean and process the historical status log data; Based on the processed historical status log data, construct a fault classification system corresponding to the network fault type; Perform model training based on the fault classification system and the time series sampling data of the switch port to obtain a fault analysis model.

4. The method according to claim 3, wherein The use of the pre-trained fault analysis model to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type includes: Based on the status change event, use a long short-term memory network to capture the status time series pattern of the switch port; Combine the status time series pattern and a random forest classifier to perform multi-label prediction on the multi-dimensional performance data to determine the network fault type.

5. The method according to claim 1, characterized in that, After using the pre-trained fault analysis model to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type, the method further includes: If the network fault type is port oscillation caused by network configuration error, restore the network configuration to the previous stable configuration version; If the network fault type is a hardware-related fault, mark the switch port as disabled and trigger the switching of the standby link.

6. The method according to claim 1, characterized in that After using the pre-trained fault analysis model to perform root cause analysis on the status change event and the multi-dimensional performance data to determine the network fault type, the method further includes: Based on the network fault type, judge the severity of the fault of the network connected to the switch port; According to the severity of the fault, find and execute the corresponding alarm policy in the hierarchical alarm system.

7. The method according to claim 1, characterized in that, The method further includes: Analyze the historical port status information of the switch port to generate a port status report; Generate optimization suggestions for the network topology structure according to the port status report.

8. A fault analysis device, characterized in that, Including: A monitoring module configured to adaptively monitor the port status information of the switch port to determine whether the port status information has changed; A receiving module, configured to receive a status change event sent by the switch port when it is detected that the port status information changes; An obtaining module, configured to obtain multi-dimensional performance data of the network connected to the switch port; A determining module, configured to perform root cause analysis on the status change event and the multi-dimensional performance data by using a pre-trained fault analysis model to determine the network fault type.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.

10. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent switch fault prediction method and system based on deep learning

    CN117834567A

  • Switch communication system with adaptive fault-tolerant capability and implementation method thereof

    CN118784449A

  • Systems and methods for diagnostic, performance and fault management of a network

    US20130232258A1

Cited By

  • Server interface detection method and system, server, equipment, medium and product

    CN120670243A

  • Server interface detection method, system, server, device, medium and product

    CN120670243B

  • Fault diagnosis method of HBA card and electronic equipment

    CN120973587A

  • A fault diagnosis method of an HBA card and an electronic device

    CN120973587B

  • Fault root cause positioning method and system based on automatic analysis

    CN121037200A