Fault analysis method and device, electronic equipment, computer readable storage medium and computer program product
By acquiring topology association information between switches and service devices through a fault analysis system, the inefficiency of switch and server anomaly analysis in high-performance computing scenarios is solved, enabling rapid and accurate fault location and optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-01-10
- Publication Date
- 2026-07-10
AI Technical Summary
In high-performance computing scenarios, existing technologies struggle to quickly and accurately analyze the causes of anomalies in switches and servers, leading to network traffic bottlenecks and low operational efficiency.
The fault analysis system obtains the topology association information between the switch and the service device, quickly locates the abnormal device based on the alarm information, and performs synchronous information analysis to obtain the fault analysis results.
It enables rapid and accurate analysis of abnormal results from switches and service devices, improving operational efficiency, quickly locating fault points, and optimizing network performance.
Smart Images

Figure CN122372392A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network technology, and in particular to a fault analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In high-performance computing scenarios, computation and network communication are two crucial components, and the quality of servers and switches has a significant impact on the availability of the high-performance environment. Servers and switches have complex structures. For example, each server has eight network interface cards (NICs), each with two network ports, resulting in 16 cables connecting a server to a switch, totaling 16 ports. If a switch port malfunctions, it can lead to problems such as training failing to run, training traffic being halved, and training traffic frequently being forwarded through the core switch, causing switch traffic bottlenecks. Therefore, a method that can quickly and accurately analyze the causes of switch and server malfunctions is urgently needed. Summary of the Invention
[0003] This application provides a fault analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can quickly and accurately analyze abnormal results of switches and service equipment.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] This application provides a fault analysis method applied to a fault analysis system, which includes a switch cluster and a service device cluster. The method includes: in response to receiving a first alarm message sent by an abnormal switch in the switch cluster, obtaining topology association information between the switches in the switch cluster and the service devices in the service device cluster; based on the first alarm message and the topology association information, determining an abnormal service device associated with the abnormal switch from the service device cluster; obtaining a second alarm message of the abnormal service device; and performing information synchronization analysis on the first alarm message and the second alarm message to obtain the fault analysis result of the fault analysis system.
[0006] This application provides a fault analysis device applied to a fault analysis system, the fault analysis system including a switch cluster and a service device cluster; the fault analysis device includes: a first information acquisition module, used to acquire topology association information between the switches in the switch cluster and the service devices in the service device cluster in response to receiving a first alarm information sent by an abnormal switch in the switch cluster; an anomaly determination module, used to determine an abnormal service device associated with the abnormal switch from the service device cluster based on the first alarm information and the topology association information; a second information acquisition module, used to acquire a second alarm information of the abnormal service device; and an information analysis module, used to perform synchronous information analysis on the first alarm information and the second alarm information to obtain the fault analysis result of the fault analysis system.
[0007] This application provides an electronic device, which includes: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the fault analysis method provided in this application.
[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the fault analysis method provided in this application.
[0009] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the fault analysis method provided in this application.
[0010] The embodiments of this application have the following beneficial effects:
[0011] Fault analysis methods can be applied to fault analysis systems, which can include switch clusters and service device clusters. This allows switches and service devices to be managed by the same system, improving operational efficiency. First, when the fault analysis system receives the first alarm information from an abnormal switch, it can obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm information and the topology association information, it identifies the abnormal service device associated with the abnormal switch from the service device cluster. This allows for rapid location of the affected abnormal service device even when the switch is malfunctioning. Next, it obtains the second alarm information from the abnormal service device and performs synchronous analysis of the first and second alarm information to obtain the fault analysis results. Since the affected abnormal service device has been quickly located, the service device side can quickly perceive the switch's anomaly. Furthermore, by combining the first and second alarm information for analysis, the abnormal results of the switch and service device can be quickly and accurately analyzed. Attached Figure Description
[0012] Figure 1 This is an optional architecture diagram of the fault analysis system provided in the embodiments of this application;
[0013] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0014] Figure 3 This is an optional flowchart illustrating the fault analysis method provided in the embodiments of this application;
[0015] Figure 4 This is another optional flowchart illustrating the fault analysis method provided in the embodiments of this application;
[0016] Figure 5 This is a schematic diagram of the implementation process for determining abnormal service devices provided in an embodiment of this application;
[0017] Figure 6 This is a schematic diagram of the implementation process for obtaining the second alarm information provided in an embodiment of this application;
[0018] Figure 7 This is another optional flowchart of the fault analysis method provided in the embodiments of this application;
[0019] Figure 8 This is a schematic diagram of a port connection between a server and a switch provided in an embodiment of this application;
[0020] Figure 9 This is a schematic diagram of another port connection between the server and the switch provided in an embodiment of this application;
[0021] Figure 10 This is a schematic diagram of an optional implementation flow of the fault analysis method provided in the embodiments of this application;
[0022] Figure 11 This is a schematic diagram of another optional implementation process of the fault analysis method provided in the embodiments of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0025] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0026] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0027] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0028] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0029] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0030] 1) Responding to: used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0031] 2) Human-computer interaction interface: The interface used to provide human-computer interaction functions / the interface to display fault analysis results.
[0032] 3) Topology diagram: A graphical representation used to show the physical or logical connections between network devices (such as switches, routers, servers, etc.). For example, it can be used to show the physical connections between a switch cluster and a service device cluster.
[0033] 4) Topology association information: This refers to specific data describing the connection relationship between devices, including device name, interface, IP address, protocol, etc.
[0034] 5) Message Queue (MQ): A communication mechanism used to pass messages between applications or system components. Message queues allow senders (producers) to put messages into a queue, and receivers (consumers) to retrieve and process these messages. The main purpose of message queues is to decouple producers and consumers, allowing them to operate independently, thereby improving the system's scalability, reliability, and flexibility. Messages are data units passed in the queue; the message queue is a buffer for storing messages, and messages are processed in a first-in, first-out (FIFO) order.
[0035] Related technologies typically employ a method of periodically fetching data from all switches to obtain relevant alarm information. Then, the server quality is assessed based on the switch anomalies. However, fetching data via a scheduled task is inefficient because there are many data centers and a large number of switch ports, resulting in low fetching efficiency. The time taken for a single fetch is unstable, and in the event of a fetch failure, manual intervention is required to replenish the data, which is time-consuming. Otherwise, data loss may occur, leading to missing information and affecting the accuracy of fault diagnosis.
[0036] Based on at least one of the aforementioned problems with methods in related technologies, embodiments of this application provide a fault analysis method, apparatus, electronic device, computer-readable storage medium, and computer program product, capable of quickly and accurately analyzing abnormal results of switches and service devices. Specifically, the fault analysis method provided in this application embodiment can be applied to a fault analysis system, which may include a switch cluster and a service device cluster, enabling switches and service devices to be managed by the same system, improving operational efficiency. When the fault analysis system receives a first alarm message sent by an abnormal switch, it can obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm message and the topology association information, it determines the abnormal service device associated with the abnormal switch from the service device cluster. Thus, in the event of a switch malfunction, the affected abnormal service device can be quickly located. Subsequently, a second alarm message of the abnormal service device is obtained. Finally, the first alarm message and the second alarm message are analyzed synchronously to obtain the fault analysis result. Since the affected abnormal service device has been quickly located, the service device side can quickly perceive the switch malfunction, and by combining the first alarm message and the second alarm message for analysis, the abnormal results of the switch and service device can be quickly and accurately analyzed.
[0037] Here, we first describe an exemplary application of the fault analysis device in this application embodiment. This fault analysis device is an electronic device used to implement a fault analysis method. In one implementation, the fault analysis device (i.e., electronic device) provided in this application embodiment can be implemented as a terminal or as a server. In one implementation, the fault analysis device provided in this application embodiment can be implemented as any terminal with data processing and fault analysis functions, such as a laptop, tablet, desktop computer, mobile phone, portable music player, personal digital assistant, dedicated messaging device, portable gaming device, intelligent robot, smart home appliance, and smart vehicle device. In another implementation, the fault analysis device provided in this application embodiment can also be implemented as a server, wherein the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application embodiment. The following will illustrate an exemplary application of the fault analysis device as a server.
[0038] See Figure 1 , Figure 1 This is an optional architecture diagram of the fault analysis system provided in this application embodiment, designed to support an application that can quickly and accurately analyze the abnormal results of switches and service devices. The fault analysis system 10 includes at least a terminal 100, a network 200, and a server 300. A fault analysis application is installed on the terminal 100, which can provide fault analysis functions.
[0039] The following description uses the installation of a fault analysis application on a terminal as an example. In this embodiment, server 300 can be a server for the fault analysis application. Server 300 can constitute the fault analysis device of this embodiment, that is, the fault analysis method of this embodiment is implemented through server 300. Terminal 100 connects to server 300 through network 200, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0040] It should be noted that the server 300 in this embodiment can be the system server of a fault analysis system. This system server can be an electronic device in the fault analysis system that is different from the switches in the switch cluster and the service devices in the service device cluster. Of course, the system server can also be implemented as any one or more service devices in the service device cluster.
[0041] See Figure 1 When a user wants to perform fault analysis on a switch cluster and service device cluster, they can input a fault analysis operation in the client of the fault analysis application. This operation is used to select the switch identifier of the switch cluster to be analyzed and the device identifier of the service device cluster. Then, the terminal 100 encapsulates the switch identifier of the switch cluster to be analyzed and the device identifier of the service device cluster into a fault analysis request and sends the fault analysis request to the server 300 through the network 200. Subsequently, the server 300 responds to the fault analysis request, receives the first alarm information sent by the abnormal switch in the switch cluster, and obtains the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm information and the topology association information, the server 300 determines the abnormal service device associated with the abnormal switch from the service device cluster. Next, the server 300 obtains the second alarm information of the abnormal service device. Finally, the server 300 performs information synchronization analysis on the first alarm information and the second alarm information to obtain the fault analysis result of the fault analysis system. After obtaining the fault analysis results, the server 300 sends the fault analysis results to the terminal 100, which can then display the fault analysis results on the current interface of the fault analysis application.
[0042] In some embodiments, the implementation steps of the fault analysis method can also be executed by the terminal 100. That is, after determining the switch cluster and service device cluster to be analyzed, the terminal 100, in response to receiving the first alarm information sent by the abnormal switch in the switch cluster, obtains the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm information and the topology association information, the terminal 100 determines the abnormal service device associated with the abnormal switch from the service device cluster. Then, the terminal 100 obtains the second alarm information of the abnormal service device. Finally, the terminal 100 performs information synchronization analysis on the first alarm information and the second alarm information to obtain the fault analysis result of the fault analysis system.
[0043] The fault analysis method provided in this application embodiment can also be implemented based on a cloud platform and through cloud technology. For example, the server 300 mentioned above can be a cloud server. The cloud server, in response to receiving a first alarm message from an abnormal switch in the switch cluster, obtains the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Alternatively, the cloud server can determine the abnormal service device associated with the abnormal switch from the service device cluster based on the first alarm message and the topology association information. Alternatively, the cloud server can obtain the second alarm message of the abnormal service device. Alternatively, the cloud server can perform synchronous analysis of the first and second alarm messages to obtain the fault analysis results of the fault analysis system.
[0044] In some embodiments, a cloud storage device may also be included, where fault analysis results and other information can be stored. In this way, when performing the next fault analysis, the fault analysis results at the current moment can be retrieved from the cloud storage device as reference information for the next fault analysis.
[0045] It's important to clarify that cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied in the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can be achieved through cloud computing.
[0046] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiment of this application. The electronic device can be a fault analysis device, that is, the electronic device can be implemented as the terminal 100 mentioned above, or as the server 300 mentioned above. Figure 2 The illustrated electronic device 400 includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.
[0047] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0048] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0049] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0050] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0051] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0052] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; network communication module 452 for reaching other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.; presentation module 453 for enabling the presentation of information (e.g., user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with user interface 430 (e.g., display screen, speaker, etc.); input processing module 454 for detecting and translating one or more user inputs or interactions from one or more input devices 432.
[0053] In some embodiments, the fault analysis apparatus provided in this application can be implemented in software. Figure 2 A fault analysis device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a first information acquisition module 4551, an anomaly determination module 4552, a second information acquisition module 4553, and an information analysis module 4554. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0054] In other embodiments, the fault analysis device provided in this application can be implemented in hardware. As an example, the fault analysis device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the fault analysis method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0055] The fault analysis methods provided in the embodiments of this application can be executed by an electronic device, which can be a server or a terminal. That is, the fault analysis methods in the embodiments of this application can be executed by a server, by a terminal, or by interaction between a server and a terminal.
[0056] Figure 3 This is an optional flowchart illustrating the fault analysis method provided in this application embodiment. The following will be combined with... Figure 3 The steps shown are explained as follows: Figure 3 As shown, taking a server as the execution subject of the fault analysis method as an example, the embodiments of this application can be applied to a fault analysis system, which may include a switch cluster and a service device cluster. The method includes the following steps S101 to S104:
[0057] Step S101: In response to receiving the first alarm information sent by the abnormal switch in the switch cluster, obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster.
[0058] In this embodiment of the application, when the fault analysis system receives the first alarm information sent by the abnormal switch in the switch cluster, it can obtain the topology association information between each switch in the switch cluster and each service device in the service device cluster.
[0059] A fault analysis system is a system that detects, analyzes, and handles network faults, typically possessing automated detection, alarm, and fault location capabilities. A switch cluster combines multiple physical switches using specific technologies or protocols, managing them as a single logical device. Switch clusters can improve network manageability, reliability, and performance. Each switch in a switch cluster can be configured through a single management interface, eliminating the need for individual switch management; switches can share resources (such as routing tables), thereby improving resource utilization. If a switch in the cluster fails, other switches can take over its operation, ensuring network continuity.
[0060] A service device cluster refers to combining multiple physical servers or virtual servers using specific technologies or protocols to form a logically unified pool of computing resources. Service device clusters can improve system reliability, performance, and scalability. An abnormal switch refers to one or more switches within a switch cluster that have failed. As critical network devices, switches may experience various hardware or software failures. The following are possible switch failure types: Power failure: For example, the switch fails to start or suddenly loses power, causing network interruption. This could be due to a damaged power module, loose power cord, or unstable power supply. Fan failure: Overheating of the switch may cause it to automatically shut down or cause hardware damage. This could be due to a damaged fan or dust accumulation leading to poor heat dissipation. Port failure: A specific port becomes unusable, preventing connected devices from communicating. This could be due to physical damage to the port (e.g., loose interface, bent pins) or electronic component failure. Backplane failure: Internal communication of the switch is interrupted. This could be due to damaged or aging backplane circuitry. Operating system crash: The switch cannot operate normally and requires a restart or system reload. This could be due to software bugs, configuration errors, or memory leaks. Configuration errors: Abnormal network functions may be due to human error or corrupted configuration files. Firmware issues: Degraded switch performance or abnormal functions may be due to outdated firmware or bugs. Bandwidth bottleneck: Increased network latency and slower data transmission speeds may be due to insufficient port bandwidth or excessive traffic. Broadcast storm: Network congestion, reduced switch performance, or even network paralysis may occur. This could be due to network loops or malicious attacks. Media Access Control (MAC) MAC address table overflow: The switch cannot forward data packets correctly, leading to communication failure. This may be due to too many devices in the network or insufficient MAC address table capacity. Physical link failure: The switch cannot communicate, which may be due to a damaged network cable, a broken fiber optic cable, or a loose interface. Duplex mode mismatch: Network performance degrades, and data packets are lost. This may be due to inconsistent duplex mode settings between the switch port and the connected devices. Network attack: Network performance degrades, data leakage, or service interruption. This may be due to the switch being attacked by Distributed Denial of Service (DDoS), Address Resolution Protocol (ARP) spoofing, or MAC address flooding attacks. The first alarm message can indicate that there is an anomaly in the switch. The first alarm message can include several key fields, which can be used to characterize the switch name, status, cause of failure, time of failure, and type of failure.
[0061] Each switch in the switch cluster can send log information to the fault analysis system in real time. The fault analysis system analyzes the logs to determine if there are any anomalies or faults in the switches. When a switch in the switch cluster experiences an anomaly (such as hardware failure, link interruption, performance degradation, etc.), the abnormal switch will send an initial alarm message to the fault analysis system. After receiving the initial alarm message, the fault analysis system will obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Alternatively, each switch in the switch cluster can analyze its own log information in real time to determine if there are any anomalies or faults in the switch. When a switch in the switch cluster experiences an anomaly (such as hardware failure, link interruption, performance degradation, etc.), the abnormal switch will send an initial alarm message to the fault analysis system. This allows the switches to proactively push alarm information to the fault analysis system, enabling the fault analysis system to detect switch anomalies in real time, thereby improving the efficiency of fault analysis. Topology association information refers to the network topology information of the connection relationships between each switch in the switch cluster and each service device in the service device cluster. For example, topology association information may include the connection method between switches and service devices, and the location of switches and service devices in the network (such as access layer, aggregation layer, core layer).
[0062] Step S102: Based on the first alarm information and topology association information, determine the abnormal service device associated with the abnormal switch from the service device cluster.
[0063] In this embodiment, the fault analysis system can identify the abnormal switch based on the first alarm information, and then, based on the topology association information, identify the service devices (such as servers, cloud servers, etc.) in the service device cluster that have topology association information with the abnormal switch. These service devices are then identified as abnormal service devices. When a switch malfunctions, it can cause various problems to the service devices connected to the switch.
[0064] The following are service device failure types caused by switch malfunctions: Network Connection Interruption: When a switch port fails, a link is broken, or the switch completely crashes, the service device may be unable to communicate with the external network, resulting in service interruption. Network Performance Degradation: When the switch has insufficient bandwidth, experiences broadcast storms, or a MAC address table overflow, the service device may experience increased network latency and slower data transmission speeds. Internet Protocol (IP) Address Conflict Failure: Incorrect switch configuration may cause the service device's IP address to conflict with other devices, leading to abnormal network communication. Link Aggregation Failure: Incorrect switch link aggregation configuration or inconsistent port states can cause the high-bandwidth link between the service device and the switch to fail, resulting in network performance degradation. Broadcast Storm Failure: A network loop or malicious attack (such as ARP spoofing) on the switch can overwhelm the service device's network interface with a large number of broadcast packets, causing a surge in the service device's Central Processing Unit (CPU) usage. Security Risk Failure: When the switch is under network attack, the service device's network interface can be overwhelmed by attack traffic, leading to service interruption or data leakage. Storage Access Failure: Switch malfunctions can cause communication interruptions between the service device and storage devices, preventing the service device from accessing network storage and resulting in data being unreadable or unwriteable.
[0065] Step S103: Obtain the second alarm information of the abnormal service device.
[0066] In this embodiment of the application, the second alarm information of the abnormal service device may include multiple key fields, which can be used to characterize the name, status, cause of failure, time of failure, and type of failure of the abnormal service device.
[0067] Step S104: Perform information synchronization analysis on the first alarm information and the second alarm information to obtain the fault analysis results of the fault analysis system.
[0068] In this embodiment, the fault analysis system can perform information synchronization analysis on the first alarm information sent by the abnormal switch and the second alarm information of the abnormal service device. It is necessary to determine that the timestamp information of the first alarm information and the second alarm information for information synchronization analysis are consistent, that is, the time when the switch and the service device fail is consistent. In this way, the first alarm information and the second alarm information can be accurately analyzed to obtain the fault analysis result.
[0069] For example, if the first alarm message from the switch occurs at "2023-10-01 10:00:00", indicating that the traffic on port 1 has reached 95% and the packet loss rate has increased, and the second alarm message from the service device occurs at "2023-10-01 10:05:00", indicating that the service device's response time exceeds 5 seconds and the CPU utilization reaches 90%, the alarms from the switch and the service device occur 5 minutes apart. If this time difference is not noticed, it might be mistakenly assumed that the two alarms occurred simultaneously, and that the high traffic on port 1 of the switch caused the performance degradation of the service device. However, the service device alarm occurring 5 minutes after the switch alarm may not be directly caused by the high traffic on the switch. The service device alarm could be caused by other reasons, such as a performance bottleneck within the service device's application (e.g., slow database queries) or a sudden surge in request pressure (e.g., a surge in user access). Therefore, the inconsistency in timing can lead to incorrect fault analysis results.
[0070] When the timestamps of the first and second alarm messages match, the fault analysis system can perform synchronized analysis to obtain fault analysis results. For example, if the switch port traffic is abnormal and the service device's network interface traffic is also abnormal, it may be due to network congestion. Network congestion slows down the service device's request processing, requiring optimization of network traffic or increased bandwidth. If the switch port is disconnected and the service device cannot access the external network, it may be a physical link failure, requiring checking the network cable or switch port. If the switch port traffic is abnormally high, and the service device's CPU utilization is too high, resulting in service response timeouts, it may be a DDoS attack, requiring the implementation of protective measures. The fault analysis system can quickly locate the fault point based on the comprehensive analysis results. The fault analysis system can also archive the first and second alarm messages to provide optimization suggestions based on historical alarm data (i.e., archived first and second alarm messages). For example, if the service device frequently experiences network congestion, it can suggest upgrading network equipment or optimizing traffic scheduling. If the service device has insufficient resources, it can suggest expanding capacity or optimizing application performance.
[0071] The fault analysis method provided in this application embodiment can be applied to a fault analysis system, which may include a switch cluster and a service device cluster. This allows switches and service devices to be managed by the same system, improving operational efficiency. When the fault analysis system receives a first alarm message from an abnormal switch, it can obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm message and the topology association information, it determines the abnormal service device associated with the abnormal switch from the service device cluster. In this way, the affected abnormal service device can be quickly located when the switch is abnormal. After that, the second alarm message of the abnormal service device is obtained. Finally, the first alarm message and the second alarm message are analyzed synchronously to obtain the fault analysis result. Since the affected abnormal service device has been quickly located, the service device side can quickly perceive the switch abnormality. By combining the first alarm message and the second alarm message for analysis, the abnormal results of the switch and service device can be quickly and accurately analyzed.
[0072] The following examples illustrate the application scenarios of the fault analysis method provided in the embodiments of this application.
[0073] Servers and switches are core components of modern data centers and network infrastructure. Their practical applications are widespread, covering areas from enterprise offices to cloud computing and big data processing. Here are some specific application scenarios: Enterprise Networks: As businesses grow, the number of servers in data centers increases. Switches act as the key link connecting these servers. Within an enterprise network, servers host critical business applications such as enterprise resource planning and customer relationship management, while switches connect computers, printers, and other network devices across departments, ensuring efficient data transmission. Data Centers: Within data centers, servers host numerous virtual machines and containers, providing computing resources, storage space, and network services. Switches form the network backbone of the data center, supporting data exchange between servers and access to external networks through high-speed connections. Cloud Computing: In cloud computing environments, server clusters provide elastic computing resources, supporting multi-tenant virtualization services. Switches, through software-defined networking technology, enable dynamic allocation and management of network resources to meet the needs of different users. Content Delivery Networks (CDNs): CDN service providers use servers to store and distribute content such as videos, images, and web pages to reduce latency and improve user experience. Switches ensure the rapid transmission of content from the source server to edge nodes and then to end users. Big Data Analytics: Big data analytics platforms rely on server clusters to process and analyze massive datasets. Switches provide high-speed network connectivity, supporting the rapid flow of data between cluster nodes and integration with other systems. High-Performance Computing: In fields such as scientific research, engineering simulation, and financial modeling, high-performance computing clusters require servers to provide powerful computing capabilities. Switches ensure efficient communication between computing nodes through low-latency, high-bandwidth network connections.
[0074] The fault analysis method of this application embodiment will be described below in conjunction with the above scenario.
[0075] Figure 4 This is another optional flowchart illustrating the fault analysis method provided in the embodiments of this application, such as... Figure 4 As shown, the method includes the following steps S201 to S213:
[0076] Step S201: The terminal receives the fault analysis operation input by the user.
[0077] Fault analysis operations include selection operations or input operations. The selection operation is used to select the switch cluster and service device cluster to be analyzed, or the input operation is used to input the switch identifier of the switch cluster to be analyzed and the device identifier of the service device cluster.
[0078] In step S202, the terminal encapsulates the switch identifier of the switch cluster to be analyzed and the device identifier of the service device cluster into the fault analysis request.
[0079] The fault analysis request is used to request the server to perform fault analysis on the switch cluster and service device cluster to be analyzed.
[0080] Step S203: The terminal sends a fault analysis request to the server.
[0081] In this embodiment, the terminal sends a fault analysis request to the server to request the server to perform fault analysis on the switch cluster and service device cluster to be analyzed. Of course, in some embodiments, the server can also actively perform fault analysis on the switch cluster and service device cluster to be analyzed. That is, there can be a preset number of switch clusters and service device clusters to be analyzed, and the server can periodically or non-periodically perform fault analysis on these switch clusters and service device clusters to be analyzed.
[0082] In step S204, the server responds to the fault analysis request by collecting log information from each switch in the switch cluster in real time.
[0083] In this embodiment, the switch generates a large amount of log information (Syslog) during operation, such as interface status changes, hardware failures, and performance metrics. This log information can be sent to the server via the Syslog protocol. Each switch in the switch cluster can upload its own log information to the server in real time. For example, a data collection agent can be deployed on the switch. The data collection agent can collect the switch's Syslog logs in real time. The server can then collect the log information from each switch in the switch cluster in real time. Log information may include: timestamp: indicating the specific date and time the log occurred; switch name: indicating the name of the switch that generated the log; log level: such as general operation information (info), warnings, and errors; system events: such as records of switch startup or restart and changes in switch configuration; interface status: such as enabling, disabling, or status changes of switch interfaces (e.g., Ethernet ports) and packet loss or conflicts; hardware status: such as records of hardware failures or anomalies, and hardware failure recovery, such as power supplies, fans, and temperature sensors; and performance metrics: CPU utilization, memory utilization, and bandwidth utilization. Network traffic: For example, interface traffic statistics (such as input / output byte count, packet count). Software events: For example, records of switch operating system or software failures or errors, and switch software upgrades.
[0084] In step S205, the server parses the log information to obtain the real-time parameters of each switch.
[0085] In this embodiment, the server parses the collected log information to obtain real-time parameters for each switch. These real-time parameters may include: CPU utilization, memory usage, interface status, traffic statistics (including the number of received and sent data packets, bytes, and error packets), error and dropped packets, temperature, power status, log level, routing table, and ARP table.
[0086] In step S206, the server responds to the fact that the real-time parameters of any switch meet the preset abnormal conditions and determines that any switch is an abnormal switch.
[0087] In this embodiment, the user can pre-set abnormal conditions for determining whether a switch is abnormal. When the real-time parameters of any switch meet the preset abnormal conditions, the switch can be identified as an abnormal switch. For example, abnormal conditions may include: CPU utilization consistently exceeding 80%, memory utilization remaining above 90%, port traffic suddenly increasing to several times the normal level, port error packets or dropped packets suddenly increasing to several times the normal level, temperature exceeding a safety threshold, unstable output voltage, fan stopping or abnormal speed, etc. For example, if switch A's CPU utilization is 90% from "2023-10-01 10:00:00 to 2023-10-01 10:30:00", and switch A's CPU utilization exceeds 80% for more than 30 minutes, and switch A's real-time parameters meet the preset abnormal conditions, switch A can be identified as an abnormal switch. The log information of switch A may include: timestamps: "2023-10-01 10:00:00 to 2023-10-01 10:30:00". Switch Name: Switch A. Log Level: Warning. System Event: CPU utilization exceeds 80%.
[0088] In step S207, the server identifies the log information of the abnormal switch as the first alarm information.
[0089] In this embodiment, the server can identify the log information of the malfunctioning switch as the first alarm message. For example, the log information of switch A may include: timestamp: "2023-10-01 10:00:00-2023-10-01 10:30:00"; switch name: switch A; log level: warning; system event: CPU utilization exceeds 80%. The log information of switch A can be identified as the first alarm message.
[0090] In step S208, the server responds to the first alarm information sent by the abnormal switch in the switch cluster by obtaining the topology association information between the switches in the switch cluster and the service devices in the service device cluster.
[0091] In this embodiment of the application, when the server receives the first alarm information sent by the abnormal switch in the switch cluster, it can obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster.
[0092] In some embodiments, obtaining the topology association information between switches in a switch cluster and service devices in a service device cluster can be achieved in the following way: First, traverse the fault analysis system within a first preset time period to obtain the switch identifier of each switch and the device identifier of each service device in the fault analysis system within the first preset time period, as well as the device identifier of at least one service device connected to each switch and the switch identifier of each switch connected to each service device; then, based on the device identifier of at least one service device connected to each switch and the switch identifier of each switch connected to each service device, determine the connection relationship between the switch and the service device; then, based on the connection relationship, draw a topology relationship diagram between the switch cluster and the service device cluster in the fault analysis system; finally, generate topology association information based on the topology relationship diagram.
[0093] In this embodiment, a switch identifier is information used to identify a switch, which may include the switch's name, serial number, MAC address, IP address, etc. Each switch has a unique switch identifier. A device identifier is information used to identify a service device, which may include the service device's name, serial number, MAC address, IP address, etc. Each service device has a unique device identifier. For example, a fault analysis system may have three switches (switch A, switch B, and switch C) and six service devices (service device 1, service device 2, service device 3, service device 4, service device 5, and service device 6). Switch A is connected to service devices 1, 2, 3, and 6; switch B is connected to service devices 1, 4, and 6; and switch C is connected to service devices 4, 5, and 6. The user can preset a time period (i.e., a first preset time period), such as 5 minutes.
[0094] The server can iterate through the fault analysis system every five minutes to obtain the switch identifier (A, B, C) of each switch and the device identifier (1, 2, 3, 4, 5, 6) of each service device in the fault analysis system. It also obtains the device identifier of at least one service device connected to each switch (e.g., switch A (1, 2, 3, 6); switch B (1, 4, 6); switch C (4, 5, 6)) and the switch identifier of the switch connected to each service device (e.g., service device 1 (A, B); service device 2 (A); service device 3 (A); service device 4 (B, C); service device 5 (C); service device 6 (A, B, C)). Then, based on the device identifiers of at least one service device connected to each switch (e.g., switch A (1, 2, 3, 6); switch B (1, 4, 6); switch C (4, 5, 6)) and the switch identifiers of the switches connected to each service device (e.g., service device 1 (A, B); service device 2 (A); service device 3 (A); service device 4 (B, C); service device 5 (C); service device 6 (A, B, C)), the connection relationships between switches and service devices can be determined (i.e., switch A is connected to service devices 1, 2, 3, and 6 respectively; switch B is connected to service devices 1, 4, and 6 respectively; switch C is connected to service devices 4, 5, and 6 respectively). Furthermore, based on these connection relationships, a topology diagram of the switch cluster and service device cluster in the fault analysis system can be drawn, demonstrating the physical connection relationships between the switch cluster and the service device cluster. Finally, topology association information is generated based on the topology diagram. For example, switch A connects to service device 1 via port 1, service device 2 via port 2, service device 3 via port 3, and service device 6 via port 4. Switch B connects to service device 1 via port 1, service device 4 via port 2, and service device 6 via port 3. Switch C connects to service device 4 via port 1, service device 5 via port 2, and service device 6 via port 3.
[0095] The above implementation method allows for the periodic updating of topology association information between switches in a switch cluster and service devices in a service device cluster within a preset time period, thereby improving the accuracy and timeliness of the topology association information.
[0096] In step S209, the server determines the abnormal service device associated with the abnormal switch from the service device cluster based on the first alarm information and topology association information.
[0097] In some embodiments, see Figure 5 , Figure 5Step S209 can be achieved through the following steps S2091 to S2093:
[0098] Step S2091: Based on the first alarm information, determine the switch identifier of the abnormal switch.
[0099] In some embodiments, step S2091 can be implemented in the following manner: First, a first message is obtained from a pre-generated first message queue, wherein the first message is generated based on the switch identifier of the abnormal switch and the first alarm information, and the first message queue includes the first message; finally, the first message is parsed to obtain the switch identifier of the abnormal switch and the first alarm information.
[0100] In this embodiment, the first message is generated based on the switch identifier of the abnormal switch and the first alarm information. If there are multiple abnormal switches, multiple first messages are generated sequentially based on the switch identifier of each abnormal switch and its respective first alarm information. These multiple first messages are sequentially placed into a first message queue and processed in a first-in, first-out (FIFO) order. First, the first message that entered the first message queue is retrieved from the pre-generated first message queue. Then, the first message is parsed to obtain the switch identifier of the abnormal switch and the first alarm information. For example, the first alarm information for abnormal switch A may include: timestamp: "2023-10-01 10:00:00-2023-10-01 10:30:00"; switch name: switch A; log level: warning; system event: CPU utilization exceeds 80%. The first message can be generated based on the switch identifier "A" of abnormal switch A and the first alarm information of abnormal switch A. Then, the first message is placed into the first message queue and processed in a first-in-first-out order. When the first message corresponding to the abnormal switch A is processed, the first message is parsed to obtain the switch identifier "A" of the abnormal switch A and the first alarm information.
[0101] By implementing the above method, a message can be generated for the switch when an anomaly occurs, so that the anomaly can be detected and handled in a timely manner, thereby improving the efficiency of anomaly handling.
[0102] Step S2092: Based on the switch identifier and topology association information, determine the local topology association information between the abnormal switch and the service devices in the service device cluster.
[0103] In this embodiment, a topology diagram can be determined based on topology association information. Based on the switch identifier of the abnormal switch and the topology diagram, a local topology diagram between the abnormal switch and the service device cluster can be drawn. Based on the local topology diagram, local topology association information between the abnormal switch and the service devices in the service device cluster can be generated. For example, the topology association information could be that switch A connects to service device 1 via port 1, service device 2 via port 2, service device 3 via port 3, and service device 6 via port 4. Switch B connects to service device 1 via port 1, service device 4 via port 2, and service device 6 via port 3. Switch C connects to service device 4 via port 1, service device 5 via port 2, and service device 6 via port 3. Based on the aforementioned topology association information, the topology diagram can be determined by removing service devices not connected to the abnormal switch from the topology diagram, such as switch A, to draw a local topology diagram between the abnormal switch and the service device cluster. Alternatively, service devices connected to the abnormal switch can be selected from the topology diagram based on the switch identifier of the abnormal switch to draw a local topology diagram between the abnormal switch and the service device cluster. Finally, local topology association information is generated based on the local topology relationship diagram. For example, the local topology association information can be that switch A connects to service device 1 through port 1, service device 2 through port 2, service device 3 through port 3, and service device 6 through port 4.
[0104] Step S2093: Based on the local topology association information, identify the abnormal service devices associated with the abnormal switch from the service device cluster.
[0105] In this embodiment, based on local topology association information, abnormal service devices associated with abnormal switches can be identified from the service device cluster. For example, the local topology association information could be that switch A connects to service device 1 via port 1, service device 2 via port 2, service device 3 via port 3, and service device 6 via port 4. Therefore, the abnormal service devices associated with switch A can be identified as service device 1, service device 2, service device 3, and service device 6.
[0106] Through steps S2091 to S2093, based on the switch identifier of the abnormal switch and combined with the topology association information, the service device associated with the abnormal switch can be quickly located, thereby quickly identifying the abnormal service device and improving the efficiency of troubleshooting.
[0107] Step S210: The server obtains the second alarm information of the abnormal service device.
[0108] In some embodiments, see Figure 6 , Figure 6 Step S210 can be achieved through the following steps S2101 to S2104:
[0109] Step S2101: Retrieve the second message from the pre-generated second message queue.
[0110] Here, the second message is generated based on the device identifier of the abnormal service device, and the second message queue includes the second message.
[0111] In this embodiment, the second message is generated based on the device identifier of the abnormal service device. If the abnormal service device is associated with multiple abnormal switches, multiple second messages are generated sequentially based on the device identifier of the abnormal service device according to the different abnormal switches. These multiple second messages are sequentially placed into a second message queue and processed in a first-in, first-out (FIFO) order. The second message that first entered the second message queue is retrieved from the pre-generated second message queue. For example, the abnormal switches are switch A and switch B, the service devices associated with switch A are service device 1, service device 2, service device 3, and service device 6, and the service devices associated with switch B are service device 1, service device 4, and service device 6. Second messages are generated based on the device identifiers of service device 1, service device 2, service device 3, and service device 6, and the device identifiers of service device 1, service device 4, and service device 6, respectively.
[0112] Step S2102: Parse the second message to obtain the device identifier of the abnormal service device.
[0113] In this embodiment of the application, parsing the second message can yield the device identifier of the abnormal service device associated with the abnormal switch.
[0114] Step S2103: Determine the handling strategy for the abnormal service device based on the device identifier of the abnormal service device.
[0115] In some embodiments, step S2103 can be implemented in the following ways: in response to determining that the abnormal service device is a first type of abnormal service device based on the device identifier of the abnormal service device, the processing strategy for the first type of abnormal service device is determined to be device isolation processing; or, in response to determining that the abnormal service device is a second type of abnormal service device based on the device identifier of the abnormal service device, the processing strategy for the second type of abnormal service device is determined to send a fault reminder message to the information processing account bound to the second type of abnormal service device; wherein, the first type of abnormal service device is a service device that is currently on sale, and the second type of abnormal service device is a service device that has already been sold.
[0116] In this embodiment, the device identifier can also be used to characterize the type of service device. For example, the type of service device may include: service devices currently on sale and service devices already sold. When an abnormal service device is determined to be a service device currently on sale based on its device identifier, device isolation processing can be performed on the abnormal service device. Device isolation processing may include: network isolation (disconnecting network connections, virtual LAN isolation, etc.), service isolation (stopping services, removing load balancers, etc.), resource isolation (resource restrictions, virtual machine or container isolation, etc.), and data isolation (data access restrictions). When an abnormal service device is determined to be a service device already sold based on its device identifier, a fault alert message can be sent to the information processing account bound to the abnormal service device. The fault alert message can be sent via email, telephone, SMS, or other methods.
[0117] Through the above implementation method, targeted handling measures can be taken for different types of abnormal service equipment (in sale or sold). For service equipment in sale, equipment isolation is carried out to prevent the spread of faults and ensure the normal operation of other equipment. For service equipment sold, fault reminder messages are sent to the bound information processing account to notify users in a timely manner. This not only improves the efficiency of fault handling, but also improves the user experience.
[0118] Step S2104: Determine the second alarm information of the abnormal service device based on the device identifier of the abnormal service device.
[0119] In some embodiments, step S2104 can be implemented in the following way: First, based on the device identifier of the abnormal service device, obtain the log information of the abnormal service device within a second preset time period; finally, parse the log information to obtain the second alarm information of the abnormal service device.
[0120] In this embodiment, firstly, based on the device identifier of the abnormal service device associated with the abnormal switch, log information of the abnormal service device within a second preset time period is obtained. The time range of the second preset time period can be determined based on the time when the abnormal switch becomes abnormal. The start time of the second preset time period must be later than or the same as the time when the abnormal switch becomes abnormal. The end time of the second preset time period is not specifically limited and can be selected according to the situation. Finally, the obtained log information is parsed to obtain the second alarm information of the abnormal service device. The second alarm information may include multiple key fields, which can be used to characterize the name, status, cause of failure, time of failure, and type of failure of the abnormal service device.
[0121] Through the above implementation method, the log information of the abnormal service device can be obtained based on the device identifier of the abnormal service device, and the log information can be parsed to obtain the second alarm information. This allows for the accurate location of key information such as the cause, status and occurrence time of the abnormal service device's failure through the second alarm information, thereby improving the efficiency and accuracy of fault diagnosis.
[0122] Through steps S2101 to S2104, by obtaining and parsing the second message from the pre-generated second message queue, the device identifier of the abnormal service device can be quickly determined, and a targeted processing strategy can be adopted for the abnormal service device based on the device identifier. The initial processing can be carried out automatically according to the processing strategy based on the device identifier, which enhances the reliability and stability of the service device cluster, and obtains the second alarm information for subsequent analysis and processing.
[0123] Step S211: The server performs information synchronization analysis on the first alarm information and the second alarm information to obtain the fault analysis results of the fault analysis system.
[0124] It should be noted that step S211 is the same as step S104 above, and the implementation details of step S211 in this embodiment will not be repeated.
[0125] In step S212, the server sends the fault analysis results to the terminal.
[0126] Step S213: The terminal outputs the fault analysis results.
[0127] In this embodiment, log information from the switch cluster is collected in real time. By setting abnormal conditions, abnormal switches are detected in a timely manner, and the topological association information between the switch cluster and the service device cluster is obtained. This allows for the identification of abnormal service devices associated with the abnormal switches, greatly improving the efficiency of network management and fault handling, ensuring network stability and reliability. Furthermore, corresponding processing strategies can be adopted for abnormal service devices to perform preliminary processing, reducing the impact of abnormal service devices on other devices. Additionally, second alarm information of abnormal service devices is obtained for subsequent analysis and processing.
[0128] Figure 7 This is another optional flowchart illustrating the fault analysis method provided in the embodiments of this application, such as... Figure 7 As shown, the method includes the following steps S301 to S307:
[0129] Step S301: The terminal receives a fault analysis request for the abnormal service device sent by the information processing account bound to the second type of abnormal service device.
[0130] The second type of faulty service equipment (i.e., equipment that has already been sold for service) can be linked to an information processing account. Users who have purchased service equipment can upload a fault analysis request for the faulty service equipment through their information processing account when they discover any abnormalities or malfunctions in the equipment. The fault analysis request will then be sent to the terminal.
[0131] In step S302, the terminal encapsulates the device identifier of the second type of abnormal service device to be analyzed into the fault analysis request.
[0132] The terminal can parse the fault analysis request for abnormal service devices to obtain the device identifier of the second type of abnormal service device to be analyzed, and then encapsulate the device identifier of the second type of abnormal service device to be analyzed into the fault analysis request.
[0133] Step S303: The terminal sends a fault analysis request to the server.
[0134] It should be noted that step S303 is the same as step S203 above, and the implementation details of step S303 in this application embodiment will not be repeated.
[0135] Step S304: The server responds to the fault analysis request and obtains the fault analysis results from the fault analysis system.
[0136] In this embodiment of the application, when the server receives a fault analysis request, it can obtain the fault analysis results of the fault analysis system. The fault analysis results may include the current fault analysis results and historical fault analysis results.
[0137] Step S305: The server filters out the fault information of the second type of abnormal service device from the fault analysis results based on the device identifier of the second type of abnormal service device.
[0138] In this embodiment, the server can filter out the fault information of the second type of abnormal service device and the fault information of the abnormal switches associated with the second type of abnormal service device from the fault analysis results based on the device identifier of the second type of abnormal service device. For example, if the second type of abnormal service device is service device 1, and the switches associated with the second type of abnormal service device include switch A, switch B, and switch C, and the abnormal switches are switch B and switch C, then the fault information of switch B and switch C can also be filtered out to facilitate collaborative analysis based on the fault information of the service device and the switches, thereby improving the accuracy of fault analysis.
[0139] Step S306: The server sends fault information to the terminal.
[0140] In step S307, the terminal sends the fault information to the information processing account.
[0141] After receiving fault information, the information processing account can present it to the user. The user can then determine the cause of the service device malfunction based on this information and take appropriate action. For example, if the service device's CPU usage is too high, with details showing CPU usage consistently exceeding 90% for 30 minutes, the user can determine whether it's due to an application consuming excessive resources or a malicious attack causing the high CPU usage. If it's an application issue, the user can try restarting the relevant service or optimizing application configuration. If it's insufficient hardware resources, the user can consider upgrading hardware (such as adding CPU or memory). If it's a network attack, the user can enable firewall rules or contact the security team for assistance. If the problem persists, the user can send a notification to the relevant operations and maintenance personnel.
[0142] In this embodiment of the application, when a fault analysis request for a second type of abnormal service device is received, the fault analysis results of the fault analysis system are obtained, and then the fault information of the second type of abnormal service device is filtered out according to the device identifier of the second type of abnormal service device. This enables comprehensive and accurate analysis of fault information, improving the accuracy and efficiency of fault analysis.
[0143] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0144] In a heterogeneous computing cluster (HCC), sub-machines (computing nodes) are networked using Remote Direct Memory Access (RDMA) technology, such as... Figure 8 As shown, Figure 8 This is a schematic diagram of a port connection between a server and a switch provided in an embodiment of this application. A server has eight Mellanox network interface cards (NICs 817, 818, 819, 820, 821, 822, 823, and 824), each with two network ports. Eight sets of switches are connected to each NIC. There are 16 cables connecting the server to the switches, with a total of 16 ports associated with multiple upstream switches (switches 801, 802, 803, 804, 805, 806, 807, 808, 809, 810, 811, 812, 813, 814, 815, and 816). For example... Figure 9 As shown, Figure 9This is a schematic diagram (i.e., a topology diagram) showing the connection between another port of the server and the switch provided in this application embodiment. The servers under a switch 901 belong to different tenants. For example, the first server 902 and the second server 903 belong to tenant A; the third server 904 is an unsold, idle machine; and the Nth server 905 belongs to tenant B.
[0145] See below Figure 10 , Figure 10 This is a schematic diagram of an optional implementation flow of the fault analysis method provided in this application embodiment. In some embodiments, the electronic device used to implement the fault analysis method may include: a bypass task, a computing controller, a network management sniper, a switch controller, and a fault analysis platform.
[0146] Step S1001: The bypass task retrieves the switch list by computer room module and saves the switch list.
[0147] Bypass tasks refer to tasks that are not processed according to the standard procedure during system operation, but are completed directly through a specific path. Bypass tasks use network management snipe (SNIPER) to pull the list of all switches under a specific data center and save the switch list.
[0148] Step S1002: The bypass task periodically pulls switch logs and alarms, as well as the status of the switch configuration management database.
[0149] The bypass task obtains relevant alarm information, log information, traffic data, etc. from the switch through the Application Programming Interface (API) based on the switch's port information and other information from the switch controller.
[0150] Step S1003: The bypass task pulls the devices under the switch.
[0151] The bypass task network management SNIPER pulls the devices (i.e. servers) under the switch, and then performs a quality assessment of the server based on the corresponding anomalies of the switch.
[0152] Step S1004: The bypass task blacklists unsold devices and prohibits their sale.
[0153] For machines in the resource pool, bypass tasks first blacklist the devices through the computing controller, and then sell them to external parties after the switch side recovers.
[0154] Step S1005: For the sold devices, the bypass task records relevant alarm information and associates it with the server.
[0155] For equipment that has already been sold, the corresponding switch alarm information is pushed to the fault analysis platform for recording and analysis, thereby enabling the screening of abnormal information at a specified time based on server information.
[0156] Figure 10 Because there are many data centers and a large number of switch ports, the data retrieval efficiency is not very high. The time taken for a single retrieval is unstable. In the event of a failed retrieval, manual intervention is required to replenish the data, which is time-consuming. Otherwise, data loss will lead to missing information and affect the accuracy of fault diagnosis.
[0157] Based on this, see Figure 11 , Figure 11 This is a schematic diagram of another optional implementation flow of the fault analysis method provided in this application embodiment. By introducing a message queue, the switch actively pushes alarm information to the message queue. The background component detects the message queue, receives the pushed switch information, filters out irrelevant information, and then forwards it to other message queues for use by the server management layer. This allows for rapid and accurate analysis of abnormal results of switches and service devices. In some embodiments, the electronic device used to implement the fault analysis method may include: a server management layer, a network analysis component, a network management Kafka, a server message queue, a fault analysis platform, and a network management sniper.
[0158] Step S1101: Switch alarm, collect, and send messages.
[0159] The switch analyzes Syslog information in real time by collecting Agent data, and determines whether there is an anomaly in the switch according to the pre-set policy (i.e. the preset anomaly conditions). If there is an anomaly, the alarm information (i.e. the first alarm information) (e.g., device name, fault cause, fault time, fault type, etc.) is placed as a message (i.e. the first message) to the Kafka message queue (i.e. the first message queue).
[0160] Step S1102: The network analysis component retrieves the topology relationship between the server and the switch.
[0161] The network analysis component uses the network management system SNIPER to periodically update the topology association information between the server (i.e., the service device) and the switch.
[0162] Step S1103: The network analysis component consumes the message.
[0163] The network analysis component associates abnormal switches with servers by consuming messages from the message queue, thus identifying servers with abnormal behavior.
[0164] In step S1104, the network analysis component associates the switch alarm with the server and then re-delivers the message.
[0165] The server with the anomaly is treated as a message (i.e., the second message) and placed in the next message queue (i.e., the second message queue) for the server management layer to consume and process.
[0166] In step S1105, the server management layer consumes messages, initiates corresponding operations, and analyzes log records.
[0167] The server management layer consumes messages from the message queue, identifies abnormal servers, and initiates corresponding operations for the abnormal servers (i.e., the handling strategy for abnormal service devices). (Among these, unsold servers are blacklisted (i.e., device isolation), and sold servers are created and customers are notified (i.e., fault alert messages are sent to the information processing accounts bound to the second type of abnormal service devices). The layer also obtains the alarm information list corresponding to the abnormal servers (i.e., the second alarm information).
[0168] Step S1106: Record alarm information on the server switch side associated with the server management layer.
[0169] The server management layer synchronizes the alarm information list corresponding to the abnormal server and the alarm information of the switch associated with the abnormal server to the fault analysis platform for fault analysis.
[0170] In this embodiment, the switch actively pushes positive information, which better decouples the data linkage between the server and the switch side. This significantly improves the integrity and accuracy of system data (i.e., the system with information linkage between the server and the switch, including all servers and switches that require fault analysis). It can quickly associate relevant detection and alarm data from the switch side, scan for potential faults through the system, and provide corresponding processing basis in the process of handling faults. It can also significantly improve the availability of cloud services (i.e., the server may include cloud servers) and provide a high-quality service experience to the outside world.
[0171] The following description continues to illustrate the exemplary structure of the fault analysis device 455 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the fault analysis device 455 in the memory 450 may include:
[0172] The first information acquisition module 4551 is used to acquire topology association information between the switches in the switch cluster and the service devices in the service device cluster in response to receiving the first alarm information sent by the abnormal switch in the switch cluster; the anomaly determination module 4552 is used to determine the abnormal service device associated with the abnormal switch from the service device cluster based on the first alarm information and the topology association information; the second information acquisition module 4553 is used to acquire the second alarm information of the abnormal service device; and the information analysis module 4554 is used to perform synchronous information analysis on the first alarm information and the second alarm information to obtain the fault analysis result of the fault analysis system.
[0173] In some embodiments, the first information acquisition module 4551 is further configured to collect log information of each switch in the switch cluster in real time; parse the log information to obtain the real-time parameters of each switch; determine any switch as an abnormal switch in response to the real-time parameters of any switch meeting a preset abnormal condition; and determine the log information of the abnormal switch as the first alarm information.
[0174] In some embodiments, the first information acquisition module 4551 is further configured to traverse the fault analysis system within a first preset time period, obtain the switch identifier of each switch and the device identifier of each service device in the fault analysis system within the first preset time period, as well as the device identifier of at least one service device connected to each switch and the switch identifier of each service device connected to the switch; determine the connection relationship between the switch and the service device based on the device identifier of at least one service device connected to each switch and the switch identifier of each service device connected to the switch; draw a topology diagram between the switch cluster and the service device cluster in the fault analysis system based on the connection relationship; and generate topology association information based on the topology diagram.
[0175] In some embodiments, the anomaly determination module 4552 is further configured to determine the switch identifier of the abnormal switch based on the first alarm information; determine the local topology association information between the abnormal switch and the service devices in the service device cluster based on the switch identifier and topology association information; and determine the abnormal service devices associated with the abnormal switch from the service device cluster based on the local topology association information.
[0176] In some embodiments, the anomaly determination module 4552 is further configured to obtain a first message from a pre-generated first message queue, wherein the first message is generated based on the switch identifier of the abnormal switch and the first alarm information, and the first message queue includes the first message; and to parse the first message to obtain the switch identifier of the abnormal switch and the first alarm information.
[0177] In some embodiments, the second information acquisition module 4553 is further configured to acquire a second message from a pre-generated second message queue, wherein the second message is generated based on the device identifier of the abnormal service device, and the second message queue includes the second message; parse the second message to obtain the device identifier of the abnormal service device; determine a processing strategy for the abnormal service device based on the device identifier of the abnormal service device; and determine a second alarm message for the abnormal service device based on the device identifier of the abnormal service device.
[0178] In some embodiments, the second information acquisition module 4553 is further configured to, in response to determining that the abnormal service device is a first type of abnormal service device based on the device identifier of the abnormal service device, determine that the processing strategy for the first type of abnormal service device is device isolation processing; and in response to determining that the abnormal service device is a second type of abnormal service device based on the device identifier of the abnormal service device, determine that the processing strategy for the second type of abnormal service device is to send a fault reminder message to the information processing account bound to the second type of abnormal service device; wherein, the first type of abnormal service device is a service device that is currently on sale, and the second type of abnormal service device is a service device that has already been sold.
[0179] In some embodiments, the apparatus further includes an information sending module, configured to, in response to receiving a fault analysis request sent by an information processing account, obtain the fault analysis results of the fault analysis system; filter out the fault information of the second type of abnormal service device from the fault analysis results according to the device identifier of the second type of abnormal service device; and send the fault information to the information processing account.
[0180] In some embodiments, the second information acquisition module 4553 is further configured to acquire log information of the abnormal service device within a second preset time period based on the device identifier of the abnormal service device; and parse the log information to obtain the second alarm information of the abnormal service device.
[0181] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the fault analysis method described above in this application.
[0182] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the fault analysis method provided in this application. For example, ... Figure 3 The fault analysis method is shown.
[0183] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0184] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0185] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0186] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0187] In summary, the embodiments of this application can be applied to a fault analysis system. This system can include a switch cluster and a service device cluster, allowing switches and service devices to be managed by the same system, improving operational efficiency. When the fault analysis system receives a first alarm message from an abnormal switch, it can obtain the topology association information between the switches in the switch cluster and the service devices in the service device cluster. Then, based on the first alarm message and the topology association information, it determines the abnormal service device associated with the abnormal switch from the service device cluster. This allows for rapid location of the affected abnormal service device in the event of a switch malfunction. Next, it obtains the second alarm message from the abnormal service device. Finally, it performs synchronous analysis of the first and second alarm messages to obtain the fault analysis results. Because the affected abnormal service device has been quickly located, the service device can quickly perceive the switch malfunction. Furthermore, by combining the first and second alarm messages, it can quickly and accurately analyze the malfunction of the switch and service device. Periodic updates of the topology association information between the switches in the switch cluster and the service devices in the service device cluster within a preset time period can improve the accuracy and timeliness of the topology association information. When a switch malfunctions, a message can be generated for the switch. This allows for the timely detection and handling of switch anomalies, thereby improving the efficiency of anomaly handling. Based on the switch identifier and topology association information, it determines the local topology association information between the abnormal switch and the service devices in the service device cluster. It enables targeted handling measures for different types of abnormal service devices (both in-sale and sold). For in-sale service devices, device isolation is implemented to prevent the fault from spreading and ensure the normal operation of other devices. For sold service devices, fault alert messages are sent to the bound information processing account to promptly notify users, thus improving both the efficiency of fault handling and the user experience. It can also be used to identify abnormal service devices based on their device information. The system identifies and retrieves log information from abnormal service devices, then parses this information to obtain a second alarm message. This second alarm message allows for precise identification of the fault cause, status, and occurrence time of the abnormal service device, improving the efficiency and accuracy of fault diagnosis. By retrieving and parsing the second message from a pre-generated second message queue, the system can quickly determine the device identifier of the abnormal service device and implement targeted processing strategies based on the device identifier. This automated initial processing according to the processing strategy enhances the reliability and stability of the service device cluster and facilitates the acquisition of the second alarm message for subsequent analysis and processing.Real-time collection of log information from the switch cluster allows for the timely detection of abnormal switches by setting anomaly conditions. It also obtains topological association information between the switch cluster and the service device cluster, thereby identifying the abnormal service devices associated with the abnormal switches. This significantly improves network management and fault handling efficiency, ensuring network stability and reliability. Furthermore, corresponding handling strategies can be implemented for abnormal service devices to reduce their impact on other devices, and secondary alarm information from these devices is acquired for subsequent analysis. The server can filter fault information for these second-type abnormal service devices, along with the associated abnormal switches, from the fault analysis results based on their device identifiers. For example, if the second-type abnormal service device is service device 1, and the associated switches include switches A, B, and C, with switches B and C being the abnormal switches, then the fault information for switches B and C can also be filtered out. This collaborative analysis based on the fault information of service devices and switches improves the accuracy of fault analysis.
[0188] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A fault analysis method, characterized in that, The method is applied to a fault analysis system, which includes a switch cluster and a service device cluster; the fault analysis system includes: In response to receiving a first alarm message from an abnormal switch in the switch cluster, the topology association information between the switches in the switch cluster and the service devices in the service device cluster is obtained; Based on the first alarm information and the topology association information, the abnormal service device associated with the abnormal switch is determined from the service device cluster; Obtain the second alarm information of the abnormal service device; The first alarm information and the second alarm information are analyzed synchronously to obtain the fault analysis results of the fault analysis system.
2. The method according to claim 1, characterized in that, The first alarm message received from the abnormal switch in the switch cluster includes: Log information of each switch in the switch cluster is collected in real time; The log information is parsed to obtain the real-time parameters of each switch; If the real-time parameters of any switch meet the preset abnormal conditions, the switch is determined to be an abnormal switch. The log information of the abnormal switch is identified as the first alarm information.
3. The method according to claim 1, characterized in that, The step of obtaining the topology association information between the switches in the switch cluster and the service devices in the service device cluster includes: The fault analysis system is traversed within a first preset time period to obtain the switch identifier of each switch and the device identifier of each service device in the fault analysis system within the first preset time period, as well as the device identifier of at least one service device connected to each switch and the switch identifier of each service device connected to the switch. The connection relationship between the switch and the service device is determined based on the device identifier of at least one service device connected to each switch and the switch identifier of the switch connected to each service device. Based on the connection relationship, draw a topology diagram between the switch cluster and the service device cluster in the fault analysis system; The topological association information is generated based on the topological relationship graph.
4. The method according to claim 1, characterized in that, The step of determining the abnormal service device associated with the abnormal switch from the service device cluster based on the first alarm information and the topology association information includes: Based on the first alarm information, the switch identifier of the abnormal switch is determined; Based on the switch identifier and the topology association information, determine the local topology association information between the abnormal switch and the service devices in the service device cluster; Based on the local topology association information, the abnormal service device associated with the abnormal switch is determined from the service device cluster.
5. The method according to claim 4, characterized in that, The step of determining the switch identifier of the abnormal switch based on the first alarm information includes: The first message is obtained from a pre-generated first message queue, wherein the first message is generated based on the switch identifier of the abnormal switch and the first alarm information, and the first message queue includes the first message; The first message is parsed to obtain the switch identifier of the abnormal switch and the first alarm information.
6. The method according to any one of claims 1 to 5, characterized in that, The acquisition of the second alarm information of the abnormal service device includes: A second message is obtained from a pre-generated second message queue, wherein the second message is generated based on the device identifier of the abnormal service device, and the second message queue includes the second message; The second message is parsed to obtain the device identifier of the abnormal service device; Based on the device identifier of the abnormal service device, determine the processing strategy for the abnormal service device; Based on the device identifier of the abnormal service device, determine the second alarm information of the abnormal service device.
7. The method according to claim 6, characterized in that, The step of determining the processing strategy for the abnormal service device based on the device identifier of the abnormal service device includes: In response to determining that the abnormal service device is a first type of abnormal service device based on the device identifier of the abnormal service device, the processing strategy for the first type of abnormal service device is determined to be device isolation processing; In response to determining that the abnormal service device is a second type of abnormal service device based on the device identifier of the abnormal service device, the processing strategy for the second type of abnormal service device is determined to be to send a fault reminder message to the information processing account bound to the second type of abnormal service device; The first type of abnormal service equipment refers to service equipment currently on sale, while the second type of abnormal service equipment refers to service equipment that has already been sold.
8. The method according to claim 7, characterized in that, The method further includes: In response to receiving a fault analysis request sent by the information processing account, the fault analysis results of the fault analysis system are obtained; Based on the device identifier of the second type of abnormal service device, the fault information of the second type of abnormal service device is filtered out from the fault analysis results; The fault information is sent to the information processing account.
9. The method according to claim 6, characterized in that, The step of determining the second alarm information of the abnormal service device based on the device identifier of the abnormal service device includes: Based on the device identifier of the abnormal service device, obtain the log information of the abnormal service device within a second preset time period; The log information is parsed to obtain the second alarm information of the abnormal service device.
10. A fault analysis device, characterized in that, The device is applied to a fault analysis system, which includes a switch cluster and a service device cluster; the device includes: The first information acquisition module is used to acquire topology association information between the switches in the switch cluster and the service devices in the service device cluster in response to receiving a first alarm information sent by an abnormal switch in the switch cluster; An anomaly determination module is used to determine, based on the first alarm information and the topology association information, an abnormal service device associated with the abnormal switch from the service device cluster; The second information acquisition module is used to acquire the second alarm information of the abnormal service device; The information analysis module is used to perform synchronous analysis on the first alarm information and the second alarm information to obtain the fault analysis results of the fault analysis system.
11. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the fault analysis method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the fault analysis method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the fault analysis method according to any one of claims 1 to 9 is implemented.