Switch fault diagnosis and recovery methods, devices and switches

CN116489001BActive Publication Date: 2026-09-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310443014.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2026-09-18
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明提供了一种交换机故障诊断及恢复方法,以解决现有技术中,时间成本比较高,且对专业人员的专业度要求较高,且对交换机的故障诊断及恢复效率较低的问题

Benefits of technology

[0015]本申请实施例提供的交换机故障诊断及恢复方法,第二检测结果中包括发生异常的第二进程功能模块的标识信息,根据中间层的软件程序与驱动层的软件程序之间的调用关系以及第二进程功能模块的标识信息,从驱动层中确定第二进程功能模块调用的至少一个第三进程功能模块的标识信息,保证了确定的至少一个第三进程功能模块的标识信息的准确性,实现了缩小驱动层的检索范围效果,从而可以提高对交换机的故障进行诊断和恢复的效率。然后,根据各个第三进程功能模块的标识信息,对各个第三进程功能模块的配置信息以及第三进程功能模块中包括的各个加载驱动进程进行检测,生成第三检测结果,保证了生成的第三检测结果的准确性。对驱动层中包括的各个第三固件版本以及各个第三软件版本进行检测,生成第三通用检测结果,保证了生成的第三通用检测结果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489001B_ABST
    Figure CN116489001B_ABST
Patent Text Reader

Abstract

This invention relates to the field of switches, and discloses a method, apparatus, switch, and storage medium for fault diagnosis and recovery of switches. Based on the functions of various software programs and hardware devices within the switch, this invention divides the software programs and hardware devices into an application layer, a middleware layer, and a driver layer. It then performs detection on each of the application layer, middleware layer, and driver layer, generating detection results for each layer. Based on these detection results, it determines the location of the switch fault. Finally, based on the fault location, it restores the switch, generates a diagnostic recovery report, and sends the report to the target personnel. This method saves time, lowers the barrier to entry for fault diagnosis and recovery of switches, and improves the efficiency of fault diagnosis and recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of switches, and more specifically to methods, devices, and switches for fault diagnosis and recovery. Background Technology

[0002] Switches are crucial network devices responsible for forwarding data packets to their correct destinations. During operation, switches may encounter problems such as network congestion, port failures, and link flickering, necessitating troubleshooting and recovery.

[0003] In existing technologies, it is usually necessary for professionals to inspect the switch hardware, check the switch logs, and use network analysis tools to analyze the switch's data flow and network topology in order to diagnose and recover from switch failures.

[0004] The above methods require professional personnel to conduct step-by-step troubleshooting and testing of the switch, which is time-consuming, requires a high level of expertise from the personnel, and has low efficiency in fault diagnosis and recovery of the switch. Summary of the Invention

[0005] In view of this, the present invention provides a method for fault diagnosis and recovery of switches, in order to solve the problems of high time cost, high professional requirements for personnel, and low efficiency in fault diagnosis and recovery of switches in the prior art.

[0006] In a first aspect, the present invention provides a method for diagnosing and recovering from switch faults, the method comprising: Based on the functions of the various software programs and hardware devices in the switch, the software programs and hardware devices in the switch are divided into the application layer, the middleware layer, and the driver layer. Among them, the software programs in the application layer can call the software programs in the middleware layer, and the software programs in the middleware layer can call the software programs in the driver layer. The application layer, middleware layer, and driver layer are tested separately, and the test results for each layer are generated. Based on the detection results of each layer—application layer, middleware layer, and driver layer—the location of the switch fault is determined. Based on the location of the switch failure, the switch failure is restored, a diagnostic recovery report is generated, and the diagnostic recovery report is sent to the target personnel.

[0007] The switch fault diagnosis and recovery method provided in this application divides the software programs and hardware devices in the switch into application layer, middleware layer, and driver layer according to their functions, ensuring the accuracy of this division. This facilitates layered monitoring of the software programs and hardware devices in the switch. The application layer, middleware layer, and driver layer are detected separately, generating corresponding detection results for each layer, ensuring the accuracy of these results. Based on the detection results for each layer, the location of the switch fault is determined, ensuring the accuracy of the determined fault location. Then, based on the fault location, the switch fault is recovered, a diagnostic recovery report is generated, and the report is sent to the target personnel. This achieves fault recovery of the switch while ensuring that the target personnel can receive the diagnostic recovery report. The above method eliminates the need for professionals to perform step-by-step troubleshooting and testing on the switch, thus saving time and lowering the barrier to diagnosing and restoring switch faults, while also improving the efficiency of fault diagnosis and restoration.

[0008] In one optional implementation, the application layer, the middleware layer, and the driver layer are detected separately, and detection results corresponding to each of the application layer, the middleware layer, and the driver layer are generated, including: Perform detection on the application layer and generate application layer detection results; Based on the application layer detection results, the intermediate layer is detected, and intermediate layer detection results are generated. Based on the intermediate layer detection results, the driver layer is detected, and the driver layer detection results are generated.

[0009] The switch fault diagnosis and recovery method provided in this application embodiment detects the application layer and generates application layer detection results, ensuring the accuracy of the generated application layer detection results. Based on the application layer detection results, the intermediate layer is detected and intermediate layer detection results are generated, ensuring the accuracy of the generated intermediate layer detection results. Based on the intermediate layer detection results, the driver layer is detected and driver layer detection results are generated, ensuring the accuracy of the generated driver layer detection results. This, in turn, ensures the accuracy of determining the fault location of the switch based on the detection results corresponding to each layer: application layer, intermediate layer, and driver layer.

[0010] In one optional implementation, the application layer is inspected to generate application layer inspection results, including: The first firmware version and the first software version included in the application layer are detected to generate a first general detection result. Each first running process in the application layer is detected, and the first detection result is generated.

[0011] The switch fault diagnosis and recovery method provided in this application embodiment detects each first firmware version and each first software version included in the application layer, generating a first general detection result, thus ensuring the accuracy of the generated first general detection result. It also detects each first running process in the application layer, generating a first detection result, ensuring the accuracy of the generated first detection result. Furthermore, it ensures the accuracy of the intermediate layer detection result generated based on the application layer detection result.

[0012] In one optional implementation, the first detection result includes the identification information of the first process functional module corresponding to the first running process that experienced the anomaly. Based on the application layer detection result, the intermediate layer is detected to generate an intermediate layer detection result, including: Based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process functional module, determine the identification information of at least one second process functional module called by the first process functional module from the middle layer. Based on the identification information of each second process functional module, the configuration information of each second process functional module and each second running process included in the second process functional module are detected, and a second detection result is generated. The various second firmware versions and various second software versions included in the intermediate layer are detected, and a second general detection result is generated.

[0013] The switch fault diagnosis and recovery method provided in this application includes, in the first detection result, the identification information of the first process functional module corresponding to the first running process that has malfunctioned. Based on the calling relationship between the application layer software program and the middle layer software program, and the identification information of the first process functional module, the identification information of at least one second process functional module called by the first process functional module is determined from the middle layer. This ensures the accuracy of the identified identification information of at least one second process functional module, effectively narrowing the search scope of the middle layer and thus improving the efficiency of fault diagnosis and recovery for the switch. Then, based on the identification information of each second process functional module, the configuration information of each second process functional module and each second running process included in the second process functional module are detected to generate a second detection result, ensuring the accuracy of the generated second detection result. The various second firmware versions and various second software versions included in the middle layer are detected to generate a second general detection result, ensuring the accuracy of the generated second general detection result. This further ensures the accuracy of the driver layer detection result generated based on the middle layer detection result.

[0014] In one optional implementation, the second detection result includes the identification information of the second process functional module that experienced the anomaly. Based on the intermediate layer detection result, the driver layer is detected to generate a driver layer detection result, including: Based on the calling relationship between the intermediate layer software program and the driver layer software program and the identification information of the second process functional module, determine the identification information of at least one third process functional module called by the second process functional module from the driver layer. Based on the identification information of each third process functional module, the configuration information of each third process functional module and the loading driver processes included in the third process functional module are detected, and a third detection result is generated. The system detects each third firmware version and each third software version included in the driver layer and generates a third general detection result.

[0015] The switch fault diagnosis and recovery method provided in this application includes the identification information of the abnormal second process functional module in the second detection result. Based on the call relationship between the intermediate layer software program and the driver layer software program, and the identification information of the second process functional module, the identification information of at least one third process functional module called by the second process functional module is determined from the driver layer. This ensures the accuracy of the determined identification information of at least one third process functional module, narrowing the search scope of the driver layer and improving the efficiency of switch fault diagnosis and recovery. Then, based on the identification information of each third process functional module, the configuration information of each third process functional module and the various loaded driver processes included in the third process functional module are detected to generate a third detection result, ensuring the accuracy of the generated third detection result. Finally, the various third firmware versions and various third software versions included in the driver layer are detected to generate a third general detection result, ensuring the accuracy of the generated third general detection result.

[0016] In one optional implementation, based on the location of the switch failure, the switch failure is recovered, a diagnostic recovery report is generated, and the diagnostic recovery report is sent to the target personnel, including: When an application layer failure occurs, the failure is recovered, and a first diagnostic recovery result is generated. Based on the initial diagnosis and recovery results, generate an initial diagnosis and recovery report and send the report to the target personnel.

[0017] The switch fault diagnosis and recovery method provided in this application embodiment recovers from an application layer fault by generating a first diagnostic recovery result, ensuring the accuracy of the generated first diagnostic recovery result. Then, based on the first diagnostic recovery result, a first diagnostic recovery report is generated and sent to the target personnel, ensuring the accuracy of the generated first diagnostic recovery report and guaranteeing that the target personnel can receive it.

[0018] In an optional implementation, the method further includes: When both the middleware layer and the application layer fail, the failure in the middleware layer is recovered. Once the fault in the intermediate layer is successfully recovered, the fault in the application layer is recovered. When the fault recovery of the intermediate layer fails, a flag is set for the intermediate layer to block it out, and the fault in the application layer is recovered to generate a second diagnostic recovery result. Based on the recovery results of the second diagnosis, a second diagnosis recovery report is generated and sent to the target personnel.

[0019] The switch fault diagnosis and recovery method provided in this application embodiment recovers from faults in the middle layer and application layer when both occur, ensuring the accuracy of the middle layer fault recovery. If the middle layer fault recovery is successful, the application layer fault is recovered; if the middle layer fault recovery fails, a flag is set on the middle layer to block it, and the application layer fault is recovered, generating a second diagnostic recovery result. This enables cross-layer calls from the application layer to the driver layer, achieving hot recovery of the switch where possible, reducing service downtime, improving system availability and robustness, and enhancing user satisfaction. Then, based on the second diagnostic recovery result, a second diagnostic recovery report is generated and sent to the target personnel, ensuring the accuracy of the generated report and guaranteeing that the target personnel receive it.

[0020] In an optional implementation, the method further includes: When the driver layer, middleware layer, and application layer all fail, the driver layer failure is recovered. When fault recovery at the driver layer fails, a third diagnostic recovery result is generated. Once the fault in the driver layer is successfully recovered, the fault in the intermediate layer is recovered. Once the fault in the intermediate layer is successfully recovered, the fault in the application layer is recovered. When the fault recovery of the intermediate layer fails, a flag is set for the intermediate layer to block it out, and the fault in the application layer is recovered to generate the fourth diagnostic recovery result. Based on the recovery results of the third or fourth diagnosis, a third diagnosis recovery report is generated and sent to the target personnel.

[0021] The switch fault diagnosis and recovery method provided in this application embodiment recovers from faults in the driver layer, middleware layer, and application layer when all three occur. If driver layer fault recovery fails, a third diagnostic recovery result is generated, ensuring its accuracy. If driver layer fault recovery is successful, the middleware layer is recovered; if middleware layer fault recovery is successful, the application layer fault is recovered, thus achieving switch fault recovery. If middleware layer fault recovery fails, a flag is set on the middleware layer to block it, and the application layer fault is recovered, generating a fourth diagnostic recovery result. This method enables cross-layer calls from the application layer to the driver layer, achieving hot recovery of the switch where possible. This reduces service downtime, improves system availability and robustness, enhances user satisfaction, and ensures the accuracy of the generated fourth diagnostic recovery result. Based on either the third or fourth diagnostic recovery result, a third diagnostic recovery report is generated and sent to the target personnel, ensuring both the accuracy and the ability of the target personnel to receive the report.

[0022] Secondly, the present invention provides a switch fault diagnosis and recovery device, the device comprising: The partitioning module is used to divide the software programs and hardware devices in the switch into application layer, middleware layer and driver layer according to the functions of each software program and hardware device in the switch. Among them, the software programs in the application layer can call the software programs in the middleware layer, and the software programs in the middleware layer can call the software programs in the driver layer. The detection module is used to detect the application layer, the middleware layer, and the driver layer respectively, and generate the corresponding detection results for each layer. The determination module is used to determine the location of the switch fault based on the detection results of each layer, including the application layer, the middleware layer, and the driver layer. The recovery module is used to recover from the fault in the switch based on the location of the fault, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0023] The switch fault diagnosis and recovery device provided in this application divides the software programs and hardware devices in the switch into application layer, middleware layer, and driver layer according to their functions, ensuring the accuracy of this division. This facilitates layered monitoring of the software programs and hardware devices in the switch. The application layer, middleware layer, and driver layer are detected separately, generating corresponding detection results for each layer, ensuring the accuracy of these results. Based on the detection results for each layer, the location of the switch fault is determined, ensuring the accuracy of the determined fault location. Then, based on the fault location, the switch fault is recovered, a diagnostic recovery report is generated, and the report is sent to the target personnel. This achieves fault recovery of the switch while ensuring that the target personnel can receive the diagnostic recovery report. The aforementioned device eliminates the need for professional personnel to perform step-by-step troubleshooting and testing of the switch, thus saving time and lowering the barrier to diagnosing and restoring switch faults, while also improving the efficiency of fault diagnosis and restoration.

[0024] Thirdly, the present invention provides a switch, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the switch fault diagnosis and recovery method of the first aspect or any corresponding embodiment described above. Attached Figure Description

[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating a switch fault diagnosis and recovery method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating another method for fault diagnosis and recovery of a switch according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating another method for fault diagnosis and recovery of a switch according to an embodiment of the present invention; Figure 4 This is a structural block diagram of a switch fault diagnosis and recovery device according to an embodiment of the present invention; Figure 5This is a schematic diagram of the hardware structure of the switch according to an embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Switches are crucial network devices responsible for forwarding data packets to their correct destinations. During operation, switches may encounter problems such as network congestion, port failures, and link flickering, necessitating troubleshooting and recovery.

[0029] In related technical cases, professionals typically need to first check the switch's hardware status, such as whether the power supply, fans, and ports are functioning properly. If a hardware fault is found, the hardware needs to be replaced or repaired. Secondly, the switch's configuration needs to be checked for correctness, such as VLAN configuration, port speeds, and link aggregation. If configuration errors are found, they need to be corrected. Next, by reviewing the switch's logs, one can understand the switch's operating status, events, and error messages, thus pinpointing the location of the fault. Logs can be viewed through a command-line interface or network management tools. Finally, by using network analysis tools, such as Wireshark and tcpdump, the switch's data flow and network topology can be analyzed to further pinpoint the root cause of the network problem.

[0030] Therefore, this invention provides a method for fault diagnosis and recovery of a switch. By detecting various software programs and hardware devices within the switch, the location of the fault is determined. Then, based on the location of the fault, the switch is restored. This achieves automatic detection and recovery of the switch.

[0031] According to an embodiment of the present invention, a method for diagnosing and recovering switch faults is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0032] This embodiment provides a method for diagnosing and recovering switch faults. The executing entity can be a device for diagnosing and recovering switch faults. This device can be implemented as part or all of the switch through software, hardware, or a combination of both. The following description uses a switch as the executing entity.

[0033] This embodiment provides a method for diagnosing and recovering from switch faults, which can be used in the aforementioned switches. Figure 1 This is a flowchart of a switch fault diagnosis and recovery method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Based on the functions of each software program and hardware device in the switch, divide the software programs and hardware devices in the switch into the application layer, the middleware layer, and the driver layer.

[0034] In this context, software programs in the application layer can call software programs in the middle layer, and software programs in the middle layer can call software programs in the driver layer.

[0035] The application layer is the top layer of the entire system, responsible for implementing the system's specific functions. It typically includes modules such as the user interface, business logic, and data processing. The application layer's task is to implement the corresponding functions according to requirements and interact with the hardware through the middleware layer and the driver layer.

[0036] The middle layer provides common software components and interfaces, shielding the differences in hardware driver layers of different products to facilitate application layer development.

[0037] The driver layer is the interface layer between hardware and software, responsible for handling access to and control of the underlying hardware. The driver layer typically interacts directly with the hardware, providing a set of APIs (Application Programming Interfaces) for upper-layer modules to call. Through these APIs, upper-layer modules can perform hardware initialization, data read / write operations, and other tasks. The main task of the driver layer is to ensure that the hardware functions correctly, while providing a set of simple and easy-to-use interfaces to facilitate development by upper-layer modules.

[0038] Specifically, the driver layer is divided based on the type and function of the hardware device. It mainly abstracts the functions of the hardware. Generally, one hardware device corresponds to one driver, such as network driver, USB driver, CPLD driver, sensor driver, etc.

[0039] The intermediate layer is a standardized driver layer interface that follows general specifications to achieve driver uniformity. Because different products have different hardware and chip designs, the driver layers vary significantly; therefore, an intermediate layer is used to standardize and unify the interface.

[0040] The application layer is divided based on the business requirements and functional modules of the software system, such as business functional modules that provide related functions to users, such as forwarding data packets and network isolation.

[0041] Step S102: Detect the application layer, middleware layer, and driver layer respectively, and generate the detection results corresponding to each layer.

[0042] Specifically, the switch can perform detection on the application layer, middle layer, and driver layer respectively based on the calling relationship between the software programs of the application layer, middle layer, and driver layer, thereby generating separate detections for the application layer, middle layer, and driver layer.

[0043] This step will be explained in detail below.

[0044] Step S103: Determine the location of the switch fault based on the detection results of each layer, including the application layer, the middle layer, and the driver layer.

[0045] Specifically, the switch can perform a horizontal comparison of the detection results corresponding to the application layer, the middle layer, and the driver layer, and determine the location of the switch fault based on the comparison results.

[0046] Step S104: Based on the location of the switch failure, restore the switch to normal operation, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0047] Specifically, after determining the location of the fault, the switch can restore the fault, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0048] The switch fault diagnosis and recovery method provided in this application divides the software programs and hardware devices in the switch into application layer, middleware layer, and driver layer according to their functions, ensuring the accuracy of this division. This facilitates layered monitoring of the software programs and hardware devices in the switch. The application layer, middleware layer, and driver layer are detected separately, generating corresponding detection results for each layer, ensuring the accuracy of these results. Based on the detection results for each layer, the location of the switch fault is determined, ensuring the accuracy of the determined fault location. Then, based on the fault location, the switch fault is recovered, a diagnostic recovery report is generated, and the report is sent to the target personnel. This achieves fault recovery of the switch while ensuring that the target personnel can receive the diagnostic recovery report. The above method eliminates the need for professionals to perform step-by-step troubleshooting and testing on the switch, thus saving time and lowering the barrier to diagnosing and restoring switch faults, while also improving the efficiency of fault diagnosis and restoration.

[0049] This embodiment provides a method for diagnosing and recovering from switch faults, which can be used in the aforementioned switches. Figure 2 This is a flowchart of a switch fault diagnosis and recovery method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Based on the functions of each software program and hardware device in the switch, divide the software programs and hardware devices in the switch into the application layer, the middleware layer, and the driver layer.

[0050] In this context, software programs in the application layer can call software programs in the middle layer, and software programs in the middle layer can call software programs in the driver layer.

[0051] Please see details Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0052] Step S202: Detect the application layer, middleware layer, and driver layer respectively, and generate the corresponding detection results for each layer.

[0053] Specifically, step S202 includes: Step S2021: Perform detection on the application layer and generate application layer detection results.

[0054] In some optional implementations, step S2021 above includes: Step a1: Detect each first firmware version and each first software version included in the application layer, and generate a first general detection result.

[0055] Specifically, the switch can call the API to read the various first firmware versions and various first software versions included in the application layer, and then read the first standard firmware version corresponding to each first firmware version and the first standard software version corresponding to each first software version from the configuration file.

[0056] The switch can compare each first firmware version with its corresponding first standard firmware version. When the difference between each first firmware version and its corresponding first standard firmware version is less than a preset difference threshold, the difference is recorded. When the difference between each first firmware version and its corresponding first standard firmware version is greater than the preset difference threshold, an alarm message is output.

[0057] Similarly, the switch can compare each first software version with its corresponding first standard software version. When the difference between each first software version and its corresponding first standard software version is less than a preset difference threshold, the difference is recorded. When the difference between each first software version and its corresponding first standard software version is greater than the preset difference threshold, an alarm message is output.

[0058] Step a2: Detect each first running process in the application layer and generate the first detection result.

[0059] Optionally, the switch can view the process logs corresponding to each first running process in the application layer in real time, thereby detecting each first running process in the application layer and generating the first detection result.

[0060] Optionally, the switch can also use process detection tools to detect each first running process in the application layer and generate a first detection result.

[0061] When all the first running processes in the application layer are running normally, the first detection result can be that all the first running processes are detected as normal.

[0062] When there is an abnormal first running process in the application layer, the switch can determine the identification information of the first process functional module corresponding to the abnormal first running process based on the correspondence between each first running process and the first process functional module. Therefore, the first detection result may include the identification information of the first process functional module corresponding to the abnormal first running process.

[0063] Step S2022: Based on the application layer detection results, the intermediate layer is detected, and the intermediate layer detection results are generated.

[0064] In some optional implementations, the first detection result includes the identification information of the first process functional module corresponding to the first running process that experienced the anomaly, and step S2022 includes: Step b1: Based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process functional module, determine the identification information of at least one second process functional module called by the first process functional module from the middle layer.

[0065] Specifically, when the first detection result includes the identification information of the first process functional module corresponding to the first running process that has an abnormality, the switch determines that the first process functional module in the application layer has a fault, in order to accurately determine the location and cause of the fault.

[0066] The switch can determine the identification information of at least one second process function module called by the first process function module from the middle layer based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process function module.

[0067] Step b2: Based on the identification information of each second process functional module, detect the configuration information of each second process functional module and the second running processes included in the second process functional module, and generate a second detection result.

[0068] Specifically, after determining the identification information of each second process functional module from the intermediate layer, the switch can detect the configuration information of each second process functional module and the second running processes included in the second process functional module, and generate a second detection result.

[0069] For example, assuming that the first detection result includes the first process function module corresponding to the first running process that has an abnormality, which is the module corresponding to the switch fan, the switch determines the identification information of at least one second process function module corresponding to the switch fan in the middle layer based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process function module.

[0070] Then, the configuration information of at least one second process functional module corresponding to the switch fan in the intermediate layer and the various second running processes included in the second process functional module are detected to generate a second detection result.

[0071] Optionally, when all the second running processes included in the second process function module are running normally, the second detection result can be that all the second running processes are detected as normal, in which case the switch determines that the fault only occurs at the application layer.

[0072] Optionally, when an abnormal second running process exists within the second process functional module, the switch determines the abnormal second process functional module based on the correspondence between the abnormal second running process and the second process functional module. Therefore, the second detection result may include the identification information of the abnormal second process functional module.

[0073] Step b3: Detect each second firmware version and each second software version included in the intermediate layer, and generate a second general detection result.

[0074] Specifically, the switch can call the API to read the various second firmware versions and various second software versions included in the middle layer, and then read the second standard firmware version corresponding to each second firmware version and the second standard software version corresponding to each second software version from the configuration file.

[0075] The switch can compare each second firmware version with its corresponding second standard firmware version. When the difference between each second firmware version and its corresponding second standard firmware version is less than a preset difference threshold, the difference is recorded. When the difference between each second firmware version and its corresponding second standard firmware version is greater than the preset difference threshold, an alarm message is output.

[0076] Similarly, the switch can compare each second software version with its corresponding second standard software version. When the difference between each second software version and its corresponding second standard software version is less than a preset difference threshold, the difference is recorded. When the difference between each second software version and its corresponding second standard software version is greater than the preset difference threshold, an alarm message is output.

[0077] Step S2023: Based on the intermediate layer detection results, the driver layer is detected, and the driver layer detection results are generated.

[0078] In some optional implementations, the second detection result includes identification information of the second process functional module that has malfunctioned, and step S2023 includes: Step c1: Based on the calling relationship between the intermediate layer software program and the driver layer software program and the identification information of the second process functional module, determine the identification information of at least one third process functional module called by the second process functional module from the driver layer.

[0079] Specifically, when the second detection result includes the identification information of the second process function module that has malfunctioned, the switch determines that the second process function module in the middle layer has a fault, in order to accurately determine the location and cause of the fault.

[0080] The switch can determine the identification information of at least one or two third process functional modules called by the second process functional module from the driver layer based on the calling relationship between the middle layer software program and the driver layer software program and the identification information of the second process functional module.

[0081] Step c2: Based on the identification information of each third process functional module, detect the configuration information of each third process functional module and the loading driver processes included in the third process functional module, and generate a third detection result.

[0082] Specifically, after determining the identification information of each third process functional module from the driver layer, the switch can detect the configuration information of each third process functional module and the various loaded driver processes included in the third process functional module, and generate a third detection result.

[0083] Optionally, when all the loading driver processes included in the third process function module are running normally, the third detection result can be that all loading driver processes are detected as normal, in which case the switch determines that the fault only occurs in the application layer and the middle layer.

[0084] Optionally, when an abnormal loading driver process exists within the third-process functional module, the switch determines the abnormal third-process functional module based on the correspondence between the abnormal loading driver process and the third-process functional module. The switch then determines that the fault occurs at the driver layer, middleware layer, and application layer, and that the abnormality may be due to the abnormal third-process functional module in the driver layer, leading to an abnormality in the middleware layer, and consequently, an abnormality in the application layer.

[0085] Step c3: Detect each third firmware version and each third software version included in the driver layer, and generate a third general detection result.

[0086] Specifically, the switch can call the API to read the various third firmware versions and various third software versions included in the driver layer, and then read the third standard firmware version corresponding to each third firmware version and the third standard software version corresponding to each third software version from the configuration file.

[0087] The switch can compare each third-party firmware version with its corresponding third-party standard firmware version. When the difference between each third-party firmware version and its corresponding third-party standard firmware version is less than a preset difference threshold, the difference is recorded. When the difference between each third-party firmware version and its corresponding third-party standard firmware version is greater than the preset difference threshold, an alarm message is output.

[0088] Similarly, the switch can compare each third-party software version with its corresponding third-party standard software version. When the difference between each third-party software version and its corresponding third-party standard software version is less than a preset difference threshold, the difference is recorded. When the difference between each third-party software version and its corresponding third-party standard software version exceeds the preset difference threshold, an alarm message is output.

[0089] Step S203: Determine the location of the switch fault based on the detection results of each layer, including the application layer, the middle layer, and the driver layer.

[0090] Please see details Figure 1 Step S103 of the illustrated embodiment will not be described again here.

[0091] Step S204: Based on the location of the switch failure, restore the switch to its original state, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0092] Please see details Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0093] The switch fault diagnosis and recovery method provided in this application embodiment detects each first firmware version and each first software version included in the application layer, generating a first general detection result, thus ensuring the accuracy of the generated first general detection result. It also detects each first running process in the application layer, generating a first detection result, ensuring the accuracy of the generated first detection result. Furthermore, it ensures the accuracy of the intermediate layer detection result generated based on the application layer detection result.

[0094] When the first detection result includes the identification information of the first process functional module corresponding to the first running process that has an anomaly, based on the call relationship between the application layer software program and the middle layer software program and the identification information of the first process functional module, the identification information of at least one second process functional module called by the first process functional module is determined from the middle layer. This ensures the accuracy of the identified identification information of at least one second process functional module and narrows the search scope of the middle layer, thereby improving the efficiency of fault diagnosis and recovery of the switch. Then, based on the identification information of each second process functional module, the configuration information of each second process functional module and each second running process included in the second process functional module are detected to generate a second detection result, ensuring the accuracy of the generated second detection result. The various second firmware versions and various second software versions included in the middle layer are detected to generate a second general detection result, ensuring the accuracy of the generated second general detection result. This, in turn, ensures the accuracy of the driver layer detection result generated based on the middle layer detection result.

[0095] When the second detection result includes the identification information of the abnormal second process functional module, based on the call relationship between the intermediate layer software program and the driver layer software program, and the identification information of the second process functional module, the identification information of at least one third process functional module called by the second process functional module is determined from the driver layer. This ensures the accuracy of the identified identification information of at least one third process functional module, narrowing the search scope of the driver layer and thus improving the efficiency of fault diagnosis and recovery of the switch. Then, based on the identification information of each third process functional module, the configuration information of each third process functional module and the various loaded driver processes included in the third process functional module are detected to generate a third detection result, ensuring the accuracy of the generated third detection result. Finally, the various third firmware versions and various third software versions included in the driver layer are detected to generate a third general detection result, ensuring the accuracy of the generated third general detection result.

[0096] This embodiment provides a method for diagnosing and recovering from switch faults, which can be used in the aforementioned switches. Figure 3 This is a flowchart of a switch fault diagnosis and recovery method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Based on the functions of each software program and hardware device in the switch, divide the software programs and hardware devices in the switch into the application layer, the middleware layer, and the driver layer.

[0097] In this context, software programs in the application layer can call software programs in the middle layer, and software programs in the middle layer can call software programs in the driver layer.

[0098] Please see details Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0099] Step S302: Detect the application layer, middleware layer, and driver layer respectively, and generate the corresponding detection results for each layer.

[0100] Please see details Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0101] Step S303: Determine the location of the switch fault based on the detection results of each layer, including the application layer, the middle layer, and the driver layer.

[0102] Please see details Figure 2 Step S203 of the illustrated embodiment will not be described again here.

[0103] Step S304: Based on the location of the switch failure, restore the switch to its original state, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0104] In some alternative implementations, step 304 above may include the following: In one scenario, step 3041, when a failure occurs at the application layer, the failure is recovered, and a first diagnostic recovery result is generated.

[0105] Step 3042: Based on the first diagnostic recovery result, generate a first diagnostic recovery report and send the first diagnostic recovery report to the target personnel.

[0106] Specifically, when an application layer failure occurs, the switch can use a preset recovery method to recover from the application layer failure and generate a first diagnostic recovery result based on the recovery result.

[0107] The pre-examination recovery method may be to restart the first running process that failed, or it may be to use other methods. This application embodiment does not specifically limit the preset recovery method.

[0108] When application-layer fault recovery is successful, the first diagnostic recovery result can be used to indicate that the fault occurred at the application layer and recovery was successful. When application-layer fault recovery fails, the first diagnostic recovery result can be used to indicate that the fault occurred at the application layer and recovery failed.

[0109] Then, based on the results of the first diagnostic recovery, the switch generates a first diagnostic recovery report and sends the first diagnostic recovery report to the target personnel.

[0110] In another scenario, step 3043, when both the intermediate layer and the application layer fail, the failure of the intermediate layer is recovered.

[0111] Step 3044: When the fault recovery in the intermediate layer is successful, the fault recovery in the application layer is performed.

[0112] Step 3045: When the fault recovery of the intermediate layer fails, set a flag bit for the intermediate layer to block it out, recover the fault in the application layer, and generate a second diagnostic recovery result.

[0113] Step 3046: Based on the second diagnostic recovery results, generate a second diagnostic recovery report and send the second diagnostic recovery report to the target personnel.

[0114] Specifically, when both the middle layer and the application layer fail, since a failure in the middle layer may affect a failure in the application layer, the switch can first use a preset recovery method to recover from the failure in the middle layer.

[0115] When the fault in the intermediate layer is successfully recovered, the switch uses a preset recovery method to recover the fault in the application layer.

[0116] When fault recovery at the intermediate layer fails, a flag is set for the intermediate layer to block it and prevent it from affecting the application layer. Then, the switch uses a preset recovery method to recover from the fault at the application layer and generates a second diagnostic recovery result.

[0117] The pre-examination recovery method may be to restart the first running process that failed, or it may be to use other methods. This application embodiment does not specifically limit the preset recovery method.

[0118] The second diagnostic recovery result can be used to characterize that the fault occurred in the intermediate layer and the application layer, with the intermediate layer recovering successfully and the application layer also recovering successfully; the second diagnostic recovery result can also be used to characterize that the fault occurred in the intermediate layer and the application layer, with the intermediate layer recovering successfully and the application layer also failing to recover; the second diagnostic recovery result can also be used to characterize that the fault occurred in the intermediate layer and the application layer, with the intermediate layer recovering failing and the application layer recovering successfully; the second diagnostic recovery result can also be used to characterize that the fault occurred in the intermediate layer and the application layer, with the intermediate layer recovering failing and the application layer recovering failing.

[0119] Then, based on the results of the second diagnostic recovery, the switch generates a second diagnostic recovery report and sends the report to the target personnel.

[0120] In another scenario, step 3047, when the driver layer, intermediate layer, and application layer all fail, the driver layer is restored from failure.

[0121] Step 3048: When the fault recovery of the driving layer fails, a third diagnostic recovery result is generated.

[0122] Step 3049: When the fault recovery of the driving layer is successful, the fault recovery of the intermediate layer is performed.

[0123] Step 30410: When the fault in the intermediate layer is successfully recovered, the fault in the application layer is recovered.

[0124] Step 30411: When the fault recovery of the intermediate layer fails, set a flag bit for the intermediate layer to block it out, recover the fault in the application layer, and generate the fourth diagnostic recovery result.

[0125] Step 30412: Based on the recovery results of the third or fourth diagnosis, generate a third diagnosis recovery report and send the third diagnosis recovery report to the target personnel.

[0126] Specifically, when the driver layer, middleware layer, and application layer all fail, a driver layer failure will affect a middleware failure, and a middleware failure may also affect an application layer failure. Therefore, the switch can first use a preset recovery method to recover from a driver layer failure.

[0127] When fault recovery at the driver layer fails, a third diagnostic recovery result is generated. This third diagnostic recovery result can be used to characterize that faults have occurred at the driver layer, middleware layer, and application layer, and that fault recovery at the driver layer has failed.

[0128] When the driver layer fault is successfully recovered, the switch can use a preset recovery method to recover the middle layer fault; when the middle layer fault is successfully recovered, the switch can use a preset recovery method to recover the application layer fault. Once the application layer fault is successfully recovered, a fourth diagnostic recovery result can be generated. This fourth diagnostic recovery result can be used to characterize that faults have occurred in the driver layer, middle layer, and application layer, and that the driver layer fault recovery, the middle layer fault recovery, and the application layer fault recovery have all been successful.

[0129] When the driver layer fault is successfully recovered, the switch can use a preset recovery method to recover from the intermediate layer fault. When the intermediate layer fault recovery fails, a flag is set for the intermediate layer to block it, preventing it from affecting the application layer. Then, the switch uses the preset recovery method to recover from the application layer fault. When the application layer recovery is successful, the switch can generate a fourth diagnostic recovery result. This fourth diagnostic recovery result indicates that faults have occurred in the driver layer, intermediate layer, and application layer, with the driver layer fault successfully recovered, the intermediate layer fault blocked, and the application layer fault successfully recovered.

[0130] When the driver layer fault is successfully recovered, the switch can use a preset recovery method to recover from the intermediate layer fault. When the intermediate layer fault recovery fails, a flag is set for the intermediate layer to block it, preventing it from affecting the application layer. Then, the switch uses the preset recovery method to recover from the application layer fault. When the application layer recovery fails, the switch can generate a fourth diagnostic recovery result. This fourth diagnostic recovery result indicates that faults have occurred in the driver layer, intermediate layer, and application layer, with the driver layer fault recovery successful, the intermediate layer fault blocked, and the application layer fault recovery failing.

[0131] Then, the switch generates a second diagnostic recovery report based on the second diagnostic recovery results and sends the second diagnostic recovery report to the target personnel.

[0132] The switch fault diagnosis and recovery method provided in this application embodiment recovers from an application layer fault by generating a first diagnostic recovery result, ensuring the accuracy of the generated first diagnostic recovery result. Then, based on the first diagnostic recovery result, a first diagnostic recovery report is generated and sent to the target personnel, ensuring the accuracy of the generated first diagnostic recovery report and guaranteeing that the target personnel can receive it.

[0133] When both the middleware layer and the application layer fail, the middleware layer is restored, ensuring the accuracy of the recovery. If the middleware layer recovery is successful, the application layer is then restored; if the middleware layer recovery fails, a flag is set on the middleware layer to block it, and the application layer is restored, generating a second diagnostic recovery result. This enables cross-layer calls from the application layer to the driver layer and, where possible, allows for hot recovery of the switch, reducing service downtime, improving system availability and robustness, and enhancing user satisfaction. Then, based on the second diagnostic recovery result, a second diagnostic recovery report is generated and sent to the target personnel, ensuring the accuracy of the generated report and guaranteeing its receipt by the target personnel.

[0134] When failures occur in the driver layer, middleware layer, and application layer, the system recovers from the driver layer failure. If driver layer recovery fails, a third diagnostic recovery result is generated, ensuring its accuracy. If driver layer recovery is successful, the system recovers from the middleware layer failure. If middleware layer recovery is successful, the system recovers from the application layer failure, thus achieving switch failure recovery. If middleware layer recovery fails, a flag is set to block the middleware layer, and the system recovers from the application layer failure, generating a fourth diagnostic recovery result. This system enables cross-layer calls from the application layer to the driver layer, achieving switch hot recovery where possible. This reduces service downtime, improves system availability and robustness, enhances user satisfaction, and ensures the accuracy of the generated fourth diagnostic recovery result. Based on either the third or fourth diagnostic recovery result, a third diagnostic recovery report is generated and sent to the target personnel, ensuring the report's accuracy and delivery.

[0135] This embodiment also provides a switch fault diagnosis and recovery device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0136] This embodiment provides a switch fault diagnosis and recovery device, such as Figure 4 As shown, it includes: The partitioning module 401 is used to partition the software programs and hardware devices in the switch according to the functions of each software program and hardware device in the switch, into the application layer, the middle layer and the driver layer; wherein, the software programs in the application layer can call the software programs in the middle layer, and the software programs in the middle layer can call the software programs in the driver layer. The detection module 402 is used to detect the application layer, the middle layer and the driver layer respectively, and generate the detection results corresponding to each layer.

[0137] The determination module 403 is used to determine the location of the switch fault based on the detection results of each layer, including the application layer, the middle layer, and the driver layer.

[0138] The recovery module 404 is used to recover the switch from the fault location, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

[0139] In some alternative implementations, the detection module 402 includes: The first detection unit 4021 is used to detect the application layer and generate application layer detection results.

[0140] The second detection unit 4022 is used to detect the intermediate layer based on the application layer detection results and generate intermediate layer detection results.

[0141] The third detection unit 4023 is used to detect the driver layer based on the intermediate layer detection results and generate driver layer detection results.

[0142] In some optional implementations, the first detection unit 4021 includes: The first detection subunit 40211 is used to detect each first firmware version and each first software version included in the application layer and generate a first general detection result.

[0143] The second detection subunit 40212 is used to detect each first running process in the application layer and generate a first detection result.

[0144] In some optional implementations, the first detection result includes the identification information of the first process functional module corresponding to the first running process that experienced the anomaly, and the second detection unit 4022 includes: The first determining subunit 40221 is used to determine the identification information of at least one second process function module called by the first process function module from the middle layer based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process function module.

[0145] The third detection subunit 40222 is used to detect the configuration information of each second process function module and the second running processes included in the second process function module according to the identification information of each second process function module, and generate a second detection result.

[0146] The fourth detection subunit 40223 is used to detect each second firmware version and each second software version included in the intermediate layer and generate a second general detection result.

[0147] In some optional implementations, the second detection result includes identification information of the second process functional module that experienced the anomaly, and the third detection unit 4023 includes: The second determining subunit 40231 is used to determine the identification information of at least one third process function module called by the second process function module from the driver layer based on the calling relationship between the software program of the intermediate layer and the software program of the driver layer and the identification information of the second process function module.

[0148] The fifth detection subunit 40232 is used to detect the configuration information of each third process functional module and the loading driver processes included in the third process functional module according to the identification information of each third process functional module, and generate a third detection result.

[0149] The sixth detection subunit 40233 is used to detect each third firmware version and each third software version included in the driver layer and generate a third general detection result.

[0150] In some alternative implementations, the recovery module 404 includes: The first recovery unit 4041 is used to recover from the failure when the application layer fails and generate a first diagnostic recovery result.

[0151] The first generation unit 4042 is used to generate a first diagnostic recovery report based on the first diagnostic recovery result and send the first diagnostic recovery report to the target personnel.

[0152] In some alternative implementations, the recovery module 404 further includes: The second recovery unit 4043 is used to recover the fault of the intermediate layer when both the intermediate layer and the application layer fail.

[0153] The third recovery unit 4044 is used to recover the fault in the application layer when the fault in the intermediate layer is successfully recovered.

[0154] The fourth recovery unit 4045 is used to set a flag bit for the intermediate layer when the fault recovery of the intermediate layer fails, to shield the intermediate layer, to recover the fault in the application layer, and to generate a second diagnostic recovery result.

[0155] The second generation unit 4046 is used to generate a second diagnostic recovery report based on the second diagnostic recovery result and send the second diagnostic recovery report to the target personnel.

[0156] In some alternative implementations, the recovery module 404 further includes: The fifth recovery unit 4047 is used to recover the fault of the driver layer when the driver layer, the intermediate layer and the application layer all fail.

[0157] The third generation unit 4048 is used to generate a third diagnostic recovery result when the fault recovery of the driving layer fails.

[0158] The sixth recovery unit 4049 is used to recover the fault in the intermediate layer when the fault recovery of the driving layer is successful.

[0159] The seventh recovery unit 40410 is used to recover the fault in the application layer when the fault in the intermediate layer is successfully recovered.

[0160] The eighth recovery unit 40411 is used to set a flag bit for the intermediate layer when the fault recovery of the intermediate layer fails, to shield the intermediate layer, to recover the fault in the application layer, and to generate the fourth diagnostic recovery result.

[0161] The fourth generation unit 40412 is used to generate a third diagnostic recovery report based on the third diagnostic recovery result or the fourth diagnostic recovery result, and send the third diagnostic recovery report to the target personnel.

[0162] In this embodiment, the switch fault diagnosis and recovery device is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0163] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0164] This invention also provides a switch having the above-described features. Figure 4 The switch fault diagnosis and recovery device shown is shown.

[0165] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a switch provided in an optional embodiment of the present invention, such as... Figure 5 As shown, the switch includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the switch, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple storage devices, if desired. Similarly, multiple switches can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.

[0166] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0167] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0168] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the switch based on the display of a mini-program landing page. Furthermore, the memory 20 may include high-speed random access memory and non-transient memory, such as at least one disk storage device, flash memory device, or other non-transient solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the switch via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0169] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0170] The switch also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0171] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the switch, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.

[0172] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0173] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for fault diagnosis and recovery of a switch, characterized in that, The method includes: Based on the functions of the various software programs and hardware devices in the switch, the software programs and hardware devices in the switch are divided into an application layer, a middleware layer, and a driver layer; wherein, the software programs in the application layer call the software programs in the middleware layer, and the software programs in the middleware layer call the software programs in the driver layer. The application layer includes each first firmware version and each first software version, which are then detected to generate a first general detection result. Each first running process in the application layer is detected, and a first detection result is generated; the first detection result includes the identification information of the first process functional module corresponding to the first running process that has an anomaly; Based on the calling relationship between the application layer software program and the middle layer software program and the identification information of the first process functional module, the identification information of at least one second process functional module called by the first process functional module is determined from the middle layer. Based on the identification information of each second process functional module, the configuration information of each second process functional module and each second running process included in the second process functional module are detected, and a second detection result is generated; the second detection result includes the identification information of the second process functional module that has an anomaly. Each second firmware version and each second software version included in the intermediate layer are detected to generate a second general detection result; Based on the calling relationship between the software program of the intermediate layer and the software program of the driver layer, and the identification information of the second process functional module, the identification information of at least one third process functional module called by the second process functional module is determined from the driver layer. Based on the identification information of each of the third process functional modules, the configuration information of each of the third process functional modules and the loading driver processes included in the third process functional modules are detected to generate a third detection result. The third firmware version and the third software version included in the driver layer are detected to generate a third general detection result. Based on the detection results of each layer, including the application layer, the intermediate layer, and the driver layer, the location of the switch fault is determined. Based on the location of the switch failure, the switch failure is restored, a diagnostic recovery report is generated, and the diagnostic recovery report is sent to the target personnel.

2. The method according to claim 1, characterized in that, The step of restoring the switch based on the location of the switch failure, generating a diagnostic recovery report, and sending the diagnostic recovery report to the target personnel includes: When a failure occurs in the application layer, the failure is recovered, and a first diagnostic recovery result is generated; Based on the first diagnostic recovery result, a first diagnostic recovery report is generated and sent to the target personnel.

3. The method according to claim 2, characterized in that, The method further includes: When both the intermediate layer and the application layer fail, the failure of the intermediate layer is recovered. When the fault in the intermediate layer is successfully recovered, the fault in the application layer is recovered; When the fault recovery of the intermediate layer fails, a flag is set for the intermediate layer to block it out, and the fault in the application layer is recovered to generate a second diagnostic recovery result. Based on the second diagnostic recovery result, a second diagnostic recovery report is generated and sent to the target personnel.

4. The method according to claim 3, characterized in that, The method further includes: When the driver layer, the intermediate layer and the application layer all fail, the driver layer is restored to its fault condition. When the fault recovery of the driving layer fails, a third diagnostic recovery result is generated; When the fault in the driving layer is successfully recovered, the fault in the intermediate layer is then recovered. When the fault in the intermediate layer is successfully recovered, the fault in the application layer is recovered; When the fault recovery of the intermediate layer fails, a flag is set for the intermediate layer to block it out, and the fault in the application layer is recovered to generate a fourth diagnostic recovery result. Based on the third or fourth diagnostic recovery result, a third diagnostic recovery report is generated and sent to the target personnel.

5. A switch fault diagnosis and recovery device, characterized in that, The device includes: The partitioning module is used to partition the software programs and hardware devices in the switch into an application layer, a middleware layer, and a driver layer based on the functions of each software program and hardware device in the switch; wherein, the software programs in the application layer can call the software programs in the middleware layer, and the software programs in the middleware layer can call the software programs in the driver layer. The detection module is used to detect each first firmware version and each first software version included in the application layer, and generate a first general detection result; to detect each first running process in the application layer, and generate a first detection result; the first detection result includes the identification information of the first process functional module corresponding to the first running process that has an anomaly; based on the calling relationship between the software program in the application layer and the software program in the middle layer and the identification information of the first process functional module, to determine the identification information of at least one second process functional module called by the first process functional module in the middle layer; based on the identification information of each second process functional module, to detect the configuration information of each second process functional module and each second running process included in the second process functional module, and generate a second detection result. The second detection result includes the identification information of the second process functional module that has malfunctioned; the second detection result is generated by detecting each second firmware version and each second software version included in the intermediate layer; based on the calling relationship between the software program of the intermediate layer and the software program of the driver layer and the identification information of the second process functional module, the identification information of at least one third process functional module called by the second process functional module is determined from the driver layer; based on the identification information of each third process functional module, the configuration information of each third process functional module and each loading driver process included in the third process functional module are detected to generate a third detection result; the third detection result is generated by detecting each third firmware version and each third software version included in the driver layer. The determination module is used to determine the location of the fault in the switch based on the detection results corresponding to each layer of the application layer, the middle layer, and the driver layer. The recovery module is used to recover the fault of the switch according to the location of the fault, generate a diagnostic recovery report, and send the diagnostic recovery report to the target personnel.

6. A switch, characterized in that, include: The system includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the switch fault diagnosis and recovery method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Hierarchical automatic recovery method and related equipment

    CN110795267A

  • Network system, method, and switch device

    US20170054591A1