Fault repair method and computer program product

By setting the watchdog during the CPU startup phase, obtaining its operating status and obtaining log files when timeout, and automatically locate and repairing faults, the problem of difficulty in positioning and repairing after CPU downtime is solved, and the fault repair efficiency is improved.

CN120353634AInactive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510845956.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the CPU goes down, the board management controller cannot perceive the fault, making it difficult to automatically locate and repair system problems, and consumes a lot of labor costs.

Method used

Automatically locate and fix failures by setting the watchdog at different stages, obtaining its running status and obtaining log files when timeout.

Benefits of technology

No manual guessing or reproduction is required, saving labor costs and improving fault location and repair efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353634A_ABST
    Figure CN120353634A_ABST
Patent Text Reader

Abstract

The invention discloses a fault repairing method and a computer program product, and relates to the technical field of fault detection. Respectively starting a first watchdog corresponding to the starting stage of the basic input / output system, a second watchdog corresponding to the starting stage of the operating system and a third watchdog corresponding to the running stage of the operating system, and acquiring the running states of the first watchdog, the second watchdog and the third watchdog; acquiring a first log file corresponding to the target watchdog under the condition that the target watchdog is determined to be overtime based on the running states of the first watchdog, the second watchdog and the third watchdog; and based on the fault information of the trigger stage corresponding to the target watchdog in the first log file, performing fault repair on the fault information of the trigger stage. According to the fault repairing method, a large amount of labor cost is saved, and the efficiency of fault positioning and fault repairing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fault detection, and in particular, to a fault repair method and a computer program product. Background Art

[0002] In the storage system, many functions are necessary and effective means for fault location and equipment maintenance. When the functions are complete, it is beneficial for R & D and maintenance personnel to locate and repair problems after a fault occurs. If the functions are missing, a large amount of energy and time are required to reproduce, locate, and solve the problems. In related technologies, in some CPUs (Central Processing Unit), when the CPU crashes due to a fault, the baseboard management controller cannot sense the fault, so it cannot collect RAS (Reliability, Availability, Serviceability) logs, and cannot automatically collect them through the large system on the CPU side, resulting in difficulty in automatically locating the fault problem. Only guessing and reproduction can be relied on to locate and repair, which is likely to cause waste of personnel and consume a large amount of labor costs. Summary of the Invention

[0003] This application aims to solve at least one of the technical problems existing in the related technologies. For this purpose, this application provides a fault repair method and a computer program product, which can automatically collect the first log file, so that the fault problem can be automatically located based on the first log file to automatically repair the system fault, without the need for manual guessing or reproduction to locate and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair. In a first aspect, this application provides a fault repair method, including: Based on the startup sequences corresponding to the basic input / output system and the operating system, respectively start the first watchdog corresponding to the basic input / output system startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage, and obtain the running states of the first watchdog, the second watchdog, and the third watchdog; When it is determined that the target watchdog times out based on the running states of the first watchdog, the second watchdog, and the third watchdog, obtain the first log file corresponding to the target watchdog; Based on the fault information of the trigger stage corresponding to the target watchdog in the first log file, perform fault repair on the fault information of the trigger stage.

[0004] According to the fault repair method provided by the embodiments of the present application, by setting different watchdog timers at different stages, then determining the trigger stage corresponding to the watchdog timer that times out based on the running states of the watchdog timers at each stage, and obtaining the first log file corresponding to the watchdog timer, so as to repair the fault information corresponding to the trigger stage based on the first log file, it is possible to automatically collect the first log file, thereby automatically locating the fault problem based on the first log file to automatically repair the system fault, without the need for manual guessing or reproduction to locate and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair.

[0005] The fault repair method according to an embodiment of the present application, which respectively starts a first watchdog timer corresponding to the basic input / output system startup stage, a second watchdog timer corresponding to the operating system startup stage, and a third watchdog timer corresponding to the operating system running stage based on the startup sequence corresponding to the basic input / output system and the operating system, includes: When the central processing unit is powered on, start the first watchdog timer; When the basic input / output system starts up successfully, turn off the first watchdog timer and start the second watchdog timer; When the operating system starts up successfully, turn off the second watchdog timer and start the third watchdog timer.

[0006] The fault repair method according to an embodiment of the present application, where obtaining the running states of the first watchdog timer, the second watchdog timer, and the third watchdog timer includes: When the central processing unit is powered on, start the first watchdog timer; When a first shutdown signal is received within a first duration, determine that the running state of the first watchdog timer is off; When the first shutdown signal is not received within the first duration, determine that the running state of the first watchdog timer is timed out.

[0007] The fault repair method according to an embodiment of the present application, where the first duration is determined based on the following steps: Obtain a first startup duration corresponding to the basic input / output system when the basic input / output system is in a full-load state; Determine the first duration as a first multiple of the first startup duration; the first multiple is a value greater than 1.

[0008] The fault repair method according to an embodiment of the present application, where obtaining the running states of the first watchdog timer, the second watchdog timer, and the third watchdog timer includes: When the basic input / output system starts up successfully, start the second watchdog timer; When a second shutdown signal is received within the second time period, determine that the operating state of the second watchdog is shutdown; When the second shutdown signal is not received within the second time period, determine that the operating state of the second watchdog is timed out.

[0009] A fault repair method according to an embodiment of the present application, the obtaining the operating states of the first watchdog, the second watchdog, and the third watchdog includes: When the operating system starts up successfully, start the third watchdog; During the operation of the operating system, when a dog-feeding signal is detected within a third time period, determine that the operating state of the third watchdog is normal; the dog-feeding signal is sent by the operating system to the third watchdog based on a target dog-feeding period; During the operation of the operating system, when the dog-feeding signal is not detected within a fourth time period, determine that the operating state of the third watchdog is timed out; the fourth time period is greater than the third time period.

[0010] A fault repair method according to an embodiment of the present application, the obtaining a first log file corresponding to the target watchdog includes: Obtain the marking information corresponding to the target watchdog; Based on the marking information, obtain a first log file corresponding to the target watchdog.

[0011] A fault repair method according to an embodiment of the present application, the obtaining a first log file corresponding to the target watchdog includes: When it is determined that the central processing unit is powered on, determine all registers corresponding to the central processing unit; Based on a target protocol, obtain the numerical content of all the registers, and write the numerical content into a first log file; Obtain a first log file corresponding to the target watchdog.

[0012] A fault repair method according to an embodiment of the present application, before respectively starting the first watchdog corresponding to the basic input / output system startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage based on the startup sequence corresponding to the basic input / output system and the operating system, the method further includes: When the single board is powered on, start the programmable logic device and the baseboard management controller respectively; When the baseboard management controller starts up successfully, start the central processing unit.

[0013] A fault repair method according to an embodiment of the present application, the starting the central processing unit includes: When it is detected that the single board is powered on for the first time, start the central processing unit.

[0014] The fault repair method according to an embodiment of the present application, the starting the central processing unit includes: When receiving the heartbeat signal sent by the baseboard management controller, start the central processing unit.

[0015] The fault repair method according to an embodiment of the present application, the method further includes: Receive the first input of the user; In response to the first input, obtain a second log file; Based on the fault information in the second log file, perform fault repair on the fault information.

[0016] In a second aspect, the present application provides a computer program product, including: A first processing module, configured to respectively start a first watchdog corresponding to the basic input / output system startup phase, a second watchdog corresponding to the operating system startup phase, and a third watchdog corresponding to the operating system running phase based on the startup order corresponding to the basic input / output system and the operating system, and obtain the running states of the first watchdog, the second watchdog, and the third watchdog; A second processing module, configured to obtain a first log file corresponding to the target watchdog when it is determined that the target watchdog times out based on the running states of the first watchdog, the second watchdog, and the third watchdog; A third processing module, configured to perform fault repair on the fault information in the trigger phase corresponding to the target watchdog based on the fault information in the trigger phase corresponding to the target watchdog in the first log file.

[0017] According to the computer program product provided by the embodiments of the present application, by setting different watchdogs in different stages, then determining the trigger stage corresponding to the watchdog that times out according to the running states of the watchdogs in each stage, and obtaining the first log file corresponding to the watchdog, so as to repair the fault information corresponding to the trigger stage based on the first log file, it is possible to automatically collect the first log file, thereby automatically locating the fault problem based on the first log file to automatically repair the system fault, without manual guessing or reproduction to locate and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair.

[0018] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the fault repair method as described in the first aspect above.

[0019] Fourthly, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the fault repair method described in the first aspect above is implemented.

[0020] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects: By setting different watchdog timers at different stages, then determining the trigger stage corresponding to the watchdog timer that times out according to the running states of the watchdog timers at each stage, and obtaining the first log file corresponding to the watchdog timer, so as to repair the fault information corresponding to the trigger stage based on the first log file, it is possible to automatically collect the first log file, thereby automatically locating the fault problem based on the first log file to automatically repair the system fault, without the need for manual guessing or reproduction to locate and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair.

[0021] Furthermore, by obtaining the first startup duration of the basic input / output system in the full-load state, and then determining the duration greater than the first startup duration as the first duration to detect whether the first watchdog timer times out based on the first duration, it is possible to provide sufficient startup duration for the basic input / output system, avoiding the situation of misjudgment caused by insufficient preset countdown, and improving the accuracy and precision of detection.

[0022] Even further, by setting corresponding watchdog timers at each stage, and then monitoring the running states of the watchdog timers based on the countdown or regular dog-feeding strategy to determine the running states of each stage based on the running states of the watchdog timers, in the case where the watchdog timer times out, it is possible to accurately and quickly locate the fault trigger stage, thereby improving the efficiency of fault repair.

[0023] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0024] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 One of the flowcharts of a fault repair method provided by an embodiment of the present application; Figure 2 Another flowchart of a fault repair method provided by an embodiment of the present application; Figure 3The third flowchart of a fault repair method provided by an embodiment of the present application; Figure 4 The fourth flowchart of a fault repair method provided by an embodiment of the present application; Figure 5 The block diagram of a computer program product provided by an embodiment of the present application; Figure 6 The structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object may be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0028] Next, in conjunction with the accompanying drawings, the fault repair method, computer program product, electronic device, and readable storage medium provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0029] Among them, the fault repair method can be applied to a terminal, and specifically can be executed by hardware or software in the terminal.

[0030] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablet computers having a touch-sensitive surface (for example, a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (for example, a touch screen display and / or a touchpad).

[0031] In the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.

[0032] The fault repair method provided in the embodiment of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the fault repair method. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The fault repair method provided in the embodiment of the present application is described below using the electronic device as an example of the execution subject.

[0033] like Figure 1 As shown, the fault repair method includes: step 110, step 120 and step 130.

[0034] Step 110: based on the startup sequence of the basic input / output system and the operating system, respectively start a first watchdog corresponding to the basic input / output system startup phase, a second watchdog corresponding to the operating system startup phase, and a third watchdog corresponding to the operating system running phase, and obtain the running status of the first watchdog, the second watchdog, and the third watchdog; In this step, the Basic Input Output System (BIOS) is a program fixed to a ROM chip on the motherboard of the computer. The Basic Input Output System stores the computer's basic input and output programs, self-test programs after power-on, and system self-starting programs.

[0035] An operating system (OS) is a built-in program that works with the computer's various hardware to interact with the user.

[0036] A watchdog is a hardware or software mechanism used to monitor and restore the normal operation of a computer system or other electronic equipment.

[0037] The watchdog may include a hardware watchdog or a software watchdog, wherein the hardware watchdog may be an independent hardware chip or a module integrated in the processor, and may be implemented by a hardware timer; the software watchdog may be implemented by a timer of an operating system or application; the type of watchdog corresponding to each stage may be selected based on the user, and is not limited in this application.

[0038] The first watchdog is a watchdog corresponding to the startup phase of the basic input and output system. The running state of the first watchdog can be used to characterize the state of the startup phase of the basic input and output system. For example, when the basic input and output system is stuck during the initialization process, the running state of the first watchdog may be timeout.

[0039] After the system is powered on, the initialization code of the basic input and output system may be executed, and the basic input and output system may be responsible for hardware initialization and loading the boot loader.

[0040] The operating status of the first watchdog can be accessed through specific I / O ports or memory-mapped registers. In the basic input / output system code, its operating status can be confirmed by reading the watchdog status register or checking whether a reset is triggered.

[0041] The second watchdog is the watchdog corresponding to the startup phase of the operating system. The operating status of the second watchdog can be used to characterize the status of the startup phase of the operating system.

[0042] The operating status of the second watchdog can be confirmed by reading the watchdog status register or checking whether a reset is triggered.

[0043] The third watchdog is the watchdog corresponding to the running phase of the operating system. The operating status of the third watchdog can be used to characterize whether the operating system is in the normal running phase.

[0044] After the system is powered on, the basic input / output system can be started first, and then the operating system can be started. Based on the startup sequence corresponding to the basic input / output system and the operating system, the first watchdog, the second watchdog, and the third watchdog can be started respectively, and the operating status of each watchdog can be obtained during the operation of each watchdog.

[0045] Step 120: When it is determined that the target watchdog has timed out based on the operating status of the first watchdog, the second watchdog, and the third watchdog, obtain the first log file corresponding to the target watchdog; In this step, the target watchdog can be determined according to the operating status of the watchdog (such as whether a reset or timeout flag is triggered).

[0046] During the operation of the watchdog, a signal (i.e., "feeding the dog") can be sent to the watchdog regularly (such as through a timer interrupt or other means). If the watchdog does not receive the "feeding the dog" signal within the set time, the watchdog can be determined as the target watchdog.

[0047] For example, when a timeout flag is received, the first log file can record which phase of the watchdog triggered, and then obtain the first log file corresponding to the target watchdog in the triggered phase.

[0048] The first log file can be a RAS log, and the RAS log is used to capture operation environment information when the component software fails to run as scheduled.

[0049] Step 130: Based on the fault information in the triggered phase corresponding to the target watchdog in the first log file, perform fault repair on the fault information in the triggered phase.

[0050] In this step, the phase in which the target watchdog triggered and the relevant fault information can be extracted from the first log file.

[0051] According to the fault information in the log file, analyze the cause of the fault. For example, the fault information may include hardware faults (such as memory errors, CPU faults, or power supply problems, etc.), software faults (such as driver errors, kernel crashes, or user space process freezes, etc.), or configuration errors (such as improper watchdog timeout setting or basic input / output system configuration errors, etc.).

[0052] According to the fault information, corresponding repair steps can be executed, such as restarting the CPU or powering on the single board after power-off.

[0053] According to the fault repair method provided by the embodiments of the present application, by setting different watchdogs at different stages, then determining the trigger stage corresponding to the watchdog that times out based on the running status of the watchdogs at each stage, and obtaining the first log file corresponding to the watchdog, so as to repair the fault information corresponding to the trigger stage based on the first log file, it is possible to automatically collect the first log file, thereby automatically locating the fault problem based on the first log file to automatically repair the system fault, without manual guessing or reproduction for location and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair.

[0054] Such as Figure 2 shown, in some embodiments, step 110 may include: When the central processing unit is powered on, start the first watchdog; When the basic input / output system starts up successfully, turn off the first watchdog and start the second watchdog; When the operating system starts up successfully, turn off the second watchdog and start the third watchdog.

[0055] In this embodiment, when the central processing unit is powered on, the first watchdog can be initialized and started by hardware or BIOS startup code.

[0056] The first watchdog can be used to monitor the initialization process of the central processing unit and the motherboard. In the case of a fault (such as memory initialization failure) occurring at this stage, the first watchdog will trigger a system reset.

[0057] After the BIOS completes its startup process, the first watchdog can be turned off, and the system will enter the operating system startup stage.

[0058] The shutdown signal of the first watchdog can be used as the enable signal of the second watchdog. When the first watchdog is turned off, the second watchdog is started.

[0059] The second watchdog is used to ensure that the operating system can start up normally. In the case of a fault (such as a kernel crash) occurring at this stage, the second watchdog will trigger a system reset.

[0060] After the operating system starts up (e.g., after user space initialization is completed), the second watchdog can be turned off. Since the operating system has fully started up, the system will enter the normal operation stage.

[0061] The signal for turning off the second watchdog can be used as the enabling signal for the third watchdog. When the second watchdog is turned off, the third watchdog can be started.

[0062] The third watchdog is used to ensure that user space processes can run normally. In the event of a failure (such as a critical service getting stuck) during this stage, the third watchdog will trigger a system reset or an alarm.

[0063] In some embodiments, obtaining the running states of the first watchdog, the second watchdog, and the third watchdog may include: When the central processing unit is powered on, start the first watchdog; When the first shutdown signal is received within the first time period, determine that the running state of the first watchdog is off; When the first shutdown signal is not received within the first time period, determine that the running state of the first watchdog is timed out.

[0064] In this embodiment, as Figure 2 shown, the first time period can be any value between 10 and 60 minutes, which can be user-defined and is not limited in this application.

[0065] After the central processing unit (CPU) is powered on, the first watchdog in the BIOS startup stage can be enabled (e.g., it can be enabled using the cpu_reset signal). After the BIOS startup is completed (after the operating system loading is completed), the first watchdog can be turned off using the first shutdown signal (such as bios_start_complete).

[0066] The running state of the first watchdog can be used to characterize the startup state of the basic input / output system. For example, within the first time period, when the running state of the first watchdog is off, it can be determined that the basic input / output system startup is completed; within the first time period, when the running state of the first watchdog is timed out, it can be determined that the basic input / output system startup fails, that is, a failure occurs during the basic input / output system startup process.

[0067] During the actual execution process, for example, a CPLD (Complex Programmable Logic Device) or an FPGA (Field Programmable Gate Array) can start a countdown (with a timing duration of the first duration) after detecting the cpu_reset signal. The CPLD or FPGA shuts down the countdown after detecting the first shutdown signal and determines that the operating state of the first watchdog is off; in the case where the first shutdown signal is not detected within the first duration, it is regarded as an overtime watchdog bark, that is, it is determined that the operating state of the first watchdog is overtime.

[0068] In some embodiments, the first duration can be determined based on the following steps: Obtain the first startup duration corresponding to the basic input / output system when the basic input / output system is in a full-load state; Determine the first multiple of the first startup duration as the first duration.

[0069] In this embodiment, the first duration can be determined based on the longest time path for the startup of the basic input / output system (BIOS) of the current single board. For example, when the BIOS is fully configured, that is, when the BIOS is in a full-load state, obtain the first startup duration of the BIOS, and then calculate the first duration based on the first startup duration in the fully configured state.

[0070] For example, the first multiple of the first startup duration can be determined as the first duration, and the first multiple is a value greater than 1. The first multiple can be 1.2, 1.3, or 1.5, or it can also be other values, which are not limited in this application.

[0071] According to the fault repair method provided by the embodiments of the present application, by obtaining the first startup duration of the basic input / output system in a full-load state and then determining a duration greater than the first startup duration as the first duration to detect whether the first watchdog times out based on the first duration, it can provide sufficient startup duration for the basic input / output system, avoid misjudgment caused by insufficient preset countdown, and improve the accuracy and precision of detection.

[0072] In some embodiments, obtaining the operating states of the first watchdog, the second watchdog, and the third watchdog may include: After the basic input / output system starts up successfully, start the second watchdog; When the second shutdown signal is received within the second duration, determine that the operating state of the second watchdog is off; When the second shutdown signal is not received within the second duration, determine that the operating state of the second watchdog is overtime.

[0073] In this embodiment, as Figure 2 shown, the second duration can be 90s, or it can also be other values, which can be based on user-defined settings and are not limited in this application.

[0074] After the basic input / output system (BIOS) finishes booting and loading the operating system (OS), the first shutdown signal (such as bios_start_complete) can be used as the enabling signal for the second watchdog during the OS startup phase. After the OS startup is completed and enters the running phase, the second shutdown signal (such as the kernel_init_end signal) can be used to shut down the second watchdog.

[0075] The running state of the second watchdog can be used to characterize the startup state of the operating system. For example, within the second duration, if it is determined that the running state of the second watchdog is off, it can be determined that the operating system startup is completed; within the second duration, if it is determined that the running state of the second watchdog is timed out, it can be determined that the operating system startup fails, that is, a fault occurs during the operating system startup process.

[0076] In the actual execution process, for example, after the CPLD detects the bios_start_complete signal, it can start counting down (the counting duration is the second duration). After the CPLD detects the second shutdown signal, it stops the counting down and determines that the running state of the second watchdog is off; if the second shutdown signal is not detected within the second duration, it is regarded as the watchdog barking due to timeout, that is, it is determined that the running state of the second watchdog is timed out.

[0077] In some embodiments, the second duration can be determined based on the following steps: Obtain the second startup duration corresponding to the operating system when the operating system is in the full-load state; Determine the second duration as the second multiple of the second startup duration.

[0078] In this embodiment, the second duration can be determined based on the longest time path of the operating system startup. For example, when the operating system is in the full-load state, obtain the second startup duration of the operating system, and then calculate the second duration based on the second startup duration in the fully configured state.

[0079] For example, the second duration can be determined as the second multiple of the second startup duration, and the second multiple is a value greater than 1. The second multiple can be 1.2, 1.3, or 1.5, or it can also be other values, which are not limited in this application.

[0080] In this application, by obtaining the second startup duration of the operating system in the full-load state, and then determining the duration greater than the second startup duration as the second duration, to detect whether the second watchdog times out based on the second duration, it can provide sufficient startup duration for the operating system, avoid misjudgment caused by insufficient preset countdown, and improve the accuracy and precision of detection.

[0081] Continue to refer to Figure 2 , in some embodiments, obtaining the operating states of the first watchdog, the second watchdog, and the third watchdog may include: When the operating system starts up successfully, start the third watchdog; During the operation of the operating system, when a dog-feeding signal is detected within the third duration, determine that the operating state of the third watchdog is normal; the dog-feeding signal is sent by the operating system to the third watchdog based on the target dog-feeding period; During the operation of the operating system, when a dog-feeding signal is not detected within the fourth duration, determine that the operating state of the third watchdog times out.

[0082] In this embodiment, after the operating system starts up successfully, that is, after the initialization of the operating system kernel ends, a second shutdown signal (kernel_init_end signal) may be issued. The second shutdown signal may be the enabling signal of the third watchdog. When the CPLD detects the second shutdown signal, the third watchdog may be started.

[0083] The operating state of the third watchdog may be used to characterize the operating state of the operating system. For example, when it is determined that the operating state of the third watchdog is normal, it may be determined that the operating system is in a normal operating state; when it is determined that the operating state of the third watchdog times out, it may be determined that the operating system is in an abnormal operating state.

[0084] A strategy of regularly feeding the dog can be adopted to monitor the operating state of the operating system.

[0085] The dog-feeding signal is sent by the operating system to the third watchdog based on the target dog-feeding period, where the target dog-feeding period may be to send a dog-feeding signal to the third watchdog once every third duration.

[0086] A fixed dog-feeding time interval (i.e., the third duration) can be set. For example, the operating system can feed the third watchdog once every 10s, that is, the operating system can send a dog-feeding signal to the third watchdog once every 10s to indicate to the third watchdog that the operating system is running normally.

[0087] A maximum allowable non-dog-feeding time limit, that is, the fourth duration, can be set. The fourth duration is greater than the third duration. For example, it can be set to 60s, or it can also be set to other values, which are not limited in this application.

[0088] For example, within 60s, if the CPLD has not detected the watchdog signal sent by the operating system, it can be considered that an abnormal situation has occurred during the operation of the operating system, that is, it is regarded as a timeout. That is, the operating state of the third watchdog can be determined to be a timeout.

[0089] According to the fault repair method provided by the embodiments of the present application, by setting corresponding watchdogs in each stage, and then monitoring the operating state of the watchdogs based on a countdown or regular watchdog feeding strategy, to determine the operating state of each stage based on the operating state of the watchdogs. In the case of a watchdog timeout, the fault trigger stage can be accurately and quickly located, thereby improving the efficiency of fault repair.

[0090] In some embodiments, the method may further include: When a watchdog signal is detected within the third time period, reset the timer corresponding to the third watchdog; Detect the watchdog signal within the third time period starting from the reset moment of the timer corresponding to the third watchdog.

[0091] In this embodiment, within the set third time period time window, when the operating system sends a watchdog signal to the third watchdog, the watchdog timer will be reset. That is, the countdown can be reset to the initial value (i.e., "the third time period"), thereby restarting the countdown process.

[0092] After the timer is reset, the system starts a new "third time period" time window. Within this new time window, the CPLD can continue to detect the watchdog signal.

[0093] When a watchdog signal is detected again within this new time window, the timer will be reset again, and the above process will be repeated.

[0094] When a watchdog signal is not detected within this new time window, the timer will time out, and the operating state of the third watchdog is determined to be a timeout.

[0095] According to the fault repair method provided by the embodiments of the present application, by resetting the timer of the third watchdog when a watchdog signal is detected within the third time period, the operating state of the operating system can be monitored in real time, ensuring that the third watchdog timer will not mis-trigger the timeout processing mechanism, and improving the stability of the system.

[0096] In some embodiments, obtaining the first log file corresponding to the target watchdog may include: Based on the operating state of the target watchdog, or based on the monitoring signal corresponding to the central processing unit, obtain the first log file corresponding to the target watchdog.

[0097] In this embodiment, central processing units (CPUs) of different series may adopt different strategies to collect RAS logs.

[0098] For example, a CPU of the Intel series may collect RAS logs by monitoring the triggering of signals (such as the caterr (Catastrophic Error) signal) to locate fault problems. When the baseboard management controller receives a timeout signal of the target watchdog or detects a monitoring signal corresponding to the CPU, the process of automatically collecting the first log file may be triggered.

[0099] According to the fault repair method provided by the embodiments of the present application, by detecting the watchdog status and caterr status, the first log file is automatically collected when receiving a timeout signal of the watchdog or detecting a caterr signal corresponding to the CPU. The same set of log collection logic is used in CPU series that support and do not support caterr, as well as CPUs from different manufacturers, improving the versatility and flexibility of the method.

[0100] In some embodiments, step 120 may include: Obtaining the marking information corresponding to the target watchdog; Based on the marking information, obtaining the first log file corresponding to the target watchdog.

[0101] In this embodiment, when the BMC (Baseboard Management Controller) receives a dog barking mark or caterr, the log can record which stage of the watchdog is triggered. Then, according to the marking information (such as a flag bit) corresponding to the timed-out target watchdog, the first log file corresponding to the target watchdog can be obtained, so that it is possible to distinguish which stage the fault occurred based on the first log file.

[0102] According to the fault repair method provided by the embodiments of the present application, by obtaining the marking information corresponding to the target watchdog to obtain the first log file corresponding to the target watchdog based on the marking information, it is possible to distinguish the stage where the fault occurred based on the first log file, accurately locate the stage where the fault occurred, and improve the efficiency of fault repair.

[0103] In some embodiments, before obtaining the marking information corresponding to the target watchdog, the method may further include: When the first watchdog is the target watchdog, determining the marking information of the target watchdog as the first marking information; When the second watchdog is the target watchdog, determining the marking information of the target watchdog as the second marking information; When the third watchdog is the target watchdog, determine that the marker information of the target watchdog is the third marker information.

[0104] In this embodiment, when the CPLD detects a watchdog timeout, it can provide different marker bits to the BMC to distinguish the trigger phase.

[0105] For example, when the first watchdog corresponding to the startup phase of the basic input / output system times out, the CPLD can send the first marker information, such as the marker bit 0x11, to the BMC. The BMC can record this marker bit and identify it as a timeout in the BIOS startup phase in the log. The system can adopt specific repair strategies based on this marker bit, such as restarting the BIOS or checking the BIOS configuration, etc.

[0106] When the second watchdog corresponding to the startup phase of the operating system times out, the CPLD can send the second marker information, such as the marker bit 0x22, to the BMC. The BMC can record this marker bit and identify it as a timeout in the OS startup phase in the log. The system can take corresponding repair measures based on this marker bit, such as reloading the operating system or checking the startup configuration, etc.

[0107] When the third watchdog corresponding to the running phase of the operating system times out, the CPLD can send the third marker information, such as the marker bit 0x55, to the BMC. The BMC can record this marker bit and identify it as a timeout in the OS running phase in the log. The system can adopt different repair strategies based on this marker bit, such as restarting services, checking system resources, or performing fault diagnosis, etc.

[0108] According to the fault repair method provided by the embodiments of the present application, by providing different timeout marker information for the watchdogs in different phases, the trigger phase can be distinguished based on the marker information, so that it can be used for log record distinction, and different fault repair strategies can be implemented after the first log information is collected, which can more accurately locate the fault phase and thus more accurately and quickly repair the fault.

[0109] As Figure 3 shown, in some embodiments, step 120 may include: When it is determined that the central processing unit is powered on, determine all the registers corresponding to the central processing unit; Based on the target protocol, obtain the numerical content of all the registers and write the numerical content into the first log file; Obtain the first log file corresponding to the target watchdog.

[0110] In this embodiment, it is possible to check whether the central processing unit is in the powered-on state to avoid the timeout barking of the watchdog triggered by abnormal power-off.

[0111] When the central processing unit is powered on, the central processing unit series and the number of central processing units on the current single board can be read. Then, according to the MCE (Machine Check Exception, a mechanism triggered by the central processing unit when detecting hardware errors for reporting and handling these errors) registers supported by the current central processing unit, the numerical contents of all MCE registers in the table are read through the target protocol (such as APML (Advanced Platform Management Link)). For example, the numerical contents of the registers corresponding to each central processing unit can be read in sequence, and a log file is recorded until all the registers are read, and the record is output to the first log file (RAS log). After the first log file is read, the BMC can execute the automatic repair system fault strategy according to the corresponding trigger policy.

[0112] According to the fault repair method provided by the embodiments of the present application, by recording the watchdog trigger stage and register values, the cause of the system fault can be quickly located, which helps users quickly diagnose problems and improves the fault repair efficiency and the user experience.

[0113] In some embodiments, before step 110, the method may further include: When the single board is powered on, the programmable logic device and the baseboard management controller are started respectively; When the baseboard management controller starts up successfully, the central processing unit is started.

[0114] In this embodiment, after the single board is powered on, the programmable logic device and the baseboard management controller are first powered on and started.

[0115] The startup speed of the programmable logic device is relatively fast, usually completing initialization in milliseconds, while the startup time of the baseboard management controller is relatively long (usually several seconds to dozens of seconds, depending on the firmware complexity).

[0116] After the baseboard management controller starts up successfully, the power module can be controlled to power on the central processing unit.

[0117] In some embodiments, starting the central processing unit may include: When it is detected that the single board is powered on for the first time, the central processing unit is started.

[0118] In this embodiment, it can be detected whether the single board is powered on for the first time based on the programmable logic device. For example, it can be detected whether the single board is powered on for the first time through a hardware signal (such as detecting the power status register or the level of a specific pin).

[0119] In some embodiments, starting the central processing unit may include: Upon receiving the heartbeat signal sent by the baseboard management controller, start the central processing unit.

[0120] In this embodiment, after the programmable logic device finishes starting up, it can enter the waiting state until the baseboard management controller finishes starting up and sends a heartbeat signal, which is used to indicate that the baseboard management controller has finished starting up.

[0121] After the baseboard management controller runs normally, it can notify the programmable logic device to power on the central processing unit through the heartbeat signal or other communication mechanisms.

[0122] After receiving the signal, the programmable logic device can control the power module to power on the central processing unit.

[0123] According to the fault repair method provided by the embodiments of the present application, by defining the single-board power-on sequence, it can be ensured that the central processing unit is started after the baseboard management controller functions completely normally, which can avoid system failures or hardware damages caused by the unreadiness of the baseboard management controller, and ensure the monitoring of the system in the normal operating state of the baseboard management controller.

[0124] As Figure 4 shown, in some embodiments, the method may further include: Receiving a first input from the user; In response to the first input, obtaining a second log file; Based on the fault information in the second log file, performing fault repair on the fault information.

[0125] In this embodiment, the first input is used to input an acquisition instruction, where the acquisition instruction may include an IPMI (Intelligent Platform Management Interface) instruction.

[0126] Among them, the first input may be at least one of the following ways: First, the first input may be a touch operation, including but not limited to click operations, swipe operations, press operations, etc.

[0127] In this embodiment, receiving the first input from the user may be receiving a touch operation of the user in the display area of the terminal display screen.

[0128] In order to reduce the user's misoperation rate, the action area of the first input may be limited to a specific area, such as the upper middle area of the interface for obtaining the second log file; or in the state of displaying the interface for obtaining the second log file, a target control is displayed on the current interface, and touching the target control can achieve the first input; or the first input is set as a continuous multiple tapping operation on the display area within a target time interval.

[0129] Second, the first input may be a physical button input.

[0130] In this embodiment, an entity button corresponding to the second log file is provided on the body of the terminal, and a first input from the user can be received, for example, receiving an operation of the user pressing the corresponding entity button; the first input can also be a combined operation of pressing multiple entity buttons simultaneously.

[0131] Thirdly, the first input can be a voice input.

[0132] In this embodiment, when the terminal receives a voice such as "obtain the second log file", it can trigger an interface for displaying the second log file.

[0133] Of course, in other embodiments, the first input can also be in other forms, including but not limited to character input, etc., which can be specifically determined according to actual needs, and the embodiments of the present application do not limit this.

[0134] The second log file can be a RAS log.

[0135] In the case of receiving an IPMI instruction, the IPMI instruction can be parsed, and in response to the IPMI instruction, the central processor series and the number of central processors on the current single board can be read, and then according to the MCE registers supported by the current central processor, the numerical contents of all MCE registers in the table can be read through the target protocol and output and recorded into the second log file. After the second log file is read, based on the fault information in the second log file, the fault can be repaired.

[0136] According to the fault repair method provided by the embodiments of the present application, by manually triggering the collection of RAS logs, it is possible to manually trigger the collection of RAS logs after a fault is discovered even when the watchdog policy or caterr signal cannot monitor, which broadens the applicable scenarios and enhances the reliability and maintainability of the system.

[0137] Next, a computer program product provided by the present application will be described, and the computer program product described below can be mutually corresponding and referred to the fault repair method described above.

[0138] For the fault repair method provided by the embodiments of the present application, the execution subject can be a computer program product. In the embodiments of the present application, taking the computer program product executing the fault repair method as an example, the computer program product provided by the embodiments of the present application is described.

[0139] The embodiments of the present application also provide a computer program product.

[0140] As Figure 5 shown, the computer program product includes: a first processing module 510, a second processing module 520, and a third processing module 530.

[0141] The first processing module 510 is configured to start a first watchdog corresponding to the basic input / output system startup phase, a second watchdog corresponding to the operating system startup phase, and a third watchdog corresponding to the operating system running phase respectively based on the startup sequence corresponding to the basic input / output system and the operating system, and obtain the running states of the first watchdog, the second watchdog, and the third watchdog; The second processing module 520 is configured to obtain a first log file corresponding to the target watchdog when it is determined that the target watchdog times out based on the running states of the first watchdog, the second watchdog, and the third watchdog; The third processing module 530 is configured to perform fault repair on the fault information in the trigger phase based on the fault information in the trigger phase corresponding to the target watchdog in the first log file.

[0142] According to the computer program product provided by the embodiments of the present application, by setting different watchdogs in different phases, then determining the trigger phase corresponding to the watchdog that times out according to the running states of the watchdogs in each phase, and obtaining the first log file corresponding to the watchdog, so as to repair the fault information corresponding to the trigger phase based on the first log file, it is possible to automatically collect the first log file, thereby automatically locating the fault problem based on the first log file to automatically repair the system fault, without manual guessing or reproduction for location and repair, saving a large amount of labor costs and improving the efficiency of fault location and fault repair.

[0143] In some embodiments, the first processing module 510 may further be configured to: Start the first watchdog when the central processing unit is powered on; Turn off the first watchdog and start the second watchdog when the basic input / output system startup is completed; Turn off the second watchdog and start the third watchdog when the operating system startup is completed.

[0144] In some embodiments, the first processing module 510 may further be configured to: Start the first watchdog when the central processing unit is powered on; Determine that the running state of the first watchdog is off when the first shutdown signal is received within the first time period; Determine that the running state of the first watchdog times out when the first shutdown signal is not received within the first time period.

[0145] In some embodiments, the first processing module 510 may further be configured to: Obtain a first startup duration corresponding to the basic input / output system when the basic input / output system is in a full-load state; Determine the first time period as the first multiple of the first startup duration; the first multiple is a value greater than 1.

[0146] In some embodiments, the first processing module 510 may further be configured to: Start the second watchdog when the basic input / output system starts up successfully; Determine that the running state of the second watchdog is off when a second shutdown signal is received within a second time period; Determine that the running state of the second watchdog is timed out when a second shutdown signal is not received within a second time period.

[0147] In some embodiments, the first processing module 510 may further be configured to: Start the third watchdog when the operating system starts up successfully; During the operation of the operating system, determine that the running state of the third watchdog is normal when a dog feeding signal is detected within a third time period; the dog feeding signal is sent by the operating system to the third watchdog based on a target dog feeding period; During the operation of the operating system, determine that the running state of the third watchdog is timed out when a dog feeding signal is not detected within a fourth time period; the fourth time period is greater than the third time period.

[0148] In some embodiments, the second processing module 520 may further be configured to: Obtain the marker information corresponding to the target watchdog; Based on the marker information, obtain the first log file corresponding to the target watchdog.

[0149] In some embodiments, the second processing module 520 may further be configured to: When it is determined that the central processing unit is powered on, determine all the registers corresponding to the central processing unit; Based on the target protocol, obtain the numerical content of all the registers and write the numerical content into the first log file; Obtain the first log file corresponding to the target watchdog.

[0150] In some embodiments, the computer program product may further include a fourth processing module, configured to, before starting the first watchdog corresponding to the basic input / output system startup phase, the second watchdog corresponding to the operating system startup phase, and the third watchdog corresponding to the operating system running phase respectively based on the startup sequence corresponding to the basic input / output system and the operating system, start the programmable logic device and the baseboard management controller respectively when the single board is powered on; Start the central processing unit when the baseboard management controller starts up successfully.

[0151] In some embodiments, the fourth processing module may further be configured to: When it is detected that the single board is powered on for the first time, start the central processing unit.

[0152] In some embodiments, the fourth processing module can also be used for: When receiving the heartbeat signal sent by the baseboard management controller, start the central processing unit.

[0153] In some embodiments, the computer program product may further include a fifth processing module for: Receive the first input from the user; In response to the first input, obtain a second log file; Based on the fault information in the second log file, perform fault repair on the fault information.

[0154] The computer program product in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than the terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0155] The computer program product in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0156] The computer program product provided by the embodiments of the present application can implement Figures 1 to 4 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0157] In some embodiments, such as Figure 6As shown in the figure, an embodiment of the present application further provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, it implements each process of the above-mentioned embodiment of the fault repair method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0158] It should be noted that the electronic device in the embodiment of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.

[0159] On the other hand, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the fault repair method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0160] On yet another aspect, an embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or an instruction to implement each process of the above-mentioned embodiment of the fault repair method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0161] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A fault repair method, characterized in that, Including: Based on the startup sequences corresponding to the Basic Input / Output System (BIOS) and the operating system, respectively start the first watchdog corresponding to the BIOS startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage, and obtain the running states of the first watchdog, the second watchdog, and the third watchdog; When it is determined that the target watchdog times out based on the running states of the first watchdog, the second watchdog, and the third watchdog, obtain the first log file corresponding to the target watchdog; Based on the fault information of the trigger stage corresponding to the target watchdog in the first log file, perform fault repair on the fault information of the trigger stage.

2. The fault repair method according to claim 1, wherein The step of respectively starting the first watchdog corresponding to the BIOS startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage based on the startup sequences corresponding to the BIOS and the operating system includes: When the central processing unit is powered on, start the first watchdog; When the BIOS startup is completed, turn off the first watchdog and start the second watchdog; When the operating system startup is completed, turn off the second watchdog and start the third watchdog.

3. The fault repair method according to claim 1, wherein The step of obtaining the running states of the first watchdog, the second watchdog, and the third watchdog includes: When the central processing unit is powered on, start the first watchdog; When the first shutdown signal is received within the first time period, determine that the running state of the first watchdog is off; When the first shutdown signal is not received within the first time period, determine that the running state of the first watchdog is timed out.

4. The fault repair method according to claim 3, wherein The first time period is determined based on the following steps: Obtain the first startup time period corresponding to the BIOS when the BIOS is in a full-load state; Determine the first time period as the first startup time period multiplied by a first multiple; The first multiple is a value greater than 1.

5. The fault repair method according to any one of claims 1-4, characterized in that, The step of obtaining the running states of the first watchdog, the second watchdog, and the third watchdog includes: When the BIOS startup is completed, start the second watchdog; When the second shutdown signal is received within the second time period, determine that the running state of the second watchdog is off; When the second shutdown signal is not received within the second time period, determine that the running state of the second watchdog is timed out.

6. The fault repair method according to any one of claims 1-4, characterized in that, The step of obtaining the running states of the first watchdog, the second watchdog, and the third watchdog includes: When the operating system startup is completed, start the third watchdog; During the running of the operating system, when a dog-feed signal is detected within the third time period, determine that the running state of the third watchdog is normal; the dog-feed signal is sent by the operating system to the third watchdog based on a target dog-feed cycle. During the operation of the operating system, when the watchdog signal is not detected within a fourth time period, it is determined that the operating state of the third watchdog is timed out; the fourth time period is greater than the third time period.

7. The fault repair method according to any one of claims 1-4, characterized in that, The obtaining of the first log file corresponding to the target watchdog includes: Obtaining the marking information corresponding to the target watchdog; Based on the marking information, obtaining the first log file corresponding to the target watchdog.

8. The fault repair method according to any one of claims 1-4, characterized in that The obtaining of the first log file corresponding to the target watchdog includes: When it is determined that the central processing unit is powered on, determining all registers corresponding to the central processing unit; Based on the target protocol, obtaining the numerical content of all the registers and writing the numerical content into the first log file; Obtaining the first log file corresponding to the target watchdog.

9. The fault repair method according to any one of claims 1-4, characterized in that, Before starting the first watchdog corresponding to the basic input / output system startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage respectively based on the startup sequences of the basic input / output system and the operating system, the method further includes: When the single board is powered on, starting the programmable logic device and the baseboard management controller respectively; When the baseboard management controller starts up successfully, starting the central processing unit.

10. The fault repair method according to claim 9, characterized in that The starting of the central processing unit includes: When it is detected that the single board is powered on for the first time, starting the central processing unit.

11. The fault repair method according to claim 9, wherein, The starting of the central processing unit includes: When receiving the heartbeat signal sent by the baseboard management controller, starting the central processing unit.

12. The fault repair method according to any one of claims 1-4, characterized in that, The method further includes: Receiving a first input from the user; In response to the first input, obtaining a second log file; Based on the fault information in the second log file, performing fault repair on the fault information.

13. A computer program product, characterized in that, Including: A first processing module, configured to start the first watchdog corresponding to the basic input / output system startup stage, the second watchdog corresponding to the operating system startup stage, and the third watchdog corresponding to the operating system running stage respectively based on the startup sequences of the basic input / output system and the operating system, and obtaining the operating states of the first watchdog, the second watchdog, and the third watchdog; A second processing module, configured to obtain the first log file corresponding to the target watchdog when it is determined that the target watchdog is timed out based on the operating states of the first watchdog, the second watchdog, and the third watchdog; A third processing module, configured to perform fault repair on the fault information of the trigger stage based on the fault information of the trigger stage corresponding to the target watchdog in the first log file.

14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the fault repair method according to any one of claims 1-12.

15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fault repair method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Log information collection method, device and equipment and readable storage medium

    CN110134540A

  • Monitoring method and system for preventing startup jamming, electronic equipment and medium

    CN116010211A

  • Method and system for eliminating faults of server

    CN1917446A

Cited By

  • Storage system process management method, electronic equipment, storage medium and program product

    CN120723586A