Fault processing method and device and readable storage medium
By centrally handling the fault processing tasks of multiple boards in the fault processing device of embedded devices, the problem of large workload and low efficiency during the configuration process is solved, and more efficient configuration and more stable and reliable board operation is achieved.
Patent Information
- Application Number
- CN202311597499.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
During the configuration of embedded devices, it is necessary to configure each board separately for troubleshooting functions, resulting in large configuration workload and low efficiency.
By centrally handling the fault processing tasks of multiple boards in the fault processing device, you only need to equip the fault processing device with a fault processing function, and during the configuration process, you only need to configure the fault processing device with a fault processing function.
It reduces the workload during the configuration process, improves configuration efficiency, and improves the stability and reliability of the board.
Smart Images

Figure CN120045390A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and particularly to a fault handling method, apparatus, and readable storage medium. Background Art
[0002] With the development of communication technologies, the application scope of embedded devices has become increasingly wide. An embedded device is composed of multiple embedded boards (usually referred to as boards) connected by communication, and is used to implement relatively complex functions. An embedded device such as a power control and protection device is generally composed of multiple boards connected by communication. The multiple boards cooperate with each other to monitor, control, and protect the power grid.
[0003] To improve the reliability of an embedded device, a corresponding fault handling function is usually configured for each board, so that when each board detects its own fault during operation, it can take corresponding measures to handle the fault. Although the method of configuring a fault handling function for each board can improve the reliability of the embedded device, during the configuration process, corresponding configurations need to be made for the fault handling functions of each board respectively, resulting in a relatively large workload and low configuration efficiency during the configuration process. Summary of the Invention
[0004] Embodiments of this application provide a fault handling method, apparatus, and readable storage medium, which can reduce the workload during the configuration process of an embedded device and improve the configuration efficiency.
[0005] In a first aspect, a fault handling method is provided, including:
[0006] Receiving a fault notification sent by a first board in an embedded device in a fault situation;
[0007] In response to the fault notification, determining a fault handling strategy corresponding to the first board from a policy configuration;
[0008] Executing the fault handling strategy to handle the fault of the first board.
[0009] In embodiments of this application, a fault notification sent by a first board in an embedded device in a fault situation can be received, then in response to the fault notification, a fault handling strategy corresponding to the first board is determined, and then the fault handling strategy is executed to handle the fault of the first board. In this way, the fault handling tasks of multiple boards can be centralized in a fault handling device, so only a fault handling function needs to be configured for the fault handling device. Therefore, during the configuration process of the embedded device, only corresponding configurations need to be made for the fault handling function of the fault handling device, thereby reducing the workload during the configuration process and improving the configuration efficiency.
[0010] In some embodiments, determining the fault handling policy corresponding to the first board from the policy configuration includes: determining the fault handling policy from the policy configuration according to the fault log of the first board.
[0011] In the embodiments of the present application, in the process of determining the fault handling policy of the first board, determining the corresponding fault handling policy according to the fault log of the first board can determine the fault handling policy that matches the fault of the first board, so that the fault of the first board can be processed more accurately, and further the stability and reliability of the first board can be improved.
[0012] In some embodiments, determining the fault handling policy from the policy configuration according to the fault log of the first board includes: determining the fault location of the first board according to the fault log; determining the fault handling policy from the policy configuration according to the fault location.
[0013] In the embodiments of the present application, when determining the fault handling policy according to the fault log at the time of the fault of the first board, the fault handling policy that matches the fault location of the first board can be determined, so that the fault of the first board can be accurately processed according to the fault handling policy.
[0014] In some embodiments, determining the fault handling policy from the policy configuration according to the fault location includes: when the fault location is a functional module in the first board, determining the fault handling policy corresponding to the functional module from the policy configuration.
[0015] In the embodiments of the present application, when it is determined according to the fault log that the fault location of the first board is a functional module in the first board, determining the fault handling policy corresponding to the functional module from the policy configuration can perform refined processing on the functional module with a fault in the first board, so that the fault can be processed more accurately.
[0016] In some embodiments, determining the fault handling policy corresponding to the functional module from the policy configuration includes: determining the fault type of the functional module according to the fault log; determining the fault handling policy corresponding to the fault type of the functional module from the policy configuration.
[0017] In the embodiments of the present application, determining the corresponding fault handling policy according to the fault location and fault type of the first board to process the fault of the first board can accurately process the fault of the first board, so that the stability and reliability of the first board can be improved.
[0018] In some embodiments, determining the fault handling policy from the policy configuration according to the fault log of the first board card includes: determining the fault type of the first board card according to the fault log; and determining the fault handling policy from the policy configuration according to the fault type of the first board card.
[0019] In the embodiments of the present application, different fault handling policies are set for different fault types of the first board card. During the fault handling process, the fault type of the first board card can be determined according to the fault log of the first board card, and then the corresponding fault handling policy can be determined according to the fault type of the first board card to handle the fault of the first board card. In this way, a relatively accurate fault handling policy can be determined, and then the fault of the first board card can be accurately handled according to the fault handling policy, improving the stability and reliability of the first board card.
[0020] In some embodiments, determining the fault handling policy from the policy configuration according to the fault type of the first board card includes: determining the fault handling policy from the policy configuration according to the fault type of the first board card and the state of the embedded device.
[0021] In the embodiments of the present application, the second board card determines the fault handling policy according to the fault type of the first board card and the state of the embedded device. During the execution of the fault handling policy, the fault handling process can be adapted to the state of the embedded device, so as to reduce the impact of the fault handling process of the first board card on the embedded device and coordinate the various boards in the embedded device.
[0022] In some embodiments, determining the fault handling policy from the policy configuration according to the fault type of the first board card and the state of the embedded device includes: determining the fault handling policy from the policy configuration according to the state of other boards and the fault type of the first board card, where the other boards include the boards associated with the first board card in the embedded device.
[0023] In the embodiments of the present application, the fault handling policy is determined according to the fault type of the first board card and the states of other boards associated with the first board card. During the fault handling process, the first board card and the associated other boards can be coordinated, the impact of the fault handling process of the first board card on the associated other boards can be reduced, and further the coordination between the first board card and the associated other boards can be improved.
[0024] In some embodiments, before determining the fault handling policy from the policy configuration according to the fault log of the first board card, the method further includes: obtaining the fault log from the first board card; or obtaining the fault log from the fault notification.
[0025] In the embodiments of the present application, the fault log can be obtained from the first board card or from the fault notification, so as to determine the fault handling strategy corresponding to the first board card according to the fault log, and thus the fault handling strategy corresponding to the first board card can be quickly determined according to the fault log.
[0026] In some embodiments, the execution of the fault handling strategy includes: when the fault handling strategy is to restart the embedded device, controlling each board card in the embedded device to restart; or, when the fault handling strategy is to restart the first board card, sending a first restart instruction to the first board card to control the first board card to restart; or, when the fault handling strategy is to restart the function module in the first board card, sending a second restart instruction to the first board card to cause the first board card to restart the function module.
[0027] In the embodiments of the present application, the fault handling strategy can indicate to restart the embedded device, the first board card or the function module in the first board card, so that the fault handling process can meet different fault handling requirements, thereby improving the flexibility in the fault handling process.
[0028] In some embodiments, the method further includes: sending the fault log to the host computer, so that the host computer performs fault analysis on the first board card according to the fault log.
[0029] In the embodiments of the present application, when the second board card sends the fault log to the host computer, it is convenient for the host computer to analyze the fault cause and fault location of the first board card, and it is convenient for the user to check, maintain and update the first board card according to the fault cause and fault location of the first board card.
[0030] In some embodiments, the fault handling method is applied to the second board card in the embedded device, and determining the fault handling strategy corresponding to the first board card from the policy configuration includes: determining the fault handling strategy from the policy configuration stored in the second board card.
[0031] In the embodiments of the present application, storing the fault handling strategy of the first board card in the policy configuration of the second board card can facilitate the second board card to quickly execute the fault handling strategy corresponding to the first board card to handle the fault of the first board card when the first board card fails.
[0032] In a second aspect, a fault handling device is provided, including:
[0033] a receiving module, configured to receive a fault notification sent by a first board card in an embedded device in case of a fault;
[0034] A determination module, configured to determine a fault handling policy corresponding to the first board from a policy configuration in response to the fault notification;
[0035] An execution module, configured to execute the fault handling policy to handle the fault of the first board.
[0036] In some embodiments, the fault handling device further includes: an acquisition module, configured to acquire a fault log from the first board; or acquire a fault log from the fault notification.
[0037] In some embodiments, the fault handling device further includes: a sending module, configured to send the fault log to a host computer, so that the host computer performs fault analysis on the first board according to the fault log.
[0038] In some embodiments, the fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.
[0039] In a third aspect, an embedded device is provided, including a first board and a second board, where the first board and the second board are communicatively connected through a backplane, and the fault handling device as described in the second aspect is configured on the second board.
[0040] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program runs on a fault handling device, the fault handling device is caused to execute the fault handling method described in the first aspect above.
[0041] In a fifth aspect, a fault handling device is provided, including: a processor; a memory; and a computer program, where the computer program is stored in the memory. When the computer program is executed by the processor, the fault handling device is caused to execute the fault handling method described in the first aspect above.
[0042] In a sixth aspect, a computer program product is provided, including: computer program code. When the computer program code runs on a fault handling device, the fault handling device is caused to execute the fault handling method described in the first aspect above.
[0043] In a seventh aspect, a chip is provided, including: a processor, configured to call and run a computer program from a memory, so that a fault handling device installed with the chip executes the fault handling method described in the first aspect above.
[0044] Understandably, the readable storage medium provided in the fourth aspect above, the fault handling devices provided in the second, third, and fifth aspects, the computer program product provided in the fifth aspect, and the chip provided in the sixth aspect are all used to execute the fault handling method described in the first aspect above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be elaborated here. Description of the Drawings
[0045] Figure 1 FIG. shows a schematic diagram of an application scenario of a fault handling method provided by an embodiment of the present application.
[0046] Figure 2 FIG. shows a schematic diagram of the system structure of an embedded device provided by an embodiment of the present application.
[0047] Figure 3 FIG. shows a flowchart of the steps of a fault handling method provided by an embodiment of the present application.
[0048] Figure 4 FIG. shows a schematic diagram of the process of a fault handling method provided by an embodiment of the present application.
[0049] Figure 5 FIG. shows a schematic diagram of the process of a fault handling method provided by an embodiment of the present application.
[0050] Figure 6 FIG. shows a schematic diagram of the process of another fault handling method provided by an embodiment of the present application.
[0051] Figure 7 FIG. shows a block diagram of the structure of a fault handling device provided by an embodiment of the present application.
[0052] Figure 8 FIG. shows a block diagram of the structure of another fault handling device provided by an embodiment of the present application. Detailed Embodiments
[0053] Next, the technical solutions in the present application will be described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0054] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0055] As used herein, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized. Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0056] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, both A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.
[0057] An embedded device refers to a system composed of multiple boards for implementing complex functions. The boards can be communicatively connected by wired or wireless communication. The boards can also be referred to as embedded systems, intelligent boards, etc. For example, the embedded device is a power control and protection device, which includes one or more boards for monitoring the voltage, current, power, etc. of the power grid, one or more boards for controlling the on / off of the lines in the power grid, and a board for communicating with the control center. The boards in the power control and protection device are communicatively connected via a bus. It should be understood that the embedded device may include but is not limited to the above examples.
[0058] In the current technology, in order to improve the reliability of the embedded device, a corresponding fault handling function is usually equipped for each board, so that each board can take corresponding measures for fault handling when it detects its own fault during operation. Although the method of equipping each board with a fault handling function can improve the reliability of the embedded device, during the configuration process of the embedded device, corresponding configurations need to be made for the fault handling functions of each board respectively, resulting in a relatively large workload and low configuration efficiency during the configuration process.
[0059] For example, in order for a board to have a fault handling function, it is necessary not only to configure a software module for fault handling for the board, but also to configure at least one fault handling strategy for the board. When a fault is detected during the operation of the board, the software module can be run to determine one of the fault handling strategies from at least one fault handling strategy based on the detected fault, and then the determined fault handling strategy is executed to handle the fault. An embedded device usually includes a relatively large number of boards. When a software module for fault handling and a fault handling strategy are separately configured for each board, the workload during the configuration process will increase and the configuration efficiency will decrease.
[0060] To solve the above technical problems, an embodiment of the present application provides a fault handling method and a fault handling device. During the operation of the embedded device, the fault handling device executes the fault handling strategy to handle the faults of the boards (hereinafter referred to as the first boards) in the embedded device. In this way, the fault handling tasks of multiple boards can be centralized in the fault handling device, and only the fault handling device needs to be equipped with the fault handling function, which can reduce the workload during the configuration process and improve the configuration efficiency.
[0061] Among them, the fault handling device may be a board (hereinafter referred to as the second board) in the embedded device. In this way, during the configuration process of the embedded device, only a software module for fault handling needs to be configured in the second board, and the fault handling strategy corresponding to each first board needs to be configured, and it is not necessary to configure a software module for fault handling and a fault handling strategy in other boards except the second board, which reduces the workload during the configuration process and improves the configuration efficiency. Of course, the fault handling device may also be a device such as a host computer or a controller communicatively connected to the embedded device.
[0062] See Figure 1 , Figure 1 shows a schematic diagram of an application scenario of a fault handling method provided by an embodiment of the present application. The embedded device in this application scenario is a power control and protection device 10, and the power control and protection device 10 includes a plurality of boards, including but not limited to Figure 1 the shown board 11, board 12, and board 13.
[0063] Among them, each board includes a main processor (central processing unit, CPU) for controlling the operation of the board and a coprocessor (also called a communication module) for communication. The coprocessor may specifically be a field programmable gate array (FPGA). It should be understood that the board also includes other components not shown, such as a storage module, a power module, and an execution element.
[0064] The power control and protection device 10 further includes a backplane bus 14, and each board is connected to the backplane bus 14 through the coprocessor it has, so as to realize the communication connection between the boards 11, 12 and 13.
[0065] Wherein, a software module (also called a fault detection module) for fault detection can be stored in the storage module on each board. During the operation of the power control and protection device 10, the main processor can obtain and run the fault detection module from the storage module to perform fault detection on the board where it is located.
[0066] Generally, one of the multiple boards is configured as the main board (also called the management board or management single board), and the others are configured as slave boards (also called application boards or application single boards). The slave boards are used to implement a specific function, and the main board is used to manage the multiple slave boards. For example, boards 12 and 13 are slave boards, which are respectively used to monitor the power grid and control the on / off of the lines in the power grid; board 11 is the main board, which is used to manage boards 12 and 13.
[0067] In some embodiments, the first board is a slave board in the embedded device, and the fault processing device can be the main board (i.e., the second board). For example, the second board can be the main board in the power control and protection device 10, then the first board is the slave board in the power control and protection device 10. Of course, the second board can also be a certain slave board in the power control and protection device 10, then the first board includes the main board and other slave boards.
[0068] Taking the second board as the main board as an example, during the operation of the power control and protection device 10, the main processor in the slave board runs the fault detection module to perform fault detection, and when a slave board fault is detected, a fault notification is sent to the main board. Correspondingly, when the main board receives the fault notification sent by the slave board, it determines and executes the fault processing strategy corresponding to the slave board to handle the fault of the slave board.
[0069] Wherein, the main processor in the main board can also perform self-check on the main board during operation, and when a fault is detected, it determines and executes the fault processing strategy of the main board to recover the fault of the main board. The main processor in the slave board can cooperate with the main board when the main processor in the main board executes the fault processing strategy corresponding to the slave board.
[0070] To facilitate the understanding of the present application, the following will introduce in detail the fault processing method and fault processing device provided by the embodiments of the present application with reference to the accompanying drawings.
[0071] See Figure 2 , Figure 2 shows a schematic diagram of the system structure of an embedded device 20 provided by an embodiment of the present application. As Figure 2As shown, the embedded device 20 includes one or more first boards 21 (i.e., slave boards), and a second board 22 (i.e., the main board, and the fault handling device is the main board). A fault detection module for performing self-checking of faults is provided in the first board 21, and a communication module for communicating with the second board 22 is provided. A communication module for communicating with the first board 21 is provided in the second board 22, and a software module (i.e., a fault handling module) for executing a fault handling strategy to handle the faults of the first board 21 is provided. It should be understood that other software modules and hardware modules may also be included in the first board 21 and the second board 22, which are not described in detail in this embodiment.
[0072] Among them, after the first board 21 is started, the fault detection module in the first board 21 starts to run, enabling the first board 21 to perform self-checking of faults, and sending a fault notification to the second board through the communication module after detecting a fault in the first board 21. Correspondingly, the second board 22 can, in response to the fault notification, determine the fault handling strategy corresponding to the first board 21 from the policy configuration, and then execute the fault handling strategy to handle the faults of the first board 21. The fault handling strategy is used to instruct the second board 22 to perform corresponding control on the first board to eliminate the faults of the first board 21 or reduce the impact of the faults on the first board 21 and the entire embedded device.
[0073] See Figure 3 , Figure 3 shows a step flowchart of a fault handling method 100 provided by an embodiment of the present application. The execution subject of the fault handling method 100 may be the second board in the above example, that is, the fault handling device. As Figure 3 shown, the fault handling method 100 may include:
[0074] Step 110, receiving a fault notification sent by a first board in an embedded device in the event of a fault.
[0075] Exemplarily, during operation, the first board may detect other software modules running in the first board except for the fault detection module to determine whether other software modules in the first board have faults. Specifically, the first board determines that a software module has a fault when it detects that the main thread of a certain software module is blocked, and determines that a software module has a fault when it detects that a certain task or process in a certain software module has not been effectively executed for a long time. It should be understood that the first board may detect some or all of the software modules in the first board, and the faults of the software modules may include but are not limited to main thread blocking, tasks and processes not being effectively executed, etc.
[0076] Exemplarily, during operation, the first board can detect the hardware modules included in the first board to determine whether a failure occurs in the hardware modules of the first board. The hardware module is, for example, a storage module, and the storage module can be a double data rate (DDR) memory that supports error checking and correcting (ECC). The first board can detect whether faults such as single-bit storage errors, double-bit storage errors, and crashes occur in the DDR memory. The hardware module is, for example, a communication module again. When the first board detects abnormal communication data, it can determine that the communication module fails. It should be understood that the first board can detect some or all of the hardware modules in the first board, and the hardware modules can include but are not limited to storage modules, communication modules, etc.
[0077] Exemplarily, during operation, the first board can also detect whether an overall failure occurs in the first board, and the overall failure is, for example, the first board crashing. For example, when a hardware watchdog can be set on the first board, when the first board detects that the hardware watchdog is not effectively reset, it can determine that an overall failure occurs in the first board. It should be understood that the method for detecting whether an overall failure occurs in the first board can include but is not limited to the above examples.
[0078] After detecting the above-mentioned software module failures, hardware module failures, and overall failures in itself, the first board can send a failure notification to the second board. The failure notification can include relevant information of the first board (such as the board identification) to notify the second board of the failure of the first board.
[0079] The above are only exemplary examples. The method for the first board to perform self-diagnosis of failures and the specific failures detected can include but are not limited to the above examples.
[0080] Step 120: In response to the failure notification, determine the failure handling policy corresponding to the first board from the policy configuration.
[0081] Step 130: Execute the failure handling policy to handle the failure of the first board.
[0082] In some embodiments, during the configuration process of the embedded device, the failure handling policy corresponding to the first board can be configured in the policy configuration of the second board. After receiving the failure notification, the second board can first determine the pre-configured failure handling policy corresponding to the first board, and then execute the failure handling policy. Figure 2For example, during the configuration process of the main board card (the second board card), the board identification of each slave board card can be sequentially stored in the configuration file of the main board card (the configuration file is used to store policy configurations), and the corresponding fault handling policy of the slave board card can be stored accordingly. After receiving a fault notification sent by a certain slave board card, the main board card can obtain the board identification of the slave board card from the fault notification, then obtain the corresponding fault handling policy from the configuration file of the main board card, and then execute the fault handling policy to handle the fault of the slave board card.
[0083] Exemplarily, the fault handling policy can be stored in the configuration file in the form of an instruction. For example, if the fault handling policy of the first board card is to restart the first board card, then the first restart instruction of the first board card can be stored in the configuration file, and the first restart instruction is used to restart the first board card. After receiving the fault notification, the second board card can obtain the corresponding first restart instruction from the configuration file according to the board identification in the fault notification, and then the second board card can send the first restart instruction to the first board card (that is, execute the fault handling policy). Correspondingly, after receiving the first restart instruction, the first board card can respond and execute the first restart instruction to restart the first board card. It should be understood that the fault handling policy can also be stored in the configuration file of the first board card in the form of identification information or description information, and the specific storage form of the fault handling policy can include but is not limited to the above examples.
[0084] In some other embodiments, the policy configuration can be pre-stored in other devices. For example, it can be stored in a host computer or a server communicatively connected to the second board card. After receiving the fault notification sent by the first board card, the second board card can forward the fault notification to the host computer. After receiving the fault notification, the host computer can obtain the corresponding fault handling policy from the pre-stored policy configuration according to the board identification in the fault notification, and then send the fault handling policy to the second board card. Then, the second board card can execute the fault handling policy.
[0085] In some other embodiments, after receiving the fault notification sent by the first board card, the second board card can first determine the state of the embedded device at the current moment, and then determine the fault handling policy corresponding to the first board card according to the state of the embedded device. For example, assume that restarting the first board card has no impact on the embedded device when the embedded device is in the sleep state and can solve the fault that occurs to the first board card. Then, the fault handling policy of the first board card when the embedded device is in the sleep state can be set to restart the first board card. During the configuration process of the second board card, the board identification of the first board card can be stored in the configuration file of the second board card, and the corresponding fault handling policy of the first board card is stored as restarting the first board card. At the same time, the state of the embedded device can be stored as the sleep state accordingly.
[0086] During the operation of the embedded device, after receiving the fault notification sent by the first board, the second board first determines the board identification of the first board from the fault notification and determines the state of the embedded device at the current moment, such as the sleep state. Then, the second board can determine the fault handling strategy corresponding to the board identification of the first board and the sleep state of the embedded device (i.e., restart the first board) from the configuration file, and then the second board can execute the fault handling strategy to restart the first board.
[0087] In some other embodiments, after receiving the fault notification sent by the first board, the second board can determine the fault handling strategy corresponding to the first board according to the state of the first board. For example, assuming that restarting the first board has no impact on the embedded device when the first board is in the sleep state and can solve the fault that occurs in the first board, the fault handling strategy when the first board is in the sleep state can be set to restart the first board. And, assuming that restarting the software module in the first board has no impact on the embedded device when the first board is in the idle state and can solve the fault that occurs in the first board, the fault handling strategy when the first board is in the idle state can be set to restart the software module in the first board.
[0088] In this way, during the configuration process of the second board, the board identification of the first board can be stored in the configuration file of the second board, and two fault handling strategies of the first board, namely restarting the first board and restarting the software module, can be stored correspondingly. At the same time, the state corresponding to the fault handling strategy (restarting the first board) can be stored as the sleep state of the first board, and the state corresponding to the fault handling strategy (restarting the software module) can be stored as the idle state of the first board.
[0089] During the operation of the embedded device, when the first board detects that it has a fault, it can send a fault notification to the second board, and the fault notification includes the state and board identification of the first board. After receiving the fault notification sent by the first board, the second board first determines the board identification and state of the first board from the fault notification. When the state of the first board is the sleep state, the corresponding fault handling strategy can be determined from the configuration file as restarting the first board according to the board identification and the sleep state. Then, the second board can send a first restart instruction to the first board, and after receiving the first restart instruction, the first board executes the first restart instruction for a hot restart to restart the entire first board.
[0090] Similarly, when the state of the first board is the idle state, the second board can determine the corresponding fault handling strategy as restarting the software module from the configuration file according to the board identification and the idle state in the fault notification. Then, the second board can send a second restart instruction to the first board, and after receiving the second restart instruction, the first board executes the second restart instruction to restart the software module in the first board.
[0091] The above is only an exemplary example. In actual applications, different fault handling strategies can be configured for different faults of the first board. The fault handling strategies can be restarting the first board, restarting the software module in the first board, restarting the hardware module in the first board, restarting the embedded device, shutting down the software module or hardware module in the first board, and controlling the first board to enter a preset state, etc., but not limited thereto.
[0092] In the embodiments of the present application, a fault notification sent by the first board in the embedded device when a fault occurs can be received, and then in response to the fault notification, a fault handling strategy corresponding to the first board can be determined, and then the fault handling strategy can be executed to handle the fault of the first board. In this way, the fault handling tasks of multiple boards can be centralized in a single fault handling device. Therefore, only the fault handling function needs to be equipped for the fault handling device, so that only the corresponding configuration of the fault handling function of the fault handling device needs to be done during the configuration process of the embedded device, thereby reducing the workload during the configuration process and improving the configuration efficiency.
[0093] Optionally, when the fault handling method is applied to the second board in the embedded device, the policy configuration can be stored in the second board. After receiving the fault notification sent by the first board, the second board can determine the fault handling strategy from the policy configuration stored in the second board.
[0094] As described above, when pre-configuring the fault handling strategy for the first board, the fault handling strategy of each first board can be stored in the configuration file of the second board. After the second board receives the fault notification sent by the first board, in response to the fault notification, it can determine the fault handling strategy of the first board from the pre-stored policy configuration, and then execute the fault handling strategy to handle the fault of the first board.
[0095] In the embodiments of the present application, storing the fault handling strategy of the first board in the policy configuration of the second board can facilitate the second board to quickly execute the fault handling strategy corresponding to the first board to handle the fault of the first board when the first board fails.
[0096] Optionally, step 130 may include:
[0097] In the case where the fault handling strategy is to restart the embedded device, controlling each board in the embedded device to restart; or,
[0098] In the case where the fault handling strategy is to restart the first board, sending a first restart instruction to the first board to control the first board to restart; or,
[0099] When the fault handling strategy is to restart the functional module in the first board, send a second restart instruction to the first board to cause the first board to restart the functional module.
[0100] In some embodiments, the fault handling strategy may instruct to restart the entire embedded device. In this case, when the second board executes the fault handling strategy, it may first send a restart instruction to each first board in the embedded device to cause each first board to execute the restart instruction for restart. Then, the second board may automatically restart, thereby restarting the entire embedded device and realizing the reset of the entire embedded device. It should be understood that the method for controlling the restart of the embedded device may include but is not limited to the above examples.
[0101] In some other embodiments, the fault handling strategy may instruct to restart the first board. In this case, when the second board executes the fault handling strategy, it may send a first restart instruction to the first board, and the first restart instruction is used to instruct the first board to restart. After receiving the first restart instruction, the first board may restart in response to the first restart instruction to realize the reset of the first board.
[0102] In some other embodiments, when the first board detects a fault in a certain functional module (including software module and hardware module) in the first board, it may send a fault notification including the module identifier of the functional module to the second board. The second board may determine the corresponding fault handling strategy according to the module identifier included in the fault notification, and the fault handling strategy may instruct to restart the functional module in the first board. At this time, the second board may send a second restart instruction to the first board, and the second restart instruction may include the module identifier of the functional module. Correspondingly, after receiving the second restart instruction, the first board may control the corresponding functional module in the first board to restart according to the module identifier in the second restart instruction.
[0103] Alternatively, when the fault handling strategy instructs to restart all software modules in the first board, the second board may send a second restart instruction to the first board, and the second restart instruction may include the module identifiers of all software modules in the first board. After receiving the second restart instruction, the first board may control all software modules in the first board to restart according to the module identifiers of the software modules included in the second restart instruction. Similarly, when the fault handling strategy instructs to restart all hardware modules in the first board, the second board may send a second restart instruction to the first board, and the second restart instruction may include the module identifiers of all hardware modules in the first board. After receiving the second restart instruction, the first board may control all hardware modules in the first board to restart according to the module identifiers of the hardware modules included in the second restart instruction. Of course, the fault handling strategy may also instruct to restart some software modules and / or some hardware modules in the first board.
[0104] In the embodiments of the present application, the fault handling strategy may instruct to restart the embedded device, the first board, or the functional modules in the first board, so that the fault handling process can meet different fault handling requirements, thereby improving the flexibility in the fault handling process.
[0105] Optionally, step 120 may include:
[0106] Determine a fault handling strategy from the policy configuration according to the fault log of the first board.
[0107] In some embodiments, when the first board detects its own fault, it may generate a fault log according to the relevant information at the time of the fault (hereinafter referred to as fault information), and send the fault log to the fault handling device. The fault handling device may determine the corresponding fault handling strategy from the policy configuration according to the fault log. Among them, the fault log may include fault information that can locate the fault location, fault cause, and fault type of the first board obtained when the first board detects its own fault, as well as other information related to the fault of the first board. The first board may send the fault notification and the fault log to the second board simultaneously, or may send the fault notification and the fault log to the second board step by step. Alternatively, the fault log may also be included in the fault notification and sent to the fault handling device by the first board uniformly.
[0108] Exemplarily, for the software module fault in the above example, when the first board performs a fault self-check, if it detects a fault in a certain software module in the first board, it may obtain fault information such as the module identifier, configuration information, running status information of the software module, and the on-site information and function stack (i.e., the program counter (PC) running pointer) during the process operation of the software module in the software module, and store the fault information to obtain a fault log.
[0109] Similarly, for a hardware module fault, when the first board detects a fault in a certain hardware module, it may obtain fault information such as the module identifier and record information of the hardware module (such as single-bit storage error record or double-bit storage error record) that can be used for fault location, and store the fault information to obtain a fault log. For the overall fault of the first board, when the first board detects an overall fault of the first board, it may obtain the fault record of the first board, such as a watchdog exception record, and store the fault record to obtain a fault log. The above are only exemplary examples, and the specific information included in the fault log can be specifically set according to requirements, and may include but are not limited to the above examples.
[0110] Among them, after receiving the fault notification sent by the first board, the fault handling device can obtain the fault log of the first board, and then determine the fault handling strategy corresponding to the first board from the policy configuration. Exemplarily, during the configuration process of the embedded device, the board identification of the first board can be stored in the configuration file of the second board, and the fault log of the possible faults of the first board can be stored correspondingly. For example, for the overall fault of the first board, if the fault handling strategy when the first board has an overall fault is to restart the first board, the board identification of the first board can be stored in the configuration file of the second board, and the watchdog exception record of the first board can be stored correspondingly, and the fault handling strategy of restarting the first board can be stored. Similarly, for the software module fault and hardware module fault of the first board, the corresponding fault logs and fault handling strategies can be stored respectively.
[0111] During the operation of the embedded device, after receiving the fault notification sent by the first board, the second board first obtains the board identification of the first board from the fault notification, and obtains the fault log of the first board, then obtains the fault handling strategy corresponding to the board identification and fault log of the first board from the configuration file, and then the second board can execute the fault handling strategy to handle the fault of the first board.
[0112] The above is only an exemplary example, and the method for specifically determining the fault handling strategy according to the fault log may include but is not limited to the above example.
[0113] In the embodiment of the present application, during the process of determining the fault handling strategy of the first board, determining the corresponding fault handling strategy according to the fault log of the first board can determine the fault handling strategy matching the fault of the first board, so that the fault of the first board can be processed more accurately, and further the stability and reliability of the first board can be improved.
[0114] Optionally, before determining the fault handling strategy according to the fault log, the method may further include:
[0115] Obtain the fault log from the first board; or obtain the fault log from the fault notification.
[0116] Exemplarily, after receiving the fault notification sent by the first board, the second board can send a request for obtaining to the first board. Correspondingly, after receiving the request for obtaining, the first board can, in response to the request for obtaining, send the pre-stored fault log to the second board. After receiving the fault log sent by the first board, the second board can determine the fault handling strategy corresponding to the first board from the policy configuration according to the fault log.
[0117] Alternatively, when a failure occurs, the first board card can store a failure log and then send a failure notification including the failure log to the second board card. After receiving the failure notification sent by the first board card, the second board card can parse the failure notification, obtain the failure log from the failure notification, and then determine the failure handling policy corresponding to the first board card from the policy configuration according to the failure log.
[0118] The above are only exemplary examples, and the specific method for obtaining the failure log may include but is not limited to the above examples.
[0119] In the embodiments of the present application, the failure log can be obtained from the first board card or from the failure notification, so as to determine the failure handling policy corresponding to the first board card according to the failure log, which can quickly determine the failure handling policy corresponding to the first board card according to the failure log.
[0120] Optionally, the step of determining the failure handling policy from the policy configuration according to the failure log of the first board card may include:
[0121] Determine the failure location of the first board card according to the failure log;
[0122] Determine the failure handling policy from the policy configuration according to the failure location.
[0123] Exemplarily, after receiving the failure notification, the second board card can analyze the failure of the first board card according to the failure log of the first board card to determine the failure location of the first board card. Then, the failure handling policy corresponding to the first board card can be determined from the policy configuration according to the failure location of the first board card. For example, it can be preset that when the failure location of the first board card is a software module, the failure handling policy is to restart the software module in the first board card, and the failure handling policy is stored in the policy configuration of the second board card. After receiving the failure notification, if the second board card determines that the failure location of the first board card is a software module according to the failure log in the failure notification, it can determine the failure handling policy corresponding to the failure location (i.e., the software module) from the policy configuration as restarting the software module in the first board card.
[0124] For another example, it can be preset that when the failure location of the first board card is a storage module, the failure handling policy is to restart the storage module in the first board card, and the failure handling policy is stored in the policy configuration of the second board card. After receiving the failure notification, if the second board card determines that the failure location of the first board card is a storage module according to the failure log in the failure notification, it can determine the failure handling policy corresponding to the storage module (i.e., the hardware module) from the policy configuration as restarting the hardware module in the first board card.
[0125] The above are only exemplary examples, and the specific method for determining the failure location according to the failure log and determining the failure handling policy according to the failure location may include but is not limited to the above examples.
[0126] In the embodiments of the present application, when determining a fault handling strategy based on the fault log during the first board card failure, a fault handling strategy matching the fault location of the first board card can be determined, so that the fault of the first board card can be accurately handled according to the fault handling strategy.
[0127] Optionally, the step of determining a fault handling strategy from the policy configuration according to the fault location may include:
[0128] In the case where the fault location is a functional module in the first board card, determine a fault handling strategy corresponding to the functional module from the policy configuration.
[0129] In some embodiments, when it is determined according to the fault log that the fault location of the first board card is a functional module, the fault handling device may determine a fault handling strategy corresponding to the functional module from the policy configuration according to the fault location. For example, for the first board card, a fault handling strategy for each functional module failure in the first board card may be preset in advance, and the module identifier of the functional module and the corresponding fault handling strategy are stored in the policy configuration of the second board card. When the second board card determines according to the fault log of the first board card that the fault location of the first board card is a certain functional module, it may determine the corresponding fault handling strategy from the policy configuration according to the module identifier of the functional module, and then execute the fault handling strategy to handle the fault of the first board card.
[0130] In the embodiments of the present application, when it is determined according to the fault log that the fault location of the first board card is a functional module in the first board card, determining a fault handling strategy corresponding to the functional module from the policy configuration can perform refined processing on the faulty functional module in the first board card, so that the fault can be handled more accurately.
[0131] Optionally, the step of determining a fault handling strategy corresponding to the functional module from the policy configuration may include:
[0132] Determine the fault type of the functional module according to the fault log;
[0133] Determine a fault handling strategy corresponding to the fault type of the functional module from the policy configuration.
[0134] In some embodiments, when the fault handling device determines that the fault location of the first board is a functional module, it can further determine the fault type of the functional module according to the fault log, and then determine the fault handling strategy corresponding to the fault type of the functional module from the policy configuration. Exemplarily, for the three fault types of single-bit storage error, double-bit storage error, and system crash of the storage module, different fault handling strategies can be set respectively, and the fault handling strategy for each fault type is stored in the policy configuration of the second board. After the second board determines that the fault location of the first board is the storage module according to the fault log, it can further determine that the fault type of the storage module is one of single-bit storage error, double-bit storage error, and system crash according to the fault log. Then, the second board can determine the fault handling strategy corresponding to the fault type of the storage module from the policy configuration.
[0135] For another example, for the fault types such as main thread blockage of the software module, tasks or processes not being effectively executed for a long time, etc., different fault handling strategies can be set respectively, and the fault handling strategy for each fault type is stored in the policy configuration of the second board. After the second board determines that the fault location of the first board is the software module according to the fault log, it can further determine that the fault type of the software module is one of main thread blockage, tasks or processes not being effectively executed for a long time according to the fault log. Then, the second board can determine the fault handling strategy corresponding to the fault type of the software module from the policy configuration.
[0136] In the embodiments of the present application, determining the corresponding fault handling strategy according to the fault location and fault type of the first board to handle the fault of the first board can accurately handle the fault of the first board, thereby improving the stability and reliability of the first board.
[0137] Optionally, the step of determining the fault handling strategy from the policy configuration according to the fault log may include:
[0138] Determine the fault type of the first board according to the fault log;
[0139] Determine the fault handling strategy from the policy configuration according to the fault type of the first board.
[0140] In some embodiments, the possible faults of the first board can be classified, and corresponding fault handling strategies are set for different fault types respectively. After the second board obtains the fault log of the first board, it can first determine the fault type of the first board according to the fault log, and then determine the fault handling strategy according to the fault type.
[0141] Exemplarily, the faults of the software modules in the foregoing examples can be classified into the first category, the faults of the hardware modules can be classified into the second category, and the overall faults of the first board can be classified into the third category. For the faults of the first category, the fault handling strategy can be set to restart the software module in the first board; for the faults of the second category, the fault handling strategy can be set to restart the first board; for the faults of the third category, the fault handling strategy can be set to restart the embedded device. In this way, when configuring the fault handling strategy of each first board for the second board, the fault handling strategy for the faults of the first category can be stored as restarting the software module in the configuration file of the second board, the fault handling strategy for the faults of the second category can be stored as restarting the first board, and the fault handling strategy for the faults of the third category can be stored as restarting the embedded device.
[0142] During the operation of the embedded device, after the second board obtains the fault log of the first board, it can first determine the fault type of the first board according to the fault log, and then determine the fault handling strategy of the first board according to the fault type. For example, if the fault log indicates that the first board has a fault of the first category, the fault handling strategy corresponding to the fault of the first category (i.e., restart the software module) can be determined from the policy configuration, and then a second restart instruction can be sent to the first board.
[0143] Similarly, if the fault log indicates that the first board has a fault of the second category, the fault handling strategy corresponding to the fault of the second category can be determined from the policy configuration according to the fault type. If the fault log indicates that the first board has a fault of the third category, the fault handling strategy corresponding to the fault of the third category can be determined from the policy configuration according to the fault type.
[0144] In the embodiments of the present application, different fault handling strategies are set for different fault types of the first board. During the fault handling process, the fault type of the first board can be determined according to the fault log of the first board, and then the corresponding fault handling strategy can be determined according to the fault type of the first board to handle the fault of the first board. In this way, a more accurate fault handling strategy can be determined, and then the fault of the first board can be accurately handled according to the fault handling strategy, improving the stability and reliability of the first board.
[0145] Optionally, the step of determining the fault handling strategy from the policy configuration according to the fault type of the first board may include:
[0146] Determine the fault handling strategy from the policy configuration according to the fault type of the first board and the state of the embedded device.
[0147] In some embodiments, a fault handling strategy corresponding to the first board can be pre-configured in the second board in combination with the state of the embedded device and the fault type of the first board. During the fault handling process, the second board can determine the fault handling strategy corresponding to the first board according to the fault type of the first board and the state of the embedded device. For example, when the embedded device is in the running state and the first board has a first type of fault and needs to restart the entire embedded device to solve the fault, the board identification of the first board can be stored in the configuration file of the second board in advance, and the state of the embedded device is stored as the running state, and the fault type is stored as the first type of fault.
[0148] During the operation of the embedded device, after the second board receives the fault log, if it determines that the fault type of the first board is the first type of fault according to the fault log and determines that the embedded device is currently in the running state, it can determine from the configuration file that the fault handling strategy is to restart the embedded device. At this time, the second board can send a first restart instruction to each first board respectively to control each first board to restart, and after the second board sends the first restart instruction to each first board, it can automatically restart, so as to control the entire embedded device to restart.
[0149] In some embodiments, after the second board obtains the fault log and determines the type information of the fault type of the first board according to the fault log, it can send the fault type and the state of the embedded device to the server, and the server determines the fault handling strategy according to the fault type of the first board and the state of the embedded device.
[0150] The above is only an exemplary example, and the method for the second board to determine the fault handling strategy according to the fault type and the state of the embedded device may include but is not limited to the above examples.
[0151] In the embodiments of the present application, the second board determines the fault handling strategy according to the fault type of the first board and the state of the embedded device. During the execution of the fault handling strategy, the fault handling process can be adapted to the state of the embedded device, so as to reduce the impact of the fault handling process of the first board on the embedded device and coordinate each board in the embedded device.
[0152] Optionally, the step of determining the fault handling strategy from the policy configuration according to the fault type of the first board and the state of the embedded device may include:
[0153] Determine the fault handling strategy from the policy configuration according to the state of other boards and the fault type of the first board, and the other boards include the boards associated with the first board in the embedded device.
[0154] Among them, the board cards associated with the first board card can be one or more, and can include other first board cards in the pre-set embedded device, or can include the second board card.
[0155] In some embodiments, the fault handling strategy corresponding to the first board card can be pre-configured in the second board card in combination with the status of other board cards associated with the first board card in the embedded device and the fault type of the first board card. During the fault handling process, the fault handling strategy corresponding to the first board card can be determined according to the fault type of the first board card and the status of other board cards associated with the first board card. Exemplarily, if the other board card is the second board card, the operation of the second board card depends on the data provided by the first board card. In the case where the first board card has a type-three fault, when the second board card is in the idle state, the first board card can be restarted to handle the fault of the first board card. Then, the board card identifier of the first board card can be stored in the configuration file of the second board card in advance, and the fault type of the first board card can be correspondingly stored as the type-three fault type, and the board card identifier of the second board card and the status of the second board card can be stored as the idle state.
[0156] Similarly, in the case where the first board card has a type-three fault, when the second board card is in the running state, the software module in the first board card can be restarted to handle the fault of the first board card. Then, the board card identifier of the first board card can be stored in the configuration file of the second board card in advance, and the fault type of the first board card can be correspondingly stored as the type-three fault type, and the board card identifier of the second board card and the status of the second board card can be stored as the running state.
[0157] During the operation of the embedded device, after receiving the fault notification, the second board card can first determine the board card identifier of the second board card (other board card) determined from the configuration file according to the board card identifier of the first board card included in the fault notification, and then determine the current status of the second board card according to the board card identifier of the second board card. If the status of the second board card is the idle state, the corresponding fault handling strategy can be determined from the configuration file according to the idle state as restarting the first board card.
[0158] Similarly, during the operation of the embedded device, after receiving the fault notification, the second board card can first determine the board card identifier of the second board card (other board card) determined from the policy configuration according to the board card identifier of the first board card included in the fault notification, and then determine the current status of the second board card. If the status of the second board card is the running state, the corresponding fault handling strategy can be determined from the policy configuration according to the running state as restarting the software module in the first board card.
[0159] The above is only an exemplary example, and the method for the second board card to determine the fault handling strategy according to the fault type and the status of other board cards can include but is not limited to the above examples.
[0160] In the embodiments of the present application, a fault handling strategy is determined according to the fault type of the first board and the status of other boards associated with the first board. During the fault handling process, the first board can be coordinated with other associated boards to reduce the impact of the fault handling process of the first board on other associated boards, thereby improving the coordination between the first board and other associated boards.
[0161] In some embodiments, during the operation of the embedded device, the second board can also perform a self-check for faults to determine whether the second board has a fault, and can determine and execute the fault handling strategy corresponding to the second board when a self-fault is detected. The process of the second board performing a self-check for faults and determining and executing the fault handling strategy can refer to the first board, and this embodiment will not be elaborated here.
[0162] See Figure 4 , Figure 4 shows a schematic flowchart of a fault handling method 200 provided by the embodiments of the present application. The execution subject of the fault handling method can be the second board in the embedded device. As Figure 4 shown, the method may include the following steps:
[0163] Step 210: The first board performs a self-check for faults.
[0164] Step 220: When the first board detects its own fault, it sends a fault notification to the second board.
[0165] In the embodiments of the present application, during the operation of the embedded device, the first board can continuously perform a self-check for faults and send a fault notification to the second board when it detects its own fault. The process of the first board performing a self-check for faults and sending a fault notification can refer to the foregoing examples, and this embodiment will not be elaborated here.
[0166] Step 230: The second board determines the fault handling strategy corresponding to the first board.
[0167] Step 240: The second board executes the fault handling strategy.
[0168] In the embodiments of the present application, during the operation of the embedded device, after receiving the fault notification sent by the first board, the second board can directly determine the fault handling strategy corresponding to the first board from the policy configuration, or can obtain the fault log of the first board from the first board, determine the fault handling strategy corresponding to the first board from the policy configuration according to the fault log of the first board, and then execute the fault handling strategy to handle the fault of the first board. The process of the second board determining and executing the fault handling strategy can refer to the foregoing examples, and this embodiment will not be elaborated here.
[0169] See Figure 5 , Figure 5The flowchart of a fault handling method 300 provided by an embodiment of the present application is shown. The execution subject of the fault handling method may be the second board in the embedded device. As Figure 5 shown, the method may include the following steps:
[0170] Step 310: The first board performs a fault self-check.
[0171] Step 320: When the first board detects its own fault, it sends a fault notification and a fault log to the second board.
[0172] Step 330: The second board determines the fault type of the first board.
[0173] Step 340: The second board determines a first fault handling strategy corresponding to the first board according to the fault type of the first board.
[0174] In this embodiment, when the first board has a fault, it can send a fault notification and a fault log to the second board. The second board can first determine the fault type of the first board according to the fault log of the first board, and then determine the fault handling strategy corresponding to the first board (i.e., the first fault handling strategy) according to the fault type of the first board.
[0175] Step 350: The second board executes the first fault handling strategy.
[0176] Step 360: The second board performs a fault self-check.
[0177] Step 370: The second board determines the fault type of the second board.
[0178] Step 380: The second board determines a second fault handling strategy corresponding to the second board according to the fault type of the second board.
[0179] Step 390: The second board executes the second fault handling strategy.
[0180] In an embodiment of the present application, during the operation of the embedded device, the second board can also perform a fault self-check, and when it detects a fault in the second board, it determines a fault handling strategy corresponding to the second board (i.e., the second fault handling strategy), and then executes the second fault handling strategy to handle the fault of the second board. As Figure 2 shown, a fault detection module for performing a fault self-check can also be set in the second board 22. The fault handling module in the second board 22 can both determine the fault handling strategy corresponding to the first board (i.e., the first fault handling strategy) and execute the first fault handling strategy, and can also determine the second fault handling strategy corresponding to the second board and execute the second fault handling strategy.
[0181] It should be understood that the execution processes of steps 310 to 350 and the execution processes of steps 360 to 390 can be carried out simultaneously or step by step. The process of the second board card performing self-check on faults, determining the second fault handling strategy and executing it can refer to the first board card, which will not be elaborated in this embodiment.
[0182] See Figure 6 , Figure 6 which shows a schematic flow chart of another fault handling method 400 provided by an embodiment of the present application. The execution subject of the fault handling method can be the second board card in the embedded device. As Figure 6 shown, the method may include the following steps:
[0183] Step 401, the first board card starts self-checking for faults.
[0184] Step 402, the first board card determines whether it is a software module fault.
[0185] Step 403, the first board card determines whether it is a storage module fault.
[0186] Step 404, the first board card determines whether the hardware watchdog is abnormal.
[0187] In this embodiment, after starting self-checking for faults, the first board card can sequentially execute step 402, step 403, and step 404. In step 402, the first board card detects and determines whether the software module has a fault, and executes step 406 when the software module has a fault. On the contrary, the first board card executes step 403 when the software module has no fault.
[0188] In step 403, the first board card detects and determines whether the storage module has a fault. When it is determined that the storage module has a fault, step 405 is executed. When it is determined that the storage module has no fault, step 404 is executed. It should be understood that during the self-checking process of faults, the first board card can also detect other hardware modules except the storage module.
[0189] In step 404, the first board card detects and determines that the hardware watchdog in the first board card is abnormal, that is, determines whether the hardware watchdog is effectively reset. If the hardware watchdog is abnormal (that is, not effectively reset), step 409 is executed. Otherwise, step 401 is executed to continue the self-checking for faults.
[0190] It should be noted that after detecting and determining that the first board has a fault, the first board can also execute steps 402, 403, and 404 simultaneously to determine whether the fault of the first board is a software module fault (i.e., the first type of fault), a storage module fault (i.e., the second type of fault), or a hardware watchdog exception (i.e., the third type of fault); alternatively, steps 402, 403, and 404 can also be executed in other execution orders.
[0191] Step 405: Determine whether it is a single-bit error.
[0192] In step 405, the first board determines whether the fault of the storage module (i.e., DDR memory) is a single-bit storage error. If it is a single-bit storage error, step 407 is executed; if it is not a single-bit storage error, step 408 is executed.
[0193] Step 406: Obtain software fault information.
[0194] Step 407: Obtain single-bit error information.
[0195] Step 408: Obtain double-bit error information.
[0196] Step 409: Obtain hardware watchdog exception information.
[0197] In this embodiment, during the execution of step 406, the first board obtains the software fault information of the software module. The software fault information includes the module identifier, configuration information, running status information, scene information, function stack, etc. as exemplified above. During the execution of step 407, the first board obtains the single-bit storage error information. During the execution of step 408, the first board obtains the double-bit storage error information. And during the execution of step 409, the first board obtains the hardware watchdog exception information. Among them, the software fault information, single-bit storage error information, double-bit storage error information, and hardware watchdog exception information belong to the information in the fault log.
[0198] Step 410: The first board sends a fault notification to the second board.
[0199] Among them, the first board can send the fault log to the second board simultaneously when sending the fault notification to the second board, or can send the fault log to the second board after sending the fault notification to the second board, or the fault log can also be directly included in the fault notification.
[0200] Step 411: The second board determines the fault handling strategy.
[0201] Optionally, the second board is also used to send the fault log to the host computer so that the host computer can perform fault analysis on the first board according to the fault log.
[0202] In some embodiments, after receiving the fault log sent by the first board, the second board may also send the fault log to the host computer. The host computer can comprehensively and accurately analyze the fault of the first board based on the fault log to determine the fault location and / or fault cause, etc. of the first board, and output the fault location and / or fault cause to the user, so as to facilitate the user to perform operations such as inspection, maintenance, and update on the first board according to the fault location and / or fault cause of the first board. In practical applications, when the second board sends the fault log to the host computer, it can facilitate the host computer to analyze the fault cause and fault location, etc. of the first board, and facilitate the user to perform inspection, maintenance, and update on the first board according to the fault cause and fault location of the first board.
[0203] Step 412: The second board determines whether to restart the function module.
[0204] Step 413: The second board determines whether to restart the first board.
[0205] Step 414: The second board determines whether to restart the embedded device.
[0206] In this embodiment, after determining the fault handling strategy, the second board first executes step 412 to determine whether the fault handling strategy indicates restarting the function module (software module and / or hardware module) in the first board. If it is determined that the fault handling strategy indicates restarting the function module in the first board, step 415 is executed. If it is determined that the fault handling strategy does not indicate restarting the function module in the first board, step 413 is executed.
[0207] In step 413, the second board determines whether the fault handling strategy indicates restarting the first board. If the fault handling strategy indicates restarting the first board, step 416 is executed. If the fault handling strategy does not indicate restarting the first board, step 414 is executed.
[0208] In step 414, the second board determines whether the fault handling strategy indicates restarting the embedded device. If the fault handling strategy indicates restarting the embedded device, step 417 is executed. If the fault handling strategy does not indicate restarting the embedded device, step 418 is executed.
[0209] It should be noted that after determining the fault handling strategy, the second board can execute step 412, step 413, and step 414 simultaneously, or can execute step 412, step 413, and step 414 in other orders.
[0210] Step 415: The second board notifies the first board to restart the function module.
[0211] Step 416: The second board controls the first board to restart.
[0212] Step 417: The second board controls the embedded device to restart.
[0213] For the understanding of Steps 415, 416, and 417, reference can be made to the foregoing examples, and details are not described herein in this embodiment.
[0214] Step 418: The second board records the fault handling information.
[0215] In this embodiment, after executing the fault handling strategy, the second board can record the fault handling information, so that the staff can determine the fault handling process of the embedded device according to the fault handling information. Exemplarily, the fault handling information may include the board identification of the first board where the fault occurs, the fault type, fault location, and fault time of the first board, and the executed fault handling strategy, etc., but is not limited thereto.
[0216] In some embodiments, during the operation of the embedded device, the second board can also execute Steps 401 to 409 to detect and determine whether the second board has a fault and obtain the fault log at the time of the fault. At the same time, when the second board detects its own fault, it can also execute Steps 411 to 418, determine and execute the fault handling strategy corresponding to the second board, and record the fault handling information.
[0217] Figure 7 shows a structural block diagram of a fault handling device 70 provided by an embodiment of the present application. As Figure 7 shown, the fault handling device includes: a receiving module 71, a determining module 72, and an executing module 73.
[0218] The receiving module 71 is configured to receive a fault notification sent by a first board in the embedded device in case of a fault;
[0219] The determining module 72 is configured to determine, in response to the fault notification, a fault handling strategy corresponding to the first board from the policy configuration;
[0220] The executing module 73 is configured to execute the fault handling strategy to handle the fault of the first board.
[0221] In some embodiments, the determining module 72 is specifically configured to determine the fault handling strategy from the policy configuration according to the fault log of the first board.
[0222] In some embodiments, the determining module 72 is specifically configured to determine the fault location of the first board according to the fault log; and determine the fault handling strategy from the policy configuration according to the fault location.
[0223] In some embodiments, the determining module 72 is specifically configured to, when the fault location is a functional module in the first board, determine the fault handling policy corresponding to the functional module from the policy configuration.
[0224] In some embodiments, the determining module 72 is specifically configured to determine the fault type of the functional module according to the fault log; and determine the fault handling policy corresponding to the fault type of the functional module from the policy configuration.
[0225] In some embodiments, the determining module 72 is specifically configured to determine the fault type of the first board according to the fault log; and determine the fault handling policy from the policy configuration according to the fault type of the first board.
[0226] In some embodiments, the determining module 72 is specifically configured to determine the fault handling policy from the policy configuration according to the fault type of the first board and the state of the embedded device.
[0227] In some embodiments, the determining module 72 is specifically configured to determine the fault handling policy from the policy configuration according to the state of other boards and the fault type of the first board, where the other boards include the boards associated with the first board in the embedded device.
[0228] In some embodiments, the fault handling device further includes an obtaining module, configured to obtain the fault log from the first board; or obtain the fault log from the fault notification.
[0229] In some embodiments, the executing module 73 is specifically configured to, when the fault handling policy is to restart the embedded device, control each board in the embedded device to restart; or, when the fault handling policy is to restart the first board, send a first restart instruction to the first board to control the first board to restart; or, when the fault handling policy is to restart a functional module in the first board, send a second restart instruction to the first board to cause the first board to restart the functional module.
[0230] In some embodiments, the fault handling device further includes a sending module, configured to send the fault log to the host computer, so that the host computer performs fault analysis on the first board according to the fault log.
[0231] In some embodiments, the fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.
[0232] An embodiment of the present application further provides an embedded device, including a first board and a second board, where the first board and the second board are communicatively connected through a backplane, as Figure 7 the fault handling device is disposed on the second board.
[0233] Figure 8 FIG. shows a structural block diagram of another fault handling device 80 provided by an embodiment of the present application. As Figure 8 shown, the fault handling device 80 includes a processor 81 and a memory 82, and the above-mentioned various devices can be connected through one or more buses 84.
[0234] Among them, the fault handling device 80 further includes a computer program 83, and the computer program 83 is stored in the memory 82. When the computer program 83 is executed by the processor 81, the fault handling device 80 is caused to execute the above-mentioned Figures 3 to 6 method. Among them, all relevant contents of the steps involved in the above method embodiment can be cited in the function description of the corresponding physical device, and will not be elaborated here.
[0235] An embodiment of the present application further provides a readable storage medium, which includes a computer program. When it runs on a computer, the computer is caused to execute the method provided by the above method embodiment.
[0236] An embodiment of the present application further provides a computer program product including instructions. When the computer program product runs on a computer, the computer is caused to execute the method provided by the above method embodiment.
[0237] An embodiment of the present application further provides a chip system, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the board installed with the chip system executes the method provided by the above method embodiment.
[0238] Among them, the chip system may include an input circuit or interface for sending information or data, and an output circuit or interface for receiving information or data.
[0239] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0240] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0241] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0242] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0243] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0244] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0245] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0246] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer functional unit (which can be a personal computer, a server, or a network functional unit, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0247] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for troubleshooting, It is characterized in that include: receiving a fault notification sent by a first board in the embedded device in a fault condition; In response to the fault notification, determining a fault handling policy corresponding to the first board from a policy configuration; The fault handling strategy is executed to handle the fault of the first board.
2. The method according to claim 1, It is characterized in that The determining, from the policy configuration, a fault handling policy corresponding to the first board includes: The fault handling strategy is determined from the strategy configuration according to the fault log of the first board.
3. The method according to claim 2, It is characterized in that The determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: Determine the fault location of the first board according to the fault log; The fault handling strategy is determined from the strategy configuration according to the fault location.
4. The method according to claim 3, It is characterized in that The determining the fault handling strategy from the strategy configuration according to the fault location includes: In the case where the fault location is a functional module in the first board, the fault handling strategy corresponding to the functional module is determined from the strategy configuration.
5. The method according to claim 4, It is characterized in that The determining the fault handling strategy corresponding to the functional module from the strategy configuration includes: Determine the fault type of the functional module according to the fault log; The fault handling strategy corresponding to the fault type of the functional module is determined from the strategy configuration.
6. The method according to claim 2, It is characterized in that The determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: Determine the fault type of the first board according to the fault log; The fault handling strategy is determined from the strategy configuration according to the fault type of the first board.
7. The method according to claim 6, It is characterized in that Determining the fault handling strategy from the strategy configuration according to the fault type of the first board includes: The fault handling strategy is determined from the strategy configuration according to the fault type of the first board and the state of the embedded device.
8. The method according to claim 7, It is characterized in that The determining the fault handling strategy from the strategy configuration according to the fault type of the first board and the state of the embedded device comprises: The fault handling strategy is determined from the strategy configuration according to the status of other boards and the fault type of the first board, wherein the other boards include boards associated with the first board in the embedded device.
9. The method according to any one of claims 2 to 8, It is characterized in that Before determining the fault handling strategy from the strategy configuration according to the fault log of the first board, the method further includes: The fault log is obtained from the first board; or, the fault log is obtained from the fault notification.
10. The method according to any one of claims 1 to 9, It is characterized in that The executing the fault handling strategy includes: When the fault handling strategy is to restart the embedded device, control each board in the embedded device to restart; or, When the fault handling strategy is to restart the first board, a first restart instruction is sent to the first board to control the first board to restart; or When the fault handling strategy is to restart the functional module in the first board, a second restart instruction is sent to the first board so that the first board restarts the functional module.
11. The method according to any one of claims 2 to 10, It is characterized in that The method further comprises: The fault log is sent to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.
12. The method according to any one of claims 1 to 11, It is characterized in that The fault handling method is applied to the second board in the embedded device, and the determining of the fault handling strategy corresponding to the first board from the strategy configuration includes: The fault handling strategy is determined from the strategy configuration stored in the second board.
13. A fault handling device, It is characterized in that include: A receiving module, used for receiving a fault notification sent by a first board in the embedded device in a fault situation; a determination module, configured to determine, in response to the fault notification, a fault handling strategy corresponding to the first board from a strategy configuration; An execution module is used to execute the fault handling strategy to handle the fault of the first board.
14. The fault handling device according to claim 13, It is characterized in that Also includes: An acquisition module, used for acquiring a fault log from the first board; Alternatively, the fault log is obtained from the fault notification.
15. The fault handling device according to any one of claims 14, It is characterized in that Also includes: The sending module is used to send the fault log to the host computer, so that the host computer performs fault analysis on the first board according to the fault log.
16. The fault handling device according to any one of claims 13 to 15, It is characterized in that The fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.
17. An embedded device, It is characterized in that It includes a first board and a second board, the first board and the second board are connected via a backplane communication, and the fault handling device as described in any one of claims 13-16 is configured on the second board.
18. A readable storage medium, It is characterized in that The readable storage medium stores a computer program, and when the computer program runs on the fault processing device, the fault processing device executes the method according to any one of claims 1 to 12.