Fault handling method and device, and readable storage medium

By centrally handling the fault tasks of multiple boards in the fault handling device of embedded devices, the problems of large workload and low efficiency in the configuration process in the prior art are solved, and a more efficient configuration process and more stable and reliable board operation are achieved.

WO2025113286A1PCT designated stage expired Publication Date: 2025-06-05CONTEMPORARY AMPEREX FUTURE ENERGY RES INST (SHANGHAI) LTD +1

Patent Information

Application Number
PCT/CN2024/133293
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-20
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

During the configuration of embedded devices, the prior art requires the configuration of fault handling functions for each board separately, resulting in large workload and low efficiency.

Method used

By centrally handling the fault tasks of multiple boards in the fault processing device, you only need to equip the fault processing device with a fault processing function, and during the configuration process, you only need to configure the fault processing function of the fault processing device accordingly.

Benefits of technology

It reduces the workload during the configuration process, improves configuration efficiency, and improves the stability and reliability of the board.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133293_05062025_PF_FP_ABST
    Figure CN2024133293_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electronics, and provides a fault handling method and device, and a readable storage medium. The method comprises: receiving a fault notification sent when a first board in an embedded device fails; in response to the fault notification, determining a fault handling strategy corresponding to the first board; and then executing the fault handling strategy to handle the fault of the first board. In the fault handling method provided by the present application, fault handling tasks of a plurality of boards can be concentrated in one fault handling device, and therefore, it is simply necessary to provide the fault handling device with a fault handling function. Thus, during the configuration of the embedded device, it is simply necessary to perform corresponding configuration on the fault handling function of the fault handling device, thereby reducing the workload during the configuration, and improving the configuration efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Fault handling method, device and readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 27, 2023, with application number 202311597499.4 and invention name “Fault handling method, device and readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of electronic technology, and in particular to a fault handling method, device, and readable storage medium. Background Art

[0003] With the advancement of communications technology, the application of embedded devices is becoming increasingly widespread. These devices consist of multiple embedded boards (commonly referred to as boards) connected by communication channels, and are used to implement complex functions. For example, embedded devices such as power control and protection devices typically consist of multiple boards connected by communication channels. These boards work together to monitor, control, and protect the power grid.

[0004] To improve the reliability of embedded devices, each board is typically equipped with a fault handling function. This allows each board to take appropriate measures to address the fault if it detects a fault during operation. While equipping each board with a fault handling function can improve the reliability of embedded devices, the configuration process requires configuring the fault handling function for each board individually, resulting in a high workload and low efficiency.

[0005] Application Contents

[0006] The embodiments of the present application provide a fault handling method, device, and readable storage medium, which can reduce the workload during the configuration of an embedded device and improve configuration efficiency.

[0007] In a first aspect, a fault handling method is provided, comprising:

[0008] receiving a fault notification sent by a first board in the embedded device in a fault situation;

[0009] In response to the fault notification, determining a fault handling policy corresponding to the first board from a policy configuration;

[0010] The fault handling strategy is executed to handle the fault of the first board.

[0011] In an embodiment of the present application, a fault notification sent by a first board in an embedded device when a fault occurs can be received. In response to the fault notification, a fault handling strategy corresponding to the first board is determined, and then the fault handling strategy is executed to handle the fault of the first board. In this way, the fault handling tasks of multiple boards can be centralized in a single fault handling device, so that only the fault handling function needs to be equipped in the fault handling device. Therefore, during the configuration process of the embedded device, only the fault handling function of the fault handling device needs to be configured accordingly, thereby reducing the workload during the configuration process and improving configuration efficiency.

[0012] In some embodiments, determining the fault handling policy corresponding to the first board from the policy configuration includes: determining the fault handling policy from the policy configuration according to a fault log of the first board.

[0013] In an embodiment of the present application, in the process of determining the fault handling strategy of the first board, the corresponding fault handling strategy is determined based on the fault log of the first board, and a fault handling strategy that matches the fault of the first board can be determined, so that the fault of the first board can be handled more accurately, thereby improving the stability and reliability of the first board.

[0014] In some embodiments, determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: determining the fault location of the first board according to the fault log; and determining the fault handling strategy from the strategy configuration according to the fault location.

[0015] In an embodiment of the present application, when determining a fault handling strategy based on the fault log when the first board fails, a fault handling strategy that matches the fault location of the first board can be determined, so that the fault of the first board can be accurately handled according to the fault handling strategy.

[0016] In some embodiments, determining the fault handling strategy from the strategy configuration according to the fault location includes: when the fault location is a functional module in the first board, determining the fault handling strategy corresponding to the functional module from the strategy configuration.

[0017] In an embodiment of the present application, when the fault location of the first board is determined to be a functional module in the first board according to the fault log, the fault handling strategy corresponding to the functional module is determined from the policy configuration, and the functional module where the fault occurs in the first board can be processed in a refined manner, so that the fault can be handled more accurately.

[0018] In some embodiments, determining the fault handling strategy corresponding to the functional module from the strategy configuration includes: determining the fault type of the functional module based on the fault log; and determining the fault handling strategy corresponding to the fault type of the functional module from the strategy configuration.

[0019] In the embodiment of the present application, a corresponding fault handling strategy is determined according to the fault location and fault type of the first board to handle the fault of the first board. This can accurately handle the fault of the first board, thereby improving the stability and reliability of the first board.

[0020] In some embodiments, determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: determining the fault type of the first board according to the fault log; and determining the fault handling strategy from the strategy configuration according to the fault type of the first board.

[0021] In the embodiment of the present application, different fault handling strategies are set for different fault types of the first board. During the fault handling process, the fault type of the first board can be determined based on the fault log of the first board. Then, a corresponding fault handling strategy is determined based on the fault type of the first board to handle the fault of the first board. In this way, a more accurate fault handling strategy can be determined, and the fault of the first board can be accurately handled according to the fault handling strategy, thereby improving the stability and reliability of the first board.

[0022] In some embodiments, determining the fault handling policy from the policy configuration according to the fault type of the first board includes: determining the fault handling policy from the policy configuration according to the fault type of the first board and the status of the embedded device.

[0023] In an embodiment of the present application, the second board determines a fault handling strategy based on the fault type of the first board and the status of the embedded device. During the execution of the fault handling strategy, the fault handling process can be adapted to the status of the embedded device, thereby reducing the impact of the fault handling process of the first board on the embedded device and coordinating the various boards in the embedded device.

[0024] In some embodiments, determining the fault handling strategy from the policy configuration based on the fault type of the first board and the status of the embedded device includes: determining the fault handling strategy from the policy configuration based on the status of other boards and the fault type of the first board, wherein the other boards include boards associated with the first board in the embedded device.

[0025] In an embodiment of the present application, a fault handling strategy is determined based on the fault type of the first board and the status of other boards associated with the first board. During the fault handling process, the first board and the other associated boards can be coordinated to reduce the impact of the fault handling process of the first board on the other associated boards, thereby improving the coordination between the first board and the other associated boards.

[0026] In some embodiments, before determining the fault handling strategy from the strategy configuration according to the fault log of the first board, the method further includes: obtaining the fault log from the first board; or obtaining the fault log from the fault notification.

[0027] In an embodiment of the present application, a fault log can be obtained from the first board or from a fault notification to determine the fault handling strategy corresponding to the first board based on the fault log. In this way, the fault handling strategy corresponding to the first board can be quickly determined based on the fault log.

[0028] In some embodiments, executing the fault handling strategy includes: when the fault handling strategy is to restart the embedded device, controlling each board in the embedded device to restart; or, when the fault handling strategy is to restart the first board, sending a first restart instruction to the first board to control the first board to restart; or, when the fault handling strategy is to restart the functional module in the first board, sending a second restart instruction to the first board to cause the first board to restart the functional module.

[0029] In an embodiment of the present application, the fault handling strategy may instruct to restart the embedded device, the first board or the functional module in the first board, so that the fault handling process can be adapted to different fault handling requirements, thereby improving the flexibility of the fault handling process.

[0030] In some embodiments, the method further includes: sending the fault log to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.

[0031] In an embodiment of the present application, when the second board sends a fault log to the host computer, it is convenient for the host computer to analyze the fault cause and fault location of the first board, and it is convenient for the user to inspect, maintain and update the first board based on the fault cause and fault location of the first board.

[0032] In some embodiments, the fault handling method is applied to a second board in the embedded device, and determining the fault handling policy corresponding to the first board from the policy configuration includes: determining the fault handling policy from the policy configuration stored in the second board.

[0033] In an embodiment of the present application, the fault handling strategy of the first board is stored in the strategy configuration of the second board, so that the second board can quickly execute the fault handling strategy corresponding to the first board to handle the fault of the first board when the first board fails.

[0034] In a second aspect, a fault handling device is provided, comprising:

[0035] a receiving module, configured to receive a fault notification sent by a first board in the embedded device in the event of a fault;

[0036] a determination module, configured to determine, in response to the fault notification, a fault handling policy corresponding to the first board from a policy configuration;

[0037] An execution module is used to execute the fault handling strategy to handle the fault of the first board.

[0038] In some embodiments, the fault handling device further includes: an acquisition module, configured to acquire a fault log from the first board; or acquire a fault log from the fault notification.

[0039] In some embodiments, the fault handling device further includes: a sending module, configured to send the fault log to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.

[0040] In some embodiments, the fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.

[0041] In a third aspect, an embedded device is provided, comprising a first board and a second board, wherein the first board and the second board are connected via a backplane communication, and the fault handling device as described in the second aspect is configured on the second board.

[0042] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program runs on a fault handling device, the fault handling device executes the fault handling method described in the first aspect.

[0043] In a fifth aspect, a fault handling device is provided, comprising: a processor; a memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the fault handling device executes the fault handling method described in the first aspect.

[0044] In a sixth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a fault handling device, enables the fault handling device to execute the fault handling method described in the first aspect.

[0045] In a seventh aspect, a chip is provided, comprising: a processor for calling and running a computer program from a memory, so that a fault handling device equipped with the chip executes the fault handling method described in the first aspect.

[0046] It can be understood that the readable storage medium provided in the fourth aspect, the fault handling devices provided in the second, third and fifth aspects, the computer program product provided in the fifth aspect, and the chip provided in the sixth aspect are all used to execute the fault handling method described in the first aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the drawings without creative work.

[0048] FIG1 shows a schematic diagram of an application scenario of a fault handling method provided in an embodiment of the present application.

[0049] FIG2 shows a schematic diagram of the system structure of an embedded device provided in an embodiment of the present application.

[0050] FIG3 shows a flowchart of the steps of a fault handling method provided in an embodiment of the present application.

[0051] FIG4 shows a flow chart of a fault handling method provided in an embodiment of the present application.

[0052] FIG5 shows a flow chart of a fault handling method provided in an embodiment of the present application.

[0053] FIG6 shows a flow chart of another fault handling method provided in an embodiment of the present application.

[0054] FIG7 shows a structural block diagram of a fault handling device provided in an embodiment of the present application.

[0055] FIG8 shows a structural block diagram of another fault handling device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The technical solutions of this application will be described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, rather than all the embodiments.

[0057] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0058] The term "comprising" herein indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof. The terms "comprising", "including", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized. In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, unless otherwise stated, "multiple" means two or more.

[0059] The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0060] An embedded device refers to a system composed of multiple boards used to implement complex functions. The boards can communicate with each other via wired or wireless communication. Boards may also be referred to as embedded systems, smart boards, etc. For example, an embedded device is a power control and protection device. The power control and protection device includes one or more boards for monitoring the voltage, current, and power of the power grid, one or more boards for controlling the on / off of power lines in the power grid, and a board for communicating with a control center. The boards in the power control and protection device are connected via a bus. It should be understood that embedded devices may include, but are not limited to, the examples listed above.

[0061] To improve the reliability of embedded devices, current technology typically equips each board with a fault handling function. This allows each board to take appropriate action to address the fault if it detects a fault during operation. While this approach improves the reliability of embedded devices, it requires configuring the fault handling function for each board individually during configuration, resulting in a high workload and low efficiency.

[0062] For example, to enable fault handling capabilities on a board, not only does the board need to be configured with a software module for fault handling, but it also needs to be configured with at least one fault handling strategy. When a board detects a fault during operation, the software module can run, determine one of at least one fault handling strategy based on the detected fault, and then execute the determined fault handling strategy to handle the fault. Embedded devices typically include a large number of boards. Configuring a separate fault handling software module and fault handling strategy for each board increases the workload during configuration and reduces efficiency.

[0063] To address the aforementioned technical issues, embodiments of the present application provide a fault handling method and apparatus. During the operation of an embedded device, the apparatus executes a fault handling strategy to handle faults on a board (hereinafter referred to as the first board) in the embedded device. This allows the fault handling tasks for multiple boards to be centralized within the apparatus, requiring only the apparatus to be equipped with fault handling functionality. This reduces the workload during configuration and improves configuration efficiency.

[0064] The fault handling device can be a board in the embedded device (hereinafter referred to as the second board). Thus, during configuration of the embedded device, only the software module for fault handling needs to be configured in the second board, as well as the fault handling strategy corresponding to each first board. There is no need to configure the software modules and fault handling strategies for fault handling in boards other than the second board, thereby reducing the workload during configuration and improving configuration efficiency. Of course, the fault handling device can also be a host computer or controller, etc., that is communicatively connected to the embedded device.

[0065] Refer to Figure 1, which shows a schematic diagram of an application scenario of a fault handling method provided in an embodiment of the present application. The embedded device in the application scenario is a power control and protection device 10, which includes multiple boards, including but not limited to board 11, board 12 and board 13 shown in Figure 1.

[0066] Each board includes a central processing unit (CPU) for controlling board operations and a coprocessor (also called a communication module) for communication. The coprocessor can be a field programmable gate array (FPGA). It should be understood that the board also includes other components (not shown), such as a storage module, a power module, and an actuator.

[0067] The power control and protection device 10 further includes a backplane bus 14 . Each board is connected to the backplane bus 14 via its coprocessor to achieve communication between the board 11 , the board 12 and the board 13 .

[0068] The storage module on each board can store a software module for fault detection (also called a fault detection module). During the operation of the power control and protection device 10, the main processor can obtain and run the fault detection module from the storage module to perform fault detection on the board.

[0069] Typically, one of the multiple boards is configured as the master board (also known as a management board or management board), while the others are configured as slave boards (also known as application boards or application boards). The slave boards are used to implement specific functions, while the master board is used to manage the multiple slave boards. For example, boards 12 and 13 are slave boards, used to monitor the power grid and control the on / off of power lines, respectively; board 11 is the master board, used to manage boards 12 and 13.

[0070] In some embodiments, the first board is a slave board in an embedded device, and the fault handling device can be a master board (i.e., a second board). For example, the second board can be the master board in the power control and protection device 10, and the first board is a slave board in the power control and protection device 10. Of course, the second board can also be a slave board in the power control and protection device 10, and the first board includes the master board and other slave boards.

[0071] For example, if the second board is the main board, during operation of the power control and protection device 10, the main processor in the slave board runs the fault detection module to perform fault detection. Upon detecting a slave board fault, the main board sends a fault notification to the main board. In response, upon receiving the fault notification from the slave board, the main board determines and executes the fault handling strategy corresponding to the slave board to address the slave board fault.

[0072] The main processor in the mainboard can also perform fault self-tests on the mainboard during operation, and when a fault is detected, determine and execute the mainboard fault handling strategy to recover the mainboard. The main processor in the slave card can cooperate with the mainboard when the main processor in the mainboard executes the fault handling strategy corresponding to the slave card.

[0073] To facilitate understanding of the present application, the fault handling method and fault handling device provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0074] Referring to Figure 2, Figure 2 shows a schematic diagram of the system structure of an embedded device 20 provided in an embodiment of the present application. As shown in Figure 2, the embedded device 20 includes one or more first boards 21 (i.e., slave boards), and a second board 22 (i.e., main board, and the fault handling device is a main board). The first board 21 is provided with a fault detection module for performing fault self-checking, and a communication module for communicating with the second board 22. The second board 22 is provided with a communication module for communicating with the first board 21, and a software module (i.e., a fault handling module) for executing a fault handling strategy to handle the fault of the first board 21. It should be understood that the first board 21 and the second board 22 may also include other software modules and hardware modules, which will not be described in detail in this embodiment.

[0075] After the first board 21 is started, the fault detection module in the first board 21 starts running, enabling the first board 21 to perform a self-check for faults. Upon detecting a fault in the first board 21, the fault detection module sends a fault notification to the second board via the communication module. Correspondingly, the second board 22 can respond to the fault notification by determining the fault handling policy corresponding to the first board 21 from the policy configuration, and then execute the fault handling policy to address the fault in the first board 21. The fault handling policy instructs the second board 22 to perform corresponding control on the first board to eliminate the fault in the first board 21 or reduce the impact of the fault on the first board 21 and the entire embedded device.

[0076] Referring to FIG3 , FIG3 shows a flowchart of the steps of a fault handling method 100 provided in an embodiment of the present application. The execution subject of the fault handling method 100 may be the second board in the above example, i.e., the fault handling device. As shown in FIG3 , the fault handling method 100 may include:

[0077] Step 110: Receive a fault notification sent by a first board in the embedded device in a fault situation.

[0078] For example, during operation, the first board can detect other software modules running on the first board, excluding the fault detection module, to determine whether other software modules in the first board have experienced faults. Specifically, the first board can determine that a software module has failed if it detects that the main thread of a software module is blocked, or if it detects that a task or process in the software module has not been effectively executed for a long period of time. It should be understood that the first board can detect some or all software modules in the first board, and software module faults may include, but are not limited to, main thread blockage, ineffective execution of tasks or processes, etc.

[0079] For example, during operation, the first board can detect hardware modules included in the first board to determine whether a hardware module in the first board has failed. The hardware module, for example, is a storage module. The storage module can be a double data rate (DDR) memory that supports error checking and correcting (ECC). The first board can detect whether the DDR memory has single-bit storage errors, double-bit storage errors, or system crashes. Another example of a hardware module is a communication module. The first board can determine that the communication module has failed when abnormal communication data is detected. It should be understood that the first board can detect some or all of the hardware modules in the first board. The hardware modules may include, but are not limited to, storage modules and communication modules.

[0080] For example, during operation, the first board can also detect whether an overall fault has occurred on the first board. For example, a first board freeze can occur. For example, if a hardware watchdog is configured on the first board, the first board can determine that an overall fault has occurred on the first board when detecting that the hardware watchdog has not been effectively reset. It should be understood that methods for detecting an overall fault on the first board may include, but are not limited to, the aforementioned examples.

[0081] After the first board detects that it has a fault such as a software module fault, a hardware module fault, or an overall fault as mentioned in the above examples, it can send a fault notification to the second board. The fault notification may include relevant information of the first board (such as a board identification) to notify the second board of the first board fault.

[0082] The above are merely illustrative examples. The method for the first board to perform fault self-detection and the specific faults detected may include but are not limited to the above examples.

[0083] Step 120: In response to the fault notification, determine a fault handling policy corresponding to the first board from the policy configuration.

[0084] Step 130: Execute the fault handling strategy to handle the fault of the first board.

[0085] In some embodiments, during the configuration process of the embedded device, the fault handling policy corresponding to the first board can be configured in the policy configuration of the second board. After receiving the fault notification, the second board can first determine the pre-configured fault handling policy corresponding to the first board, and then execute the fault handling policy. Taking Figure 2 as an example, during the configuration process of the main board (second board), the board identification of each slave board can be stored in the configuration file of the main board (the configuration file is used to store the policy configuration) in sequence, and the fault handling policy corresponding to the slave board can be stored accordingly. After receiving the fault notification sent by a slave board, the main board can obtain the board identification of the slave board from the fault notification, and then obtain the fault handling policy corresponding to the board identification from the configuration file of the main board, and then execute the fault handling policy to handle the fault of the slave board.

[0086] Exemplarily, the fault handling strategy can be stored in a configuration file in the form of an instruction. For example, if the fault handling strategy of the first board is to restart the first board, the first restart instruction of the first board can be stored in the configuration file, and the first restart instruction is used to restart the first board. After receiving the fault notification, the second board can obtain the corresponding first restart instruction from the configuration file according to the board identifier in the fault notification, and then the second board can send the first restart instruction to the first board (i.e., execute the fault handling strategy). Correspondingly, after receiving the first restart instruction, the first board can respond and execute the first restart instruction to restart the first board. It should be understood that the fault handling strategy can also be stored in the configuration file of the first board in the form of identification information or description information, and the specific storage form of the fault handling strategy can include but is not limited to the above examples.

[0087] In other embodiments, the policy configuration may be pre-stored in another device, such as a host computer or server that is communicatively connected to the second board. After receiving the fault notification sent by the first board, the second board may forward the fault notification to the host computer. After receiving the fault notification, the host computer may retrieve the corresponding fault handling policy from the pre-stored policy configuration based on the board identifier in the fault notification, and then send the fault handling policy to the second board. The second board may then execute the fault handling policy.

[0088] In other embodiments, after receiving the fault notification sent by the first board, the second board can first determine the state of the embedded device at the current moment, and then determine the fault handling strategy corresponding to the first board according to the state of the embedded device. For example, assuming that restarting the first board when the embedded device is in a dormant state has no effect on the embedded device and can solve the fault of the first board, the fault handling strategy of the first board when the embedded device is in a dormant state can be set to restart the first board. During the configuration process of the second board, the board identifier of the first board can be stored in the configuration file of the second board, and the fault handling strategy of the first board can be stored as restarting the first board, and the state of the embedded device can be stored as a dormant state.

[0089] During operation of the embedded device, after receiving a fault notification sent by the first board, the second board first determines the board identifier of the first board from the fault notification and determines the current state of the embedded device, such as whether it is in a dormant state. The second board can then determine a fault handling strategy (i.e., restarting the first board) corresponding to the board identifier of the first board and the dormant state of the embedded device from a configuration file, and then execute the fault handling strategy to restart the first board.

[0090] In other embodiments, after receiving a fault notification sent by the first board, the second board can determine a fault handling strategy corresponding to the first board based on the state of the first board. For example, assuming that restarting the first board when the first board is in a dormant state has no impact on the embedded device and can resolve the fault of the first board, the fault handling strategy for the first board in the dormant state can be set to restart the first board. Furthermore, assuming that restarting the software module in the first board when the first board is in an idle state has no impact on the embedded device and can resolve the fault of the first board, the fault handling strategy for the first board in the idle state can be set to restart the software module in the first board.

[0091] Thus, during the configuration of the second board, the board identifier of the first board can be stored in the configuration file of the second board, and two fault handling strategies for the first board, namely, restarting the first board and restarting the software module, can be stored accordingly. Furthermore, the state corresponding to the fault handling strategy (restarting the first board) can be stored as the dormant state of the first board, and the state corresponding to the fault handling strategy (restarting the software module) can be stored as the idle state of the first board.

[0092] During the operation of the embedded device, when the first board detects that it has a fault, it can send a fault notification to the second board, and the fault notification includes the status and board identification of the first board. After receiving the fault notification sent by the first board, the second board first determines the board identification and status of the first board from the fault notification. When the state of the first board is a dormant state, the corresponding fault handling strategy can be determined from the configuration file based on the board identification and dormant state to restart the first board. Then, the second board can send a first restart instruction to the first board. After receiving the first restart instruction, the first board executes the first restart instruction to perform a hot restart, restarting the entire first board.

[0093] Similarly, when the first board is in an idle state, the second board can determine from the configuration file, based on the board identifier and idle state in the fault notification, that the corresponding fault handling strategy is to restart the software module. The second board can then send a second restart instruction to the first board. Upon receiving the second restart instruction, the first board executes the second restart instruction, restarting the software module on the first board.

[0094] The above are merely illustrative examples. In actual applications, different fault handling strategies can be configured for different faults of the first board. The fault handling strategies may include restarting the first board, restarting the software module in the first board, restarting the hardware module in the first board, restarting the embedded device, shutting down the software module or hardware module in the first board, and controlling the first board to enter a preset state, etc., but are not limited to these.

[0095] In an embodiment of the present application, a fault notification sent by a first board in an embedded device when a fault occurs can be received. In response to the fault notification, a fault handling strategy corresponding to the first board is determined, and then the fault handling strategy is executed to handle the fault of the first board. In this way, the fault handling tasks of multiple boards can be centralized in a single fault handling device, so that only the fault handling function needs to be equipped in the fault handling device. Therefore, during the configuration process of the embedded device, only the fault handling function of the fault handling device needs to be configured accordingly, thereby reducing the workload during the configuration process and improving configuration efficiency.

[0096] Optionally, when the fault handling method is applied to a second board in an embedded device, the policy configuration may be stored in the second board. After receiving the fault notification sent by the first board, the second board may determine the fault handling policy from the policy configuration stored in the second board.

[0097] As described above, when a fault handling policy is pre-configured for a first board, the fault handling policy for each first board can be stored in a configuration file of a second board. Upon receiving a fault notification sent by a first board, the second board can, in response to the fault notification, determine the fault handling policy for the first board from the pre-stored policy configuration, and then execute the fault handling policy to handle the fault of the first board.

[0098] In an embodiment of the present application, the fault handling strategy of the first board is stored in the strategy configuration of the second board, so that the second board can quickly execute the fault handling strategy corresponding to the first board to handle the fault of the first board when the first board fails.

[0099] Optionally, step 130 may include:

[0100] When the fault handling strategy is to restart the embedded device, control each board in the embedded device to restart; or,

[0101] In the case where the fault handling strategy is to restart the first board, a first restart instruction is sent to the first board to control the first board to restart; or

[0102] When the fault handling strategy is to restart the functional module in the first board, a second restart instruction is sent to the first board to enable the first board to restart the functional module.

[0103] In some embodiments, the fault handling strategy may instruct the entire embedded device to restart. In this case, when executing the fault handling strategy, the second board may first send a restart instruction to each first board in the embedded device, causing each first board to execute the restart instruction and restart. The second board may then automatically restart, thereby restarting the entire embedded device and resetting the entire embedded device. It should be understood that methods for controlling the restart of an embedded device may include, but are not limited to, the above examples.

[0104] In other embodiments, the fault handling strategy may instruct the first board to restart. In this case, when executing the fault handling strategy, the second board may send a first restart instruction to the first board, where the first restart instruction is used to instruct the first board to restart. After receiving the first restart instruction, the first board may restart in response to the first restart instruction to reset the first board.

[0105] In other embodiments, when the first board detects a fault in a functional module (including a software module and a hardware module) in the first board, it can send a fault notification including a module identifier of the functional module to the second board. The second board can determine a corresponding fault handling strategy based on the module identifier included in the fault notification, and the fault handling strategy can instruct to restart the functional module in the first board. At this time, the second board can send a second restart instruction to the first board, and the second restart instruction can include the module identifier of the functional module. Correspondingly, after receiving the second restart instruction, the first board can control the corresponding functional module in the first board to restart according to the module identifier in the second restart instruction.

[0106] Alternatively, when the fault handling strategy instructs to restart all software modules in the first board, the second board may send a second restart instruction to the first board, and the second restart instruction may include the module identifications of all software modules in the first board. After receiving the second restart instruction, the first board may control the restart of all software modules in the first board according to the module identifications of the software modules included in the second restart instruction. Similarly, when the fault handling strategy instructs to restart all hardware modules in the first board, the second board may send a second restart instruction to the first board, and the second restart instruction may include the module identifications of all hardware modules in the first board. After receiving the second restart instruction, the first board may control the restart of all hardware modules in the first board according to the module identifications of the hardware modules included in the second restart instruction. Of course, the fault handling strategy may also instruct to restart some software modules and / or some hardware modules in the first board.

[0107] In an embodiment of the present application, the fault handling strategy may instruct to restart the embedded device, the first board or the functional module in the first board, so that the fault handling process can be adapted to different fault handling requirements, thereby improving the flexibility of the fault handling process.

[0108] Optionally, step 120 may include:

[0109] A fault handling policy is determined from the policy configuration according to the fault log of the first board.

[0110] In some embodiments, when the first board detects its own fault, it can generate a fault log based on the relevant information at the time of the fault (hereinafter referred to as fault information), and send the fault log to the fault handling device. The fault handling device can determine the corresponding fault handling strategy from the policy configuration based on the fault log. Among them, the fault log may include fault information that can locate the fault location, fault cause and fault type of the first board when the first board detects its own fault, as well as other information related to the fault of the first board. The first board can send a fault notification and a fault log to the second board at the same time, or it can send a fault notification and a fault log to the second board in steps. Alternatively, the fault log can also be included in the fault notification and sent uniformly by the first board to the fault handling device.

[0111] For example, for the software module failure in the above example, when the first board performs fault self-test, if a software module failure in the first board is detected, the module identification, configuration information, running status information of the software module and the field information and function stack (i.e., program counter (PC) running pointer) of the process in the software module during execution, etc., which can be used to locate the fault, can be obtained, and the fault information can be stored to obtain a fault log.

[0112] Similarly, for hardware module failures, when the first board detects a hardware module failure, it can obtain the module identification and record information of the hardware module (such as a single-bit storage error record or a double-bit storage error record) and other fault information that can be used to locate the fault, and store the fault information to obtain a fault log. For an overall failure of the first board, when the first board detects an overall failure of the first board, it can obtain the fault record of the first board, such as a watchdog exception record, and store the fault record to obtain a fault log. The above are only illustrative examples, and the specific information included in the fault log can be set according to specific needs, and can include but is not limited to the above examples.

[0113] Among them, after receiving the fault notification sent by the first board, the fault handling device can obtain the fault log of the first board, and then determine the fault handling strategy corresponding to the first board from the strategy configuration according to the fault log. For example, in the configuration process of the embedded device, the board identification of the first board can be stored in the configuration file of the second board, and the fault log of the possible faults of the first board can be stored accordingly. For example, for the overall fault of the first board, if the fault handling strategy of the first board when the overall fault occurs is to restart the first board, the board identification of the first board can be stored in the configuration file of the second board, and the watchdog exception record of the first board can be stored accordingly, and the fault handling strategy can be stored as restarting the first board. Similarly, for the software module fault and hardware module fault of the first board, the corresponding fault log and fault handling strategy can be stored respectively.

[0114] During the operation of the embedded device, after receiving the fault notification sent by the first board, the second board first obtains the board identifier of the first board from the fault notification and obtains the fault log of the first board, and then obtains the fault handling strategy corresponding to the board identifier and fault log of the first board from the configuration file. After that, the second board can execute the fault handling strategy to handle the fault of the first board.

[0115] The above examples are merely illustrative. Specific methods for determining a fault handling strategy based on a fault log may include, but are not limited to, the above examples.

[0116] In an embodiment of the present application, in the process of determining the fault handling strategy of the first board, the corresponding fault handling strategy is determined based on the fault log of the first board, and a fault handling strategy that matches the fault of the first board can be determined, so that the fault of the first board can be handled more accurately, thereby improving the stability and reliability of the first board.

[0117] Optionally, before determining the fault handling strategy according to the fault log, the method may further include:

[0118] Obtain the fault log from the first board; or obtain the fault log from the fault notification.

[0119] For example, after receiving the fault notification sent by the first board, the second board may send a request to the first board. Correspondingly, after receiving the request, the first board may send a pre-stored fault log to the second board in response to the request. After receiving the fault log sent by the first board, the second board may determine the fault handling policy corresponding to the first board from the policy configuration based on the fault log.

[0120] Alternatively, the first board can store a fault log when a fault occurs and then send a fault notification including the fault log to the second board. After receiving the fault notification from the first board, the second board can parse the fault notification, obtain the fault log from the fault notification, and then determine the fault handling policy corresponding to the first board from the policy configuration based on the fault log.

[0121] The above examples are merely illustrative, and specific methods for obtaining fault logs may include but are not limited to the above examples.

[0122] In an embodiment of the present application, a fault log can be obtained from the first board or from a fault notification to determine the fault handling strategy corresponding to the first board based on the fault log. In this way, the fault handling strategy corresponding to the first board can be quickly determined based on the fault log.

[0123] Optionally, the step of determining a fault handling strategy from a strategy configuration according to the fault log of the first board may include:

[0124] Determine the fault location of the first board according to the fault log;

[0125] The fault handling policy is determined from the policy configuration based on the fault location.

[0126] Exemplarily, after receiving the fault notification, the second board can analyze the fault of the first board according to the fault log of the first board to determine the fault location of the first board. Then, a fault handling strategy corresponding to the first board can be determined from the policy configuration according to the fault location of the first board. For example, the fault handling strategy when the fault location of the first board is a software module can be pre-set to restart the software module in the first board, and the fault handling strategy can be stored in the policy configuration of the second board. After receiving the fault notification, if the second board determines that the fault location of the first board is a software module according to the fault log in the fault notification, then the fault handling strategy corresponding to the fault location (i.e., software module) can be determined from the policy configuration to be restarting the software module in the first board.

[0127] For another example, a fault handling policy for a first board's fault location being a storage module can be pre-set to restart the storage module in the first board, and the fault handling policy can be stored in the policy configuration of the second board. After receiving the fault notification, if the second board determines, based on the fault log in the fault notification, that the fault location of the first board is the storage module, it can determine from the policy configuration that the fault handling policy corresponding to the storage module (i.e., the hardware module) is restarting the hardware module in the first board.

[0128] The above examples are merely illustrative. Methods for determining a fault location based on a fault log and determining a fault handling strategy based on the fault location may include but are not limited to the above examples.

[0129] In an embodiment of the present application, when determining a fault handling strategy based on the fault log when the first board fails, a fault handling strategy that matches the fault location of the first board can be determined, so that the fault of the first board can be accurately handled according to the fault handling strategy.

[0130] Optionally, the step of determining a fault handling strategy from a strategy configuration according to the fault location may include:

[0131] In the case that the fault location is a functional module in the first board, a fault handling policy corresponding to the functional module is determined from the policy configuration.

[0132] In some embodiments, when the fault location of the first board is determined to be a functional module based on the fault log, the fault handling device can determine the fault handling strategy corresponding to the functional module from the policy configuration based on the fault location. For example, for the first board, a fault handling strategy for each functional module failure in the first board can be pre-set, and the module identifier of the functional module and the corresponding fault handling strategy can be stored in the policy configuration of the second board. When the second board determines that the fault location of the first board is a certain functional module based on the fault log of the first board, the second board can determine the corresponding fault handling strategy from the policy configuration based on the module identifier of the functional module, and then execute the fault handling strategy to handle the fault of the first board.

[0133] In an embodiment of the present application, when the fault location of the first board is determined to be a functional module in the first board according to the fault log, the fault handling strategy corresponding to the functional module is determined from the policy configuration, and the functional module where the fault occurs in the first board can be processed in a refined manner, so that the fault can be handled more accurately.

[0134] Optionally, the step of determining the fault handling policy corresponding to the functional module from the policy configuration may include:

[0135] Determine the fault type of the functional module based on the fault log;

[0136] A fault handling strategy corresponding to the fault type of the functional module is determined from the strategy configuration.

[0137] In some embodiments, when the fault location of the first board is determined to be a functional module, the fault handling device can further determine the fault type of the functional module based on the fault log, and then determine the fault handling strategy corresponding to the fault type of the functional module from the policy configuration. For example, different fault handling strategies can be set for the three fault types of the storage module: single-bit storage error, double-bit storage error, and system crash, and the fault handling strategy for each fault type can be stored in the policy configuration of the second board. After the second board determines that the fault location of the first board is the storage module based on the fault log, it can further determine, based on the fault log, whether the fault type of the storage module is one of single-bit storage error, double-bit storage error, and system crash. The second board can then determine the fault handling strategy corresponding to the fault type of the storage module from the policy configuration.

[0138] For another example, different fault handling strategies can be set for fault types such as a software module's main thread being blocked or a task or process not being effectively executed for a long period of time. Each fault handling strategy can be stored in the second board's policy configuration. After determining that the fault location of the first board is the software module based on the fault log, the second board can further determine from the fault log that the software module's fault type is one of the following: a main thread being blocked or a task or process not being effectively executed for a long period of time. The second board can then determine, from the policy configuration, the fault handling strategy corresponding to the software module's fault type.

[0139] In the embodiment of the present application, a corresponding fault handling strategy is determined according to the fault location and fault type of the first board to handle the fault of the first board. This can accurately handle the fault of the first board, thereby improving the stability and reliability of the first board.

[0140] Optionally, the step of determining the fault handling policy from the policy configuration according to the fault log may include:

[0141] Determine the fault type of the first board according to the fault log;

[0142] A fault handling strategy is determined from the strategy configuration according to the fault type of the first board.

[0143] In some embodiments, possible faults of the first board can be categorized, and corresponding fault handling strategies can be set for different fault types. After obtaining the fault log of the first board, the second board can first determine the fault type of the first board based on the fault log, and then determine the fault handling strategy based on the fault type.

[0144] For example, the software module failures in the aforementioned example can be categorized as the first category, the hardware module failures as the second category, and the overall failure of the first board as the third category. For the first category of failures, the fault handling strategy can be set to restart the software module in the first board; for the second category of failures, the fault handling strategy can be set to restart the first board; and for the third category of failures, the fault handling strategy can be set to restart the embedded device. Thus, when configuring the fault handling strategy for each first board for the second board, the fault handling strategy for the first category of failures can be stored in the configuration file of the second board as restarting the software module, the fault handling strategy for the second category of failures can be stored as restarting the first board, and the fault handling strategy for the third category of failures can be stored as restarting the embedded device.

[0145] During operation of the embedded device, after obtaining the fault log of the first board, the second board can first determine the fault type of the first board based on the fault log, and then determine the fault handling strategy for the first board based on the fault type. For example, if the fault log indicates that the first board has a first type of fault, the fault handling strategy corresponding to the first type of fault (i.e., restarting the software module) can be determined from the strategy configuration, and then a second restart instruction can be sent to the first board.

[0146] Similarly, if the fault log indicates that the first board has a second type of fault, the fault handling policy corresponding to the second type of fault can be determined from the policy configuration based on the fault type. If the fault log indicates that the first board has a third type of fault, the fault handling policy corresponding to the third type of fault can be determined from the policy configuration based on the fault type.

[0147] In the embodiment of the present application, different fault handling strategies are set for different fault types of the first board. During the fault handling process, the fault type of the first board can be determined based on the fault log of the first board. Then, a corresponding fault handling strategy is determined based on the fault type of the first board to handle the fault of the first board. In this way, a more accurate fault handling strategy can be determined, and the fault of the first board can be accurately handled according to the fault handling strategy, thereby improving the stability and reliability of the first board.

[0148] Optionally, the step of determining a fault handling strategy from a strategy configuration according to the fault type of the first board may include:

[0149] A fault handling policy is determined from the policy configuration according to the fault type of the first board and the status of the embedded device.

[0150] In some embodiments, a fault handling strategy corresponding to the first board can be pre-configured in the second board in combination with the state of the embedded device and the fault type of the first board. During the fault handling process, the second board can determine the fault handling strategy corresponding to the first board based on the fault type of the first board and the state of the embedded device. For example, when the embedded device is in an operating state and a first-class fault occurs on the first board, the entire embedded device needs to be restarted to resolve the fault. In this case, the board identifier of the first board can be pre-stored in the configuration file of the second board, and the state of the embedded device can be stored as an operating state and the fault type can be stored as a first-class fault.

[0151] During operation of the embedded device, after receiving the fault log, if the second board determines from the fault log that the fault type of the first board is a first type fault and that the embedded device is currently in operation, the second board can determine from the configuration file that the fault handling strategy is to restart the embedded device. In this case, the second board can send a first restart instruction to each first board respectively to control each first board to restart. After sending the first restart instruction to each first board, the second board can automatically restart, thereby controlling the entire embedded device to restart.

[0152] In some embodiments, after the second board obtains the fault log and determines the type information of the fault type of the first board based on the fault log, the fault type and the status of the embedded device can be sent to the server, and the server determines the fault handling strategy based on the fault type of the first board and the status of the embedded device.

[0153] The above are merely illustrative examples. The method for the second board to determine the fault handling strategy according to the fault type and the status of the embedded device may include but is not limited to the above examples.

[0154] In an embodiment of the present application, the second board determines a fault handling strategy based on the fault type of the first board and the status of the embedded device. During the execution of the fault handling strategy, the fault handling process can be adapted to the status of the embedded device, thereby reducing the impact of the fault handling process of the first board on the embedded device and coordinating the various boards in the embedded device.

[0155] Optionally, the step of determining the fault handling strategy from the strategy configuration according to the fault type of the first board and the state of the embedded device may include:

[0156] A fault handling strategy is determined from a strategy configuration according to states of other boards and a fault type of the first board, wherein the other boards include boards associated with the first board in the embedded device.

[0157] The board associated with the first board may be one or more boards, and may include other first boards in a pre-set embedded device, or may include a second board.

[0158] In some embodiments, the fault handling strategy corresponding to the first board can be pre-configured in the second board in combination with the status of other boards associated with the first board in the embedded device and the fault type of the first board. During the fault handling process, the fault handling strategy corresponding to the first board can be determined based on the fault type of the first board and the status of other boards associated with the first board. For example, if the other board is the second board, the operation of the second board depends on the data provided by the first board. In the case of a third type of fault on the first board, the first board can be restarted when the second board is in an idle state to handle the fault of the first board. Then, the board identifier of the first board can be stored in the configuration file of the second board in advance, and the fault type of the first board can be stored as the third type of fault correspondingly, and the board identifier of the second board and the status of the second board can be stored as the idle state.

[0159] Similarly, if a Type III fault occurs on the first board, the software module in the first board can be restarted while the second board is in the running state to address the fault. To do so, the first board's board identifier can be pre-stored in the second board's configuration file, along with the corresponding information that the first board's fault type is Type III, and the second board's board identifier and the second board's running state can be stored.

[0160] During operation of the embedded device, after receiving a fault notification, the second board may first determine the board identifier of the second board (or other board) from a configuration file based on the board identifier of the first board included in the fault notification, and then determine the current state of the second board based on the board identifier of the second board. If the state of the second board is idle, the corresponding fault handling strategy may be determined from the configuration file based on the idle state to be restarting the first board.

[0161] Similarly, during operation of the embedded device, after receiving a fault notification, the second board may first determine the board identifier of the second board (or another board) from the policy configuration based on the board identifier of the first board included in the fault notification, and then determine the current state of the second board. If the state of the second board is running, the corresponding fault handling policy may be determined from the policy configuration based on the running state to be restarting the software module in the first board.

[0162] The above examples are merely illustrative. The method for the second board to determine a fault handling strategy based on the fault type and the status of other boards may include but is not limited to the above examples.

[0163] In an embodiment of the present application, a fault handling strategy is determined based on the fault type of the first board and the status of other boards associated with the first board. During the fault handling process, the first board and the other associated boards can be coordinated to reduce the impact of the fault handling process of the first board on the other associated boards, thereby improving the coordination between the first board and the other associated boards.

[0164] In some embodiments, during operation of the embedded device, the second board may also perform a fault self-check to determine whether the second board has experienced a fault. Upon detecting a fault, the second board may determine and execute a fault handling strategy corresponding to the second board. The process of the second board performing a fault self-check and determining and executing a fault handling strategy may be referenced to that of the first board and will not be further described in this embodiment.

[0165] Referring to FIG4 , FIG4 shows a flow chart of a fault handling method 200 provided in an embodiment of the present application. The execution subject of the fault handling method may be a second board in an embedded device. As shown in FIG4 , the method may include the following steps:

[0166] Step 210: The first board performs a fault self-check.

[0167] Step 220: When the first board detects a fault, the first board sends a fault notification to the second board.

[0168] In the embodiment of the present application, during operation of the embedded device, the first board can continuously perform fault self-detection and send a fault notification to the second board when a fault is detected. The process of the first board performing fault self-detection and sending fault notification can be referred to the above example and will not be described in detail in this embodiment.

[0169] Step 230: The second board determines a fault handling strategy corresponding to the first board.

[0170] Step 240: The second board executes the fault handling strategy.

[0171] In an embodiment of the present application, during operation of the embedded device, after receiving a fault notification sent by the first board, the second board can directly determine the fault handling policy corresponding to the first board from the policy configuration, or obtain the first board's fault log from the first board, determine the fault handling policy corresponding to the first board from the policy configuration based on the first board's fault log, and then execute the fault handling policy to handle the fault of the first board. The process of the second board determining and executing the fault handling policy can be referred to the aforementioned example and will not be described in detail in this embodiment.

[0172] Referring to FIG5 , FIG5 shows a flow chart of a fault handling method 300 provided in an embodiment of the present application. The execution subject of the fault handling method may be a second board in an embedded device. As shown in FIG5 , the method may include the following steps:

[0173] Step 310: The first board performs a fault self-check.

[0174] Step 320: When the first board detects a fault on itself, the first board sends a fault notification and a fault log to the second board.

[0175] Step 330: The second board determines the fault type of the first board.

[0176] Step 340: The second board determines a first fault handling strategy corresponding to the first board according to the fault type of the first board.

[0177] In this embodiment, when a fault occurs, the first board can send a fault notification and a fault log to the second board. The second board can first determine the fault type of the first board based on the fault log of the first board, and then determine a fault handling strategy (i.e., the first fault handling strategy) corresponding to the first board based on the fault type of the first board.

[0178] Step 350: The second board executes the first fault handling strategy.

[0179] Step 360: The second board performs a fault self-check.

[0180] Step 370: The second board determines the fault type of the second board.

[0181] Step 380: The second board determines a second fault handling strategy corresponding to the second board according to the fault type of the second board.

[0182] Step 390: The second board executes the second fault handling strategy.

[0183] In an embodiment of the present application, during operation of the embedded device, the second board can also perform a fault self-check and, upon detecting a fault on the second board, determine a fault handling strategy corresponding to the second board (i.e., a second fault handling strategy), and then execute the second fault handling strategy to handle the fault on the second board. As shown in FIG2 , a fault detection module for performing a fault self-check can also be provided in the second board 22. The fault handling module in the second board 22 can either determine a fault handling strategy corresponding to the first board (i.e., a first fault handling strategy) and execute the first fault handling strategy, or determine a second fault handling strategy corresponding to the second board and execute the second fault handling strategy.

[0184] It should be understood that the execution process of steps 310 to 350 and the execution process of steps 360 to 390 can be performed simultaneously or in separate steps. The process of the second board performing fault self-detection and determining and executing the second fault handling strategy can be referred to the first board, and this embodiment will not be described in detail here.

[0185] Referring to FIG6 , FIG6 shows a flow chart of another fault handling method 400 provided in an embodiment of the present application. The execution subject of the fault handling method may be a second board in an embedded device. As shown in FIG6 , the method may include the following steps:

[0186] Step 401: The first board starts a fault self-check.

[0187] Step 402: The first board determines whether the fault is a software module fault.

[0188] Step 403: The first board determines whether the fault is in the storage module.

[0189] Step 404: The first board determines whether the hardware watchdog is abnormal.

[0190] In this embodiment, after the first board initiates the fault self-check, it may sequentially execute steps 402, 403, and 404. In step 402, the first board detects and determines whether a software module has failed. If so, step 406 is executed. Conversely, if the software module has not failed, the first board executes step 403.

[0191] In step 403, the first board detects and determines whether the storage module is faulty. If the storage module is faulty, step 405 is executed. If the storage module is not faulty, step 404 is executed. It should be understood that during the fault self-detection process, the first board may also detect other hardware modules in addition to the storage module.

[0192] In step 404, the first board detects and determines whether the hardware watchdog in the first board is abnormal, that is, whether the hardware watchdog is effectively reset. If the hardware watchdog is abnormal (that is, not effectively reset), step 409 is executed. Otherwise, step 401 is executed to continue the fault self-check.

[0193] It should be noted that after detecting and determining that the first board has a fault, the first board may also simultaneously execute step 402, step 403, and step 404 to determine whether the fault of the first board is one of the following types of faults: a software module fault (i.e., a first type of fault), a storage module fault (i.e., a second type of fault), and a hardware watchdog abnormality (i.e., a third type of fault); alternatively, steps 402, 403, and 404 may be executed in another execution order.

[0194] Step 405: Determine whether it is a single-bit error.

[0195] In step 405 , the first board determines whether the fault of the storage module (ie, DDR memory) is a single-bit storage error. If so, step 407 is executed; otherwise, step 408 is executed.

[0196] Step 406: Obtain software fault information.

[0197] Step 407: Obtain single-bit error information.

[0198] Step 408: Obtain double-bit error information.

[0199] Step 409: Obtain hardware watchdog exception information.

[0200] In this embodiment, during step 406, the first board obtains software fault information of the software module. The software fault information includes the module identification, configuration information, operating status information, field information, and function stack, as described in the aforementioned examples. During step 407, the first board obtains single-bit storage error information. During step 408, the first board obtains double-bit storage error information. Furthermore, during step 409, the first board obtains hardware watchdog exception information. Software fault information, single-bit storage error information, double-bit storage error information, and hardware watchdog exception information are included in the fault log.

[0201] Step 410: The first board sends a fault notification to the second board.

[0202] The first board may send the fault log to the second board at the same time as sending the fault notification to the second board, or may send the fault log to the second board after sending the fault notification to the second board, or the fault log may be directly included in the fault notification.

[0203] Step 411: The second board determines a fault handling strategy.

[0204] Optionally, the second board is further configured to send a fault log to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.

[0205] In some embodiments, after receiving the fault log sent by the first board, the second board may also send the fault log to a host computer. The host computer may then conduct a comprehensive and accurate analysis of the fault of the first board based on the fault log to determine the fault location and / or fault cause of the first board, and output the fault location and / or fault cause to the user, so that the user can perform operations such as inspection, maintenance, and update the first board based on the fault location and / or fault cause of the first board. In actual applications, when the second board sends the fault log to the host computer, it can facilitate the host computer to analyze the fault cause and fault location of the first board, and can facilitate the user to inspect, maintain, and update the first board based on the fault cause and fault location of the first board.

[0206] Step 412: The second board determines whether to restart the functional module.

[0207] Step 413: The second board determines whether to restart the first board.

[0208] Step 414: The second board determines whether to restart the embedded device.

[0209] In this embodiment, after determining the fault handling strategy, the second board first executes step 412 to determine whether the fault handling strategy indicates restarting the functional modules (software modules and / or hardware modules) in the first board. If it is determined that the fault handling strategy indicates restarting the functional modules in the first board, step 415 is executed; if it is determined that the fault handling strategy does not indicate restarting the functional modules in the first board, step 413 is executed.

[0210] In step 413, the second board determines whether the fault handling policy indicates restarting the first board. If so, step 416 is executed. If not, step 414 is executed.

[0211] In step 414 , the second board determines whether the fault handling policy indicates restarting the embedded device. If so, step 417 is executed. If not, step 418 is executed.

[0212] It should be noted that after determining the fault handling strategy, the second board may execute step 412, step 413 and step 414 simultaneously, or may execute step 412, step 413 and step 414 in another order.

[0213] Step 415: The second board notifies the first board to restart the functional module.

[0214] Step 416: The second board controls the first board to restart.

[0215] Step 417: The second board controls the embedded device to restart.

[0216] For understanding of step 415, step 416 and step 417, please refer to the above examples, which will not be described in detail in this embodiment.

[0217] Step 418: The second board records the fault handling information.

[0218] In this embodiment, after executing the fault handling strategy, the second board can record fault handling information, so that personnel can determine the fault handling process of the embedded device based on the fault handling information. Exemplarily, the fault handling information may include the board identifier of the faulty first board, the fault type, location, and time of the first board, and the executed fault handling strategy, but is not limited thereto.

[0219] In some embodiments, during operation of the embedded device, the second board may also execute steps 401 to 409 to detect whether the second board has experienced a fault and obtain a fault log in the event of a fault. Furthermore, upon detecting a fault on the second board itself, the second board may also execute steps 411 to 418 to determine and execute a fault handling strategy corresponding to the second board and record fault handling information.

[0220] FIG7 shows a block diagram of a fault handling device 70 provided in an embodiment of the present application. As shown in FIG7 , the fault handling device includes: a receiving module 71 , a determining module 72 , and an executing module 73 .

[0221] The receiving module 71 is configured to receive a fault notification sent by a first board in the embedded device in the event of a fault;

[0222] a determination module 72 for determining, in response to the fault notification, a fault handling policy corresponding to the first board from a policy configuration;

[0223] The execution module 73 is configured to execute the fault handling strategy and handle the fault of the first board.

[0224] In some embodiments, the determination module 72 is specifically configured to determine the fault handling strategy from the strategy configuration according to the fault log of the first board.

[0225] In some embodiments, the determination module 72 is specifically configured to determine the fault location of the first board according to the fault log; and determine the fault handling strategy from the strategy configuration according to the fault location.

[0226] In some embodiments, the determining module 72 is specifically configured to determine the fault handling policy corresponding to the functional module from the policy configuration when the fault location is a functional module in the first board.

[0227] In some embodiments, the determination module 72 is specifically configured to determine the fault type of the functional module according to the fault log; and determine the fault handling strategy corresponding to the fault type of the functional module from the strategy configuration.

[0228] In some embodiments, the determination module 72 is specifically configured to determine the fault type of the first board according to the fault log; and determine the fault handling strategy from the strategy configuration according to the fault type of the first board.

[0229] In some embodiments, the determination module 72 is specifically configured to determine the fault handling policy from the policy configuration according to the fault type of the first board and the status of the embedded device.

[0230] In some embodiments, the determination module 72 is specifically configured to determine the fault handling strategy from the strategy configuration according to the status of other boards and the fault type of the first board, wherein the other boards include boards associated with the first board in the embedded device.

[0231] In some embodiments, the fault handling apparatus further includes: an acquisition module configured to acquire the fault log from the first board; or acquire the fault log from the fault notification.

[0232] In some embodiments, the execution module 73 is specifically used to control the restart of each board in the embedded device when the fault handling strategy is to restart the embedded device; or, when the fault handling strategy is to restart the first board, send a first restart instruction to the first board to control the restart of the first board; or, when the fault handling strategy is to restart the functional module in the first board, send a second restart instruction to the first board to enable the first board to restart the functional module.

[0233] In some embodiments, the fault handling device further includes: a sending module, configured to send the fault log to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.

[0234] In some embodiments, the fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.

[0235] An embodiment of the present application also provides an embedded device, including a first board and a second board, wherein the first board and the second board are connected via a backplane communication, and the fault handling device described in FIG7 is configured on the second board.

[0236] FIG8 shows a block diagram of another fault handling device 80 provided in an embodiment of the present application. As shown in FIG8 , the fault handling device 80 includes a processor 81 and a memory 82 , and the above components can be connected via one or more buses 84 .

[0237] The fault handling device 80 further includes a computer program 83, which is stored in the memory 82. When the computer program 83 is executed by the processor 81, the fault handling device 80 performs the method shown in Figures 3 to 6. All relevant contents of each step involved in the above method embodiment can be referred to in the functional description of the corresponding physical device and will not be repeated here.

[0238] An embodiment of the present application further provides a readable storage medium, which includes a computer program. When the computer program is run on a computer, the computer executes the method provided in the above method embodiment.

[0239] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method provided in the above method embodiment.

[0240] An embodiment of the present application also provides a chip system, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the method provided by the above method embodiment of the board card installed with the chip system can be implemented.

[0241] The chip system may include an input circuit or interface for sending information or data, and an output circuit or interface for receiving information or data.

[0242] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0243] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0244] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0245] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0246] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0247] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0248] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0249] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer functional unit (which can be a personal computer, a server, or a network functional unit, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0250] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A fault handling method, comprising: receiving a fault notification sent by a first board in the embedded device in a fault condition; In response to the fault notification, determining a fault handling policy corresponding to the first board from a policy configuration; The fault handling strategy is executed to handle the fault of the first board.

2. The method of claim 1, wherein: The determining, from the policy configuration, a fault handling policy corresponding to the first board includes: The fault handling strategy is determined from the strategy configuration according to the fault log of the first board.

3. The method of claim 2, wherein: The determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: Determine the fault location of the first board according to the fault log; The fault handling strategy is determined from the strategy configuration according to the fault location.

4. The method of claim 3, wherein: The determining the fault handling strategy from the strategy configuration according to the fault location includes: In the case where the fault location is a functional module in the first board, the fault handling strategy corresponding to the functional module is determined from the strategy configuration.

5. The method of claim 4, wherein: The determining the fault handling strategy corresponding to the functional module from the strategy configuration includes: Determine the fault type of the functional module according to the fault log; The fault handling strategy corresponding to the fault type of the functional module is determined from the strategy configuration.

6. The method of claim 2, wherein: The determining the fault handling strategy from the strategy configuration according to the fault log of the first board includes: Determine the fault type of the first board according to the fault log; The fault handling strategy is determined from the strategy configuration according to the fault type of the first board.

7. The method of claim 6, wherein: Determining the fault handling strategy from the strategy configuration according to the fault type of the first board includes: The fault handling strategy is determined from the strategy configuration according to the fault type of the first board and the state of the embedded device.

8. The method of claim 7, wherein: The determining the fault handling strategy from the strategy configuration according to the fault type of the first board and the state of the embedded device comprises: The fault handling strategy is determined from the strategy configuration according to the status of other boards and the fault type of the first board, wherein the other boards include boards associated with the first board in the embedded device.

9. The method according to any one of claims 2 to 8, wherein: Before determining the fault handling strategy from the strategy configuration according to the fault log of the first board, the method further includes: The fault log is obtained from the first board; or, the fault log is obtained from the fault notification.

10. The method according to any one of claims 1 to 9, wherein: The executing the fault handling strategy includes: When the fault handling strategy is to restart the embedded device, control each board in the embedded device to restart; or, When the fault handling strategy is to restart the first board, a first restart instruction is sent to the first board to control the first board to restart; or When the fault handling strategy is to restart the functional module in the first board, a second restart instruction is sent to the first board so that the first board restarts the functional module.

11. The method according to any one of claims 2 to 10, wherein: The method further comprises: The fault log is sent to a host computer, so that the host computer performs a fault analysis on the first board according to the fault log.

12. The method according to any one of claims 1 to 11, wherein: The fault handling method is applied to the second board in the embedded device, and the determining of the fault handling strategy corresponding to the first board from the strategy configuration includes: The fault handling strategy is determined from the strategy configuration stored in the second board.

13. A fault handling device, comprising: A receiving module, used for receiving a fault notification sent by a first board in the embedded device in a fault situation; a determination module, configured to determine, in response to the fault notification, a fault handling strategy corresponding to the first board from a strategy configuration; An execution module is used to execute the fault handling strategy to handle the fault of the first board.

14. The fault handling device according to claim 13, wherein: Also includes: An acquisition module, used for acquiring a fault log from the first board; Alternatively, the fault log is obtained from the fault notification.

15. The fault handling device according to any one of claims 14, wherein: Also includes: The sending module is used to send the fault log to the host computer, so that the host computer performs fault analysis on the first board according to the fault log.

16. The fault handling device according to any one of claims 13 to 15, wherein: The fault handling device includes a second board in the embedded device, and the policy configuration is stored in the second board.

17. An embedded device, comprising a first board and a second board, wherein the first board and the second board are connected via a backplane communication, and the fault handling device according to any one of claims 13 to 16 is configured on the second board.

18. A readable storage medium having a computer program stored thereon, wherein when the computer program is executed on a fault handling device, the fault handling device is caused to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Fault processing system and method, server and readable storage medium

    CN114253799A

  • Parallel Multiplex Storage Systems

    US20110239040A1

Cited By

  • Fault processing method and electronic equipment

    CN121050927A