Fault diagnosis method and device, storage medium, electronic equipment and vehicle

By receiving and analyzing global checkpoints from nodes, the master node conducts centralized diagnosis of multiple controllers, solving the problems of complex diagnosis process and high resource consumption in the existing technology, and achieving the effect of simplifying diagnosis and resource conservation.

CN120428683AActive Publication Date: 2025-08-05BEIJING CO WHEELS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410168197.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2025-08-05
Estimated Expiration
2044-02-05

AI Technical Summary

Technical Problem

In the process of troubleshooting multiple controllers, the local checkpoints and target checkpoints of other controllers need to be diagnosed, resulting in complex diagnosis process and large system resource consumption.

Method used

The slave node receives the global checkpoint sent by the master node, judges whether there is a global checkpoint by parsing the configuration data of the local checkpoint, and sends the target global checkpoint to the master node for troubleshooting. The master node conducts centralized diagnosis of multiple controllers.

Benefits of technology

Simplifies the fault diagnosis process, reduces the system resource consumption of the controller to which the slave node belongs, and improves diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428683A_ABST
    Figure CN120428683A_ABST
Patent Text Reader

Abstract

The invention relates to a fault diagnosis method and device, a storage medium, electronic equipment and a vehicle, and relates to the technical field of electronics and electricities, and the method comprises the steps that a slave node receives a global check point sent by a master node, and the global check point is an associated check point of a plurality of controllers; the slave node judges whether the global check point exists in the local check points or not by analyzing the configuration data of the local check points; and if a target global check point in the global check points exists in the local check points, the slave node sends the target global check point to a master node, and the master node is used for performing fault diagnosis on a plurality of controllers related to the target global check point. According to the technical scheme, the local check points in the slave nodes can be further classified, so that the master node diagnoses the global check points, the slave nodes diagnoses the local check points, the fault diagnosis process is simplified, and the system resource consumption of the controller to which the slave nodes belong is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronics and electrical engineering, and particularly to a fault diagnosis method, device, storage medium, electronic device and vehicle. Background Art

[0002] With the development of intelligent connected vehicles and the improvement of hardware capabilities, the concept of software-defined vehicles has gradually penetrated into every component of intelligent vehicles. Automobile manufacturers must attach sufficient importance to the safety requirements in new scenarios and design corresponding safety mechanisms to avoid harm to personal safety caused by intelligent vehicles.

[0003] Currently, the traditional method connects each controller in the form of a daisy chain. Each controller stores the diagnostic rules of local checkpoints and the target checkpoints of other controllers, and determines the diagnostic results of complex processes involving multiple controllers by diagnosing local checkpoints and target checkpoints.

[0004] However, when diagnosing complex processes involving multiple controllers in this way, each controller not only needs to diagnose local checkpoints but also needs to diagnose the target checkpoints of other controllers. The diagnostic process is complex and the system resources consumption is large. Summary of the Invention

[0005] In view of this, this application provides a fault diagnosis method, device, storage medium, electronic device and vehicle, mainly aiming to improve the technical problem that in the prior art when diagnosing complex processes involving multiple controllers, each controller not only needs to diagnose local checkpoints but also needs to diagnose the target checkpoints of other controllers, the diagnostic process is complex and the system resources consumption is large.

[0006] In a first aspect, this application provides a fault diagnosis method, which can be applied to a slave node for execution, including:

[0007] Receiving a global checkpoint sent by a master node, where the global checkpoint is an associated checkpoint of multiple controllers;

[0008] Judging whether the global checkpoint exists in the local checkpoint by parsing the configuration data of the local checkpoint;

[0009] If a target global checkpoint in the global checkpoint exists in the local checkpoint, sending the target global checkpoint to the master node, where the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0010] In a second aspect, this application provides a fault diagnosis method, which can be applied to a master node for execution, including:

[0011] Send a global checkpoint to the slave nodes deployed by the controller, where the global checkpoint is the associated checkpoint of multiple controllers;

[0012] Receive the target global checkpoint in the global checkpoint sent by the slave node, where the target global checkpoint is determined by the slave node from the local checkpoint based on the configuration data of the local checkpoint;

[0013] Perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0014] In a third aspect, the present application provides a fault diagnosis device, which can be applied to a slave node and includes:

[0015] An acquisition module, configured to receive a global checkpoint sent by a master node, where the global checkpoint is the associated checkpoint of multiple controllers;

[0016] An analysis module, configured to determine whether the global checkpoint exists in the local checkpoint by analyzing the configuration data of the local checkpoint;

[0017] A diagnosis module, configured to send the target global checkpoint to the master node if the target global checkpoint exists in the local checkpoint, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0018] In a fourth aspect, the present application provides a fault diagnosis device, which can be applied to a master node and includes:

[0019] An acquisition module, configured to send a global checkpoint to the slave nodes deployed by the controller, where the global checkpoint is the associated checkpoint of multiple controllers;

[0020] An analysis module, configured to receive the target global checkpoint in the global checkpoint sent by the slave node, where the target global checkpoint is determined by the slave node from the local checkpoint based on the configuration data of the local checkpoint;

[0021] A diagnosis module, configured to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0022] In a fifth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in the first aspect or the method described in the second aspect.

[0023] In a sixth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, it implements the method described in the first aspect or the method described in the second aspect.

[0024] In a seventh aspect, the present application provides a vehicle, comprising: the electronic device as described in the sixth aspect;

[0025] The slave node executes the method as described in the first aspect; the master node executes the method as described in the second aspect.

[0026] By means of the above technical solution, a fault diagnosis method, device, storage medium, electronic device and vehicle provided by the present application can receive, by a slave node, a global checkpoint sent by a master node, where the global checkpoint is an associated checkpoint of multiple controllers; by analyzing configuration data of a local checkpoint, it is determined whether the global checkpoint exists in the local checkpoint; if a target global checkpoint in the global checkpoint exists in the local checkpoint, the target global checkpoint is sent to the master node, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint. By applying the technical solution of the present application, the slave node can receive the global checkpoint sent by the master node, determine whether the global checkpoint exists in the local checkpoint by analyzing the configuration data of the local checkpoint, and if the target global checkpoint in the global checkpoint exists in the local checkpoint, send the target global checkpoint to the master node, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint, so as to further classify the local checkpoint in the slave node, enable the master node to diagnose the global checkpoint and the slave node to diagnose the local checkpoint, simplify the fault diagnosis process, and effectively reduce the system resource consumption of the controllers to which the slave node belongs.

[0027] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0030] Figure 1 It shows a schematic flowchart of a fault diagnosis method provided by an embodiment of the present application;

[0031] Figure 2 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0032] Figure 3 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0033] Figure 4 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0034] Figure 5 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0035] Figure 6 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0036] Figure 7 It shows a schematic flowchart of a fault diagnosis method provided by an embodiment of the present application;

[0037] Figure 8 It shows a schematic flowchart of an example provided by an embodiment of the present application;

[0038] Figure 9 It shows a schematic structural diagram of a fault diagnosis device provided by an embodiment of the present application;

[0039] Figure 10 It shows a schematic structural diagram of a fault diagnosis device provided by an embodiment of the present application. Detailed implementation manners

[0040] In order to more clearly understand the above objects, features and advantages of the present application, the solution of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0041] In order to improve the technical problem that in the current prior art when diagnosing a complex process involving multiple controllers, each controller not only needs to diagnose the local checkpoint, but also needs to diagnose the target checkpoints of other controllers, the diagnosis process is complex, and the system resources consumption is relatively large. This embodiment provides a fault diagnosis method, which can be applied to be executed by a slave node, such as Figure 1 As shown, the method includes:

[0042] Step 101, the slave node receives the global checkpoint sent by the master node, and the global checkpoint is the associated checkpoint of multiple controllers.

[0043] Among them, the global checkpoint can be an associated checkpoint of multiple controllers, which is used to detect whether the global tasks involving multiple applications in multiple controllers are executed as required, and can be represented as a tuple, the content of which is the controller number and the local checkpoint number corresponding to the controller number; the local checkpoint can be a checkpoint corresponding to each local task (for example: liveness monitoring task, real-time monitoring task, logical sequence monitoring task) included in each application in the slave node, which is used to detect whether each local task included in each application is executed as required, and can be represented as a number, and this number can be any number.

[0044] Exemplarily, as Figure 2 shown, slave nodes (diagnostic slave nodes) can be deployed on each area / zone controller, and a master node (diagnostic master node) can be deployed on a high-performance computer. A star topology structure can be adopted to establish a communication connection between the master node and each slave node, so that the slave node can receive the global checkpoint sent by the master node. If a certain slave node has a problem, it will only affect the communication link between this slave node and the master node, and will not involve the entire network. Therefore, it is relatively easy to determine the location of the fault, simplify the fault diagnosis process, and facilitate network reconfiguration.

[0045] Step 102: The slave node determines whether there is a global checkpoint in the local checkpoint by parsing the configuration data of the local checkpoint.

[0046] In some embodiments, the local configuration file of the controller to which the slave node belongs can be parsed to obtain the configuration data of the local checkpoint, query whether there is a global checkpoint in the local checkpoint, and filter out the local checkpoints in the slave node that involve global tasks, so as to further classify the local checkpoints in the slave node, which helps to improve the fault diagnosis efficiency.

[0047] Step 103: If there is a target global checkpoint in the global checkpoints in the local checkpoint, the slave node sends the target global checkpoint to the master node.

[0048] Among them, the master node can be used to perform fault diagnosis on multiple controllers involved in the target global checkpoint, and the slave node can be used to monitor the local checkpoint and diagnose the controller to which it belongs.

[0049] In some embodiments, if there is a target global checkpoint in the global checkpoints in the local checkpoint, the target global checkpoint can be sent to the master node, and then the master node performs the diagnosis, so that the slave node only needs to diagnose the local checkpoint and does not need to diagnose the checkpoints involved in other controllers, reducing the system resource consumption of the controller where the slave node is located. Exemplarily, as Figure 3As shown, the working mode of the system in this embodiment is shown. After the vehicle electrical / electronic (E / E) system is started, first, a communication connection can be established between the diagnostic master node and each diagnostic slave node. Then, the application programs in the diagnostic slave nodes can report the local checkpoints corresponding to each task to the distributed diagnostic system, and the distributed diagnostic system forwards the global checkpoints in the local checkpoints to the master node.

[0050] Compared with the current existing technologies, in this embodiment, the slave node can receive the global checkpoint sent by the master node, and by analyzing the configuration data of the local checkpoint, determine whether there is a global checkpoint in the local checkpoint. If there is a target global checkpoint in the global checkpoint in the local checkpoint, the target global checkpoint is sent to the master node. The master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint, so as to further classify the local checkpoints in the slave node, enable the master node to diagnose the global checkpoint and the slave node to diagnose the local checkpoint, simplify the fault diagnosis process, and effectively reduce the system resource consumption of the controller where the slave node is located.

[0051] Further, as a refinement and extension of the above embodiment, in order to fully illustrate the specific implementation process of the method in this embodiment, optionally, the method in this embodiment may further include: obtaining the local checkpoint rule corresponding to the local checkpoint from the configuration data, and the local checkpoint rule includes at least one of the reporting activity, reporting period, reporting times, and reporting order of the local checkpoint; performing fault diagnosis on the local checkpoint according to the local checkpoint rule.

[0052] In some embodiments, local configuration files can be arranged in each controller to store the local checkpoints corresponding to each local task in each application program in the controller and the expected reporting methods of each local checkpoint. The local configuration file may specifically include the parameters shown in Table 1. Among them, the reporting activity can be used to indicate whether the local checkpoint can report information normally. If the reporting activity information of the local checkpoint is reported unsuccessfully, it means that the local checkpoint has lost connection and an activity fault has occurred; the reporting period can be used to indicate the period of reporting information of the local checkpoint, and thus the interval time of reporting information of multiple checkpoints can be determined; the reporting times can be used to indicate the number of times of reporting information of the local checkpoint within the reporting period. For example, if the reporting period of the local checkpoint is 1000 ms and the reporting times is 10 times, it can indicate that the local checkpoint should report 10 times within 1000 ms; the reporting order can be used to indicate the order of reporting information of multiple local checkpoints.

[0053] Table 1

[0054]

[0055]

[0056] Exemplarily, the local checkpoint rules are as follows:

[0057] For the local checkpoint rules of the activity monitoring task: the local checkpoint can be 1, the reporting period can be 1000 ms, the number of reports within the period is 10 times, the upper boundary of the number of reports can be 2 times (reporting 2 more times is also correct), and the lower boundary of the number of reports can be 1 time (reporting 1 less time is also correct);

[0058] For the local checkpoint rules of the real-time monitoring task: the first local checkpoint can be 3, the second local checkpoint can be 4, the maximum time interval is 52 ms, and the minimum time interval can be 48 ms;

[0059] For the local checkpoint rules of the logical order monitoring task: the starting local checkpoint of the task can be 7, the ending local checkpoint of the task can be 10, and the logical process can be reported according to the following sequence process (the local checkpoints in the sequence are between the starting local checkpoint 7 and the ending local checkpoint 10 of the task): the starting local checkpoint is 7 and the ending local checkpoint can be 8, or the starting local checkpoint is 8 and the ending local checkpoint can be 10.

[0060] Optionally, according to the local checkpoint rules, the fault diagnosis of the local checkpoint can specifically include: obtaining the first expected value corresponding to the reported activity from the local checkpoint rules, and comparing the first actual value corresponding to the reported activity reported by the local checkpoint with the first expected value; and / or, obtaining the second expected value corresponding to the reporting period from the local checkpoint rules, and comparing the second actual value corresponding to the reporting period reported by the local checkpoint with the second expected value; and / or, obtaining the third expected value corresponding to the number of reports from the local checkpoint rules, and comparing the third actual value corresponding to the number of reports reported by the local checkpoint with the third expected value; and / or, obtaining the fourth expected value corresponding to the reporting order from the local checkpoint rules, and comparing the fourth actual value corresponding to the reporting order reported by the local checkpoint with the fourth expected value; if the first actual value does not conform to the first expected value, and / or the second actual value does not conform to the second expected value, and / or the third actual value does not conform to the third expected value, and / or the fourth actual value does not conform to the fourth expected value, it is determined that there is a fault in the local checkpoint, and the corresponding recovery process is executed.

[0061] In some embodiments, a slave node may first receive local checkpoints reported by each application, monitor the actual reporting process of the local checkpoints, obtain the actual reported data of the local checkpoints, and then compare the actual reported data with the expected data in the local checkpoint rules. If the actual reported data of the local checkpoint is inconsistent with the expected data, it can be determined that the local checkpoint is abnormal, and then the corresponding local recovery process can be executed. Among them, the first actual value may be the reporting activity in the actual reporting process of the local checkpoint, the second actual value may be the reporting period in the actual reporting process of the local checkpoint, the third actual value may be the number of reports in the actual reporting process of the local checkpoint, the fourth actual value may be the reporting order in the actual reporting process of the local checkpoint, the first expected value may be the expected reporting activity in the local checkpoint rules corresponding to the local checkpoint, the second expected value may be the expected reporting period in the local checkpoint rules corresponding to the local checkpoint, the third expected value may be the expected number of reports in the local checkpoint rules corresponding to the local checkpoint, and the fourth expected value may be the expected reporting order in the local checkpoint rules corresponding to the local checkpoint.

[0062] In a specific application scenario, each application (App) can set multiple activity monitoring tasks. By monitoring the reporting activities of each local checkpoint, it can be determined whether the threads corresponding to each local checkpoint are working properly. Exemplarily, as Figure 4 shown, Checkpoint 1 of the App can be set. The App reports Checkpoint 1 (reportCheckpoint1) to the distributed monitoring program (lihdm) in the distributed diagnostic system. When a fault (such as deadlock) occurs in the application, the reporting stops. At this time, the distributed monitoring program cannot receive Checkpoint 1, and it can be found that there is an activity fault in Checkpoint 1. The distributed monitoring program can execute a recovery operation to receive Checkpoint 1 again. When an activity fault is detected, the distributed monitoring program can send a set Diagnostic Trouble Code (DTC) instruction (set DTC) to the diagnostic program (Diag), so that the diagnostic program can perform corresponding fault prompt operations according to the DTC (for example: turn on the fault light). When the activity fault is eliminated, a clear DTC instruction (clear DTC) can be sent to the diagnostic program, so that the diagnostic program can revoke the corresponding fault prompt operation according to the fault code. After the elimination is successful, a success instruction (success) is returned to the distributed monitoring program.

[0063] Exemplarily, as Figure 5As shown, each App can set multiple real-time monitoring tasks (reportCheckpoint1, reportCheckpoint2). By monitoring the actual reporting period of local checkpoints, it is determined whether each task is completed within the specified time. The expected reporting period (interval time) between checkpoint 1 (Checkpoint1) and checkpoint 2 (Checkpoint2) of the App can be set. When the actual reporting period received by the distributed monitoring program (lihdm) exceeds the expected reporting period, it can be determined that the task times out and there is a real-time fault.

[0064] Exemplarily, as Figure 6 As shown, each App can set multiple logical sequence monitoring tasks (reportCheckpoint1, reportCheckpoint2, reportCheckpoint3) to monitor whether the actual execution process is executed in the expected reporting order. Checkpoint 1 (Checkpoint1), checkpoint 2 (Checkpoint2) and checkpoint 3 (Checkpoint3) of the App can be set, and the expected reporting order of each checkpoint can be set. When there is a resource conflict in the App (for example, multiple Apps control a certain actuator or execute a certain process at the same time), if the actual reporting order received by the distributed monitoring program (lihdm) is inconsistent with the expected reporting order, it can be determined that there is a logical sequence fault in this task.

[0065] Optionally, the method of this embodiment may further include: if the diagnosis result that the target global checkpoint sent by the master node is faulty is received, then execute the recovery process corresponding to the target global checkpoint.

[0066] In some embodiments, when the diagnosis result that the target global checkpoint sent by the master node is faulty is received, the recovery process corresponding to the target global checkpoint in the local configuration file can be executed. After the abnormal recovery, the local checkpoints reported by the corresponding application program can be received again. Among them, the recovery operations may include setting the corresponding DTC, restarting the process, sending a fault broadcast, recording and ignoring, etc. Exemplarily, the recovery process can be set to a fault tolerance time of 1000 ms, the recovery method is restart, and the fault code is 0xD61040.

[0067] To illustrate the processing process of the master node, further, this embodiment also provides a fault diagnosis method applicable to the master node execution, as Figure 7 shown, this method includes:

[0068] Step 201, the master node sends a global checkpoint to the slave nodes deployed by the controller, and the global checkpoint is the associated checkpoint of multiple controllers.

[0069] Exemplarily, asFigure 8 As shown, the system working process of this embodiment is shown, which may specifically include the following steps:

[0070] 1. The distributed diagnostic system on each controller reads the system configuration to determine the controller identity (master node, slave node);

[0071] 2. Establish a heartbeat connection between the master node and each slave node. After all nodes are online, they start working. When a certain node loses the connection, a node loss error can be reported. At this time, the distributed diagnostic system pauses working until the connection is restored;

[0072] 3. The master node sends a global checkpoint to the slave nodes;

[0073] 4. When the slave node detects that the local checkpoint is the global checkpoint, the slave node forwards this checkpoint to the master node;

[0074] 5. The master node verifies the global checkpoint. If an exception occurs, it notifies the relevant slave nodes to execute the global recovery process.

[0075] Among them, the system configuration can be used to determine whether the controller undertakes the master node function. The system configuration parameters may specifically include: the master node function identifier, sub-node information, controller number, Internet Protocol (IP) address, port, etc.

[0076] Step 202: The master node receives the target global checkpoint in the global checkpoint sent by the slave node. The target global checkpoint is determined by the slave node from the local checkpoint based on the configuration data of the local checkpoint.

[0077] In some embodiments, after receiving the global checkpoint, the slave node searches for the local checkpoint corresponding to the global checkpoint based on the local configuration file, determines it as the target global checkpoint, and sends the reporting process of the App reporting the target global checkpoint to the master node, so that the master node can monitor the actual reporting process of the target global checkpoint.

[0078] Step 203: The master node performs fault diagnosis on multiple controllers involved in the target global checkpoint.

[0079] In some embodiments, the global checkpoint can be stored in the global checkpoint configuration file in the master node. After the master node monitors the actual reporting process of the target global checkpoint, it can obtain the global checkpoint rules corresponding to the target global checkpoint in the global checkpoint configuration file. The global checkpoint rules may include the reporting activity, reporting period, reporting times, reporting order, etc. of the global checkpoint. Then, according to the global checkpoint rules, fault diagnosis is performed on the target global checkpoint to reduce the system consumption of the controllers corresponding to the slave nodes.

[0080] Exemplarily, the expected reporting activity corresponding to the target global checkpoint can be obtained from the global checkpoint rule, and the actual reporting activity reported by the target global checkpoint is compared with the expected reporting activity. If the actual reporting activity does not meet the expected reporting activity, it can be determined that the target global checkpoint has a fault, and the corresponding recovery process is executed.

[0081] Exemplarily, the global checkpoint rule is as follows:

[0082] For the global checkpoint rule of the activity monitoring task: Global checkpoint (Controller 1, Local checkpoint 1 in Controller 1), the reporting period can be 1000 ms, the number of reports within the period is 10 times, the upper boundary of the number of reports can be 2 times (reporting 2 more times is also correct), and the lower boundary of the number of reports can be 1 time (reporting 1 less time is also correct);

[0083] For the global checkpoint rule of the real-time monitoring task: Starting point global checkpoint (Controller 1, Local checkpoint 3 in Controller 1), ending point global checkpoint (Controller 2, Local checkpoint 4 in Controller 2), the maximum time interval is 52 ms, and the minimum time interval can be 48 ms;

[0084] For the global checkpoint rule of the logical sequence monitoring task: Task starting global checkpoint (Controller 1, Local checkpoint 7 in Controller 1), task ending global checkpoint (Controller 4, Local checkpoint 10 in Controller 4), and the logical process can be reported according to the following sequence process (the global checkpoint numbers in the sequence are between the task starting global checkpoint (1, 7) and the task ending global checkpoint (4, 10)): Starting global checkpoint (1, 7), ending global checkpoint (2, 6), or starting global checkpoint (2, 7), ending global checkpoint (4, 10).

[0085] Compared with the current existing technologies, in this embodiment, the master node can send the global checkpoint to the slave nodes deployed by the controller, and then receive the target global checkpoint sent by the slave nodes, and perform fault diagnosis on multiple controllers involved in the target global checkpoint, enabling the master node to centrally monitor and diagnose the global checkpoint, reducing the diagnostic work of the slave nodes on the global checkpoint, simplifying the fault diagnosis process, and thus reducing the system resource consumption of the controllers to which the slave nodes belong.

[0086] Optionally, after step 203, the method of this embodiment may further include: If the target global checkpoint has a fault, the diagnosis result that the target global checkpoint has a fault is sent to the slave node corresponding to the target global checkpoint, for instructing the slave node to execute the recovery process corresponding to the target global checkpoint.

[0087] In some embodiments, if the master node determines that there is a fault in the target global checkpoint after quasi-diagnosing the actual reported data of the target global checkpoint, the diagnosis result can be sent to the slave node corresponding to the target global checkpoint, so that the slave node can find the application program corresponding to the target global checkpoint and execute the corresponding recovery process. Exemplarily, the recovery process can be set with a fault tolerance time of 1000 ms, a recovery method of restart, and a fault code of 0xD61040.

[0088] Optionally, before step 201, the method of this embodiment may further include: determining the slave node corresponding to the global checkpoint by parsing the configuration data of the global checkpoint.

[0089] In some embodiments, the master node can obtain the slave nodes involved in the global checkpoint by parsing the global configuration file of the global checkpoint, and then send the global checkpoint to the corresponding slave nodes, so that the slave nodes can identify the target global checkpoint in the local checkpoint and classify the local checkpoint, thereby improving the fault diagnosis efficiency.

[0090] Further, as Figure 1 a specific implementation of the method shown, this embodiment provides a fault diagnosis device, which can be applied to the slave node. As Figure 9 shown, the device includes: an acquisition module 31, an analysis module 32, and a diagnosis module 33.

[0091] The acquisition module 31 is configured to receive the global checkpoint sent by the master node, and the global checkpoint is an associated checkpoint of multiple controllers;

[0092] The analysis module 32 is configured to determine whether there is a global checkpoint in the local checkpoint by parsing the configuration data of the local checkpoint;

[0093] The diagnosis module 33 is configured to, if there is a target global checkpoint in the local checkpoint, send the target global checkpoint to the master node, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0094] In some examples, the diagnosis module 33 is specifically further configured to obtain the local checkpoint rule corresponding to the local checkpoint from the configuration data, and the local checkpoint rule includes at least one of the reporting activity, reporting period, reporting times, and reporting order of the local checkpoint; and perform fault diagnosis on the local checkpoint according to the local checkpoint rule.

[0095] In some examples, the diagnosis module 33 is further specifically configured to obtain a first expected value corresponding to the reporting activity from the local checkpoint rule, and compare the first actual value corresponding to the reporting activity reported by the local checkpoint with the first expected value; and / or obtain a second expected value corresponding to the reporting period from the local checkpoint rule, and compare the second actual value corresponding to the reporting period reported by the local checkpoint with the second expected value; and / or obtain a third expected value corresponding to the reporting times from the local checkpoint rule, and compare the third actual value corresponding to the reporting times reported by the local checkpoint with the third expected value; and / or obtain a fourth expected value corresponding to the reporting order from the local checkpoint rule, and compare the fourth actual value corresponding to the reporting order reported by the local checkpoint with the fourth expected value; if the first actual value does not conform to the first expected value, and / or the second actual value does not conform to the second expected value, and / or the third actual value does not conform to the third expected value, and / or the fourth actual value does not conform to the fourth expected value, it is determined that there is a fault in the local checkpoint, and the corresponding recovery process is executed.

[0096] In some examples, the diagnosis module 33 is further specifically configured to execute the recovery process corresponding to the target global checkpoint if it receives the diagnosis result that the target global checkpoint sent by the master node is faulty.

[0097] It should be noted that for other corresponding descriptions of each functional unit involved in the fault diagnosis device provided in this embodiment, reference can be made to Figure 1 the corresponding description therein, which will not be elaborated herein.

[0098] Furthermore, as Figure 7 a specific implementation of the method shown, this embodiment provides a fault diagnosis device, which can be applied to the master node. As Figure 10 shown, the device includes: an acquisition module 41, an analysis module 42, and a diagnosis module 43.

[0099] The acquisition module 41 is configured to send a global checkpoint to the slave nodes deployed by the controller, and the global checkpoint is an associated checkpoint of multiple controllers;

[0100] The analysis module 42 is configured to receive the target global checkpoint in the global checkpoint sent by the slave node, and the target global checkpoint is determined from the local checkpoint by the slave node based on the configuration data of the local checkpoint;

[0101] The diagnosis module 43 is configured to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

[0102] In some examples, the diagnosis module 43 is specifically further configured to, if a target global checkpoint fails, send the diagnosis result indicating that the target global checkpoint has failed to the slave node corresponding to the target global checkpoint, for instructing the slave node to execute the recovery process corresponding to the target global checkpoint.

[0103] In some examples, the parsing module 42 is specifically further configured to determine the slave node corresponding to the global checkpoint by parsing the configuration data of the global checkpoint.

[0104] It should be noted that for other corresponding descriptions of the various functional units involved in the fault diagnosis device provided in this embodiment, reference can be made to Figure 7 the corresponding description therein, which will not be elaborated here.

[0105] Based on the above method as Figure 1 or Figure 7 shown, correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method as Figure 1 or Figure 7 shown is implemented.

[0106] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.), including several instructions for causing a computer device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0107] Based on the above method as Figure 1 or Figure 7 shown, and Figure 9 or Figure 10 shown in the virtual device embodiment, for achieving the above purpose, this embodiment of the application further provides an electronic device, which can be configured on the vehicle (such as a new energy vehicle) side, and the device includes a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to implement the above method as Figure 1 or Figure 7 shown.

[0108] Optionally, the above physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc., and optionally the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0109] Furthermore, this embodiment also provides a vehicle, which may include devices as shown in Figure 9 and Figure 10 , or include the above-mentioned electronic device.

[0110] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0111] The storage medium may also include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can be implemented by hardware. By applying the solution of this embodiment, the slave node can receive the global checkpoint sent by the master node, and by analyzing the configuration data of the local checkpoint, determine whether there is a global checkpoint in the local checkpoint. If there is a target global checkpoint in the global checkpoint in the local checkpoint, the target global checkpoint is sent to the master node. The master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint, so as to further classify the local checkpoint in the slave node, enable the master node to diagnose the global checkpoint and the slave node to diagnose the local checkpoint, simplify the fault diagnosis process, and effectively reduce the system resource consumption of the controller to which the slave node belongs.

[0113] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0114] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A fault diagnosis method, characterized in that: include: Receive a global checkpoint sent by a master node, where the global checkpoint is a related checkpoint of multiple controllers; Determining whether the global checkpoint exists in the local checkpoint by parsing configuration data of the local checkpoint; If the target global checkpoint in the global checkpoints exists in the local checkpoint, the target global checkpoint is sent to the master node, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

2. The method according to claim 1, characterized in that The method further comprises: Acquire a local checkpoint rule corresponding to the local checkpoint from the configuration data, the local checkpoint rule comprising at least one of a reporting activity, a reporting period, a reporting number, and a reporting order of the local checkpoint; Perform fault diagnosis on the local checkpoint according to the local checkpoint rule.

3. The method according to claim 2, characterized in that The performing fault diagnosis on the local checkpoint according to the local checkpoint rule includes: Obtaining a first expected value corresponding to the reported activity from the local checkpoint rule, and comparing a first actual value corresponding to the reported activity reported by the local checkpoint with the first expected value; and / or, Obtaining a second expected value corresponding to the reporting period from the local checkpoint rule, and comparing a second actual value corresponding to the reporting period reported by the local checkpoint with the second expected value; and / or, Obtaining a third expected value corresponding to the number of reports from the local checkpoint rule, and comparing a third actual value corresponding to the number of reports reported by the local checkpoint with the third expected value; and / or, Obtaining a fourth expected value corresponding to the reporting order from the local checkpoint rule, and comparing a fourth actual value corresponding to the reporting order reported by the local checkpoint with the fourth expected value; If the first actual value does not meet the first expected value, and / or the second actual value does not meet the second expected value, and / or the third actual value does not meet the third expected value, and / or the fourth actual value does not meet the fourth expected value, it is determined that there is a failure in the local checkpoint and the corresponding recovery process is executed.

4. The method according to claim 1, wherein The method further comprises: If a diagnosis result indicating that a fault exists in the target global checkpoint is received from the master node, a recovery process corresponding to the target global checkpoint is executed.

5. A fault diagnosis method, characterized in that: include: Sending a global checkpoint to slave nodes deployed by the controller, where the global checkpoint is a related checkpoint of multiple controllers; receiving a target global checkpoint in the global checkpoints sent by the slave node, where the target global checkpoint is determined by the slave node from the local checkpoint based on configuration data of the local checkpoint; Fault diagnosis is performed on multiple controllers involved in the target global checkpoint.

6. The method according to claim 5, characterized in that After performing fault diagnosis on multiple controllers involved in the target global checkpoint, the method further includes: If the target global checkpoint has a fault, a diagnosis result of the fault in the target global checkpoint is sent to a slave node corresponding to the target global checkpoint to instruct the slave node to execute a recovery process corresponding to the target global checkpoint.

7. The method according to claim 5, characterized in that Before sending the global checkpoint to the slave node deployed by the controller, the method further includes: The slave node corresponding to the global checkpoint is determined by parsing the configuration data of the global checkpoint.

8. A fault diagnosis device, characterized in that: include: an acquisition module configured to receive a global checkpoint sent by a master node, wherein the global checkpoint is a related checkpoint of multiple controllers; a parsing module configured to parse configuration data of a local checkpoint to determine whether the global checkpoint exists in the local checkpoint; The diagnosis module is configured to send the target global checkpoint to the master node if the target global checkpoint exists in the local checkpoint, and the master node is used to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

9. A fault diagnosis device, characterized in that: include: an acquisition module configured to send a global checkpoint to slave nodes deployed by the controller, wherein the global checkpoint is a related checkpoint of multiple controllers; a parsing module configured to receive a target global checkpoint in the global checkpoints sent by the slave node, where the target global checkpoint is determined by the slave node from the local checkpoint based on configuration data of the local checkpoint; The diagnosis module is configured to perform fault diagnosis on multiple controllers involved in the target global checkpoint.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

12. A vehicle, characterized in that: include: The apparatus according to any one of claims 8 to 9, or the electronic device according to claim 11.

Citation Information

Patent Citations

  • Fault processing method and device of database cluster, and terminal

    CN108599996A

  • Node fault detection method and device

    CN110474787A

  • Fault detection method and device for distributed database system and electronic equipment

    CN112783792A

  • Fuel consumption abnormal data detection method and device, electronic equipment and storage medium

    CN113268524A

  • Communication method and system based on security controller, gateway equipment and storage medium

    CN114866263A