Fault Location Method, Device, Electronic Device, Storage Medium and Program Product

By preset heterogeneous graphs, the controller state transition relationship is recorded, combined with the current execution stage and initial operating state, the problem of low controller fault positioning efficiency is solved, and the controller fault positioning is achieved quickly and accurately.

CN120029812BActive Publication Date: 2025-08-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510488833.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-05
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, the controller fault positioning efficiency is low, mainly due to the large amount of log information, which leads to the inefficient parsing and processing efficiency.

Method used

Through the preset heterogeneous diagram, multiple state transition relationships of the controller are recorded, and the current execution stage and initial operation state are determined using the fault location request, and the target fault information is positioned in the preset heterogeneous diagram.

Benefits of technology

Improves the efficiency of fault location, reduces dependence on a large amount of log information, and achieves rapid and accurate positioning of controller failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029812B_ABST
    Figure CN120029812B_ABST
Patent Text Reader

Abstract

The present application discloses a fault location method, device, electronic device, storage medium and program product, which relate to the field of controller technology. A preset heterogeneous graph can be pre-set, and the preset heterogeneous graph includes multiple state transition relationships of the controller. The state transition relationship can indicate the transition situation of the operating state of the controller affected by the preset fault event. After a fault occurs in the controller, the current execution stage of the controller when the fault occurs and the initial operating state of the controller before the fault occur are determined according to the fault location request. The target fault information corresponding to the target fault is determined in the preset heterogeneous graph through the current execution stage and the initial operating state. It is possible to locate the target fault through the preset heterogeneous graph, obtain the target fault information, and improve the efficiency of fault location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of controller technology, and in particular to a fault location method, device, electronic device, storage medium, and program product. Background Art

[0002] The controller can manage a large number of storage devices such as hard disks and solid-state drives. During continuous and high-intensity operation, the controller may experience various failures.

[0003] In related technologies, controller log information can be collected and parsed to obtain corresponding controller fault information. However, the controller processes a large amount of data, and the number of logs increases accordingly, making fault location inefficient. Summary of the Invention

[0004] The present application provides a fault location method, apparatus, electronic device, storage medium, and program product to at least solve the problem of low efficiency of fault location in related technologies.

[0005] This application provides a fault location method, including:

[0006] receiving a fault location request, wherein the fault location request is used to request target fault information corresponding to the controller;

[0007] Determining, according to the fault location request, a current execution phase corresponding to the controller, where the current execution phase is the execution phase of the controller when the fault occurs;

[0008] Acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a fault occurs;

[0009] Acquire a preset heterogeneous graph, where the preset heterogeneous graph includes a plurality of state transition relationships, where the state transition relationships are used to indicate a transition situation in which a preset fault event affects an operating state of the controller;

[0010] The target fault information is determined in the preset heterogeneous graph according to the current execution stage and the initial operation state.

[0011] The present application also provides a fault location device, including a receiving module, a first determining module, a first acquiring module, a second acquiring module, and a second determining module:

[0012] The receiving module is used to receive a fault location request, where the fault location request is used to request fault information corresponding to the controller;

[0013] The first determining module is configured to determine, according to the fault location request, a current execution stage corresponding to the controller, where the current execution stage is the execution stage of the controller when the fault occurs;

[0014] The first acquisition module is used to acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a fault occurs;

[0015] The second acquisition module is used to acquire a preset heterogeneous graph, where the preset heterogeneous graph includes a plurality of state transition relationships, and the state transition relationship is used to indicate a transition situation in which a preset fault event affects the operating state of the controller;

[0016] The second determining module is configured to determine the target fault information in the preset heterogeneous graph according to the current execution stage and the initial operation state.

[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned fault location methods when executing the computer program.

[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned fault location methods are implemented.

[0019] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned fault location methods when executed by a processor.

[0020] Through the fault location method, device, electronic device, storage medium and program product provided by the present application, a preset heterogeneous graph can be pre-set, and the preset heterogeneous graph includes multiple state transition relationships of the controller. The state transition relationship can indicate the transition situation of the operating state of the controller affected by the preset fault event. After a fault occurs in the controller, the current execution stage of the controller at the time of the fault and the initial operating state of the controller before the fault occur can be determined based on the fault location request. The target fault information corresponding to the target fault can be determined in the preset heterogeneous graph based on the current execution stage and the initial operating state. It is possible to locate the target fault through the preset heterogeneous graph and obtain the target fault information, which can improve the efficiency of fault location. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0023] Figure 2 A flowchart of a fault location method provided in an embodiment of the present application;

[0024] Figure 3 A schematic diagram of the structure of a preset heterogeneous graph provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of a process for constructing a preset heterogeneous graph provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of a process for determining the current execution stage provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of a process for determining target fault information provided in an embodiment of the present application;

[0028] Figure 7 A schematic structural diagram of a fault location device provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0032] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application. Figure 1 , including a controller 101 and a detection device 102.

[0033] Controller 101 may be a storage controller that manages and coordinates data transmission between storage devices and a host system. During continuous and high-intensity operation, controller 101 inevitably experiences various faults. If controller 101 faults are not promptly addressed, they can trigger a chain reaction of faults in other related devices or systems. It is worth noting that the present embodiment uses a storage controller as an example to illustrate the fault location method. In reality, the controller may also be other storage devices, such as a server controller.

[0034] After a fault occurs in controller 101, controller 101 may generate a fault location request and send it to detection device 102. Based on the fault location request, detection device 102 may locate the fault in controller 101 and determine target fault information. Detection device 102 may display the target fault information and, based on the target fault information, address the target fault to ensure stable operation of the data storage system.

[0035] In related technologies, controller log information can be collected and parsed to obtain corresponding fault information of the controller. However, the amount of data processed by the controller is large, and the number of log information increases accordingly, making fault location inefficient.

[0036] The fault location method provided in the embodiment of the present application can pre-set a preset heterogeneous graph, which includes multiple state transition relationships of the controller. The state transition relationship can indicate the transition situation of the operating state of the controller affected by the preset fault event. After a fault occurs in the controller, the current execution stage of the controller at the time of the fault and the initial operating state of the controller before the fault occur can be determined based on the fault location request. The target fault information corresponding to the target fault can be determined in the preset heterogeneous graph based on the current execution stage and the initial operating state. It is possible to locate the target fault through the preset heterogeneous graph and obtain the target fault information, which can improve the efficiency of fault location.

[0037] Figure 2 This is a flowchart of a fault location method provided in an embodiment of the present application. Figure 2 , the method may include:

[0038] S201: Receive a fault location request.

[0039] The fault location request can be used to request the target fault information corresponding to the controller.

[0040] The target fault information may include the target fault and target fault type corresponding to the controller, etc., where the target fault can be a description text of the controller fault, and the fault type can be a hard disk failure, memory failure, network interface failure, operating system failure, database failure, application failure, power supply failure, heat dissipation failure, etc.

[0041] S202: Determine the current execution phase corresponding to the controller according to the fault location request.

[0042] The controller may include multiple execution stages during the operation process, and the multiple execution stages have corresponding execution orders.

[0043] In some possible embodiments, the multiple execution stages may be waiting to join the cluster, in the cluster, waiting for abnormal exit from the cluster, and upgrading the cluster, wherein the execution order is: waiting to join the cluster, in the cluster, waiting for abnormal exit from the cluster, and upgrading the cluster.

[0044] Waiting to join the cluster: After the storage controller completes the system startup process and passes the self-test mechanism, it can confirm that its hardware and software components are in normal working condition. The storage controller can enter the ready state, waiting to join the cluster.

[0045] Specifically, when waiting to join the cluster, the storage controller can be divided into two situations: one is that after the first startup, it is ready to interact with the cluster and wait for the cluster's acceptance instructions and related configuration information; the other is that the storage controller has actively exited the cluster (for example, after performing maintenance operations), and after the self-check is normal, it re-enters the state of waiting to join the cluster again. During this period, it continues to listen to the cluster's joining permission signal, waiting to establish a connection and complete the joining process.

[0046] In a cluster: The storage controller successfully establishes a reliable communication connection with other nodes in the cluster and interacts with the cluster through a series of communication handshake protocols (for example, specific application-layer protocols based on TCP / IP). This connection synchronizes the cluster's configuration information, including but not limited to network parameters, security policies, and resource allocation rules. It also initializes various cluster-related services and processes (for example, distributed storage services and data synchronization processes), officially becoming an active member of the cluster and participating in cluster activities such as resource sharing, task scheduling, and data processing, collaborating to maintain the cluster's normal operation.

[0047] Abnormal exit waiting: The cluster management module identifies a storage controller as an abnormal node and removes it from the cluster due to violations of pre-defined cluster rules (e.g., resource usage violations, security policy conflicts, etc.) or because it no longer meets the necessary conditions for cluster operation (e.g., hardware failure, software version incompatibility, etc.). At this point, the storage controller is in an abnormal exit state, its connection to the cluster is severed, and due to its abnormal state, it cannot rejoin the cluster. The node remains in this waiting state while the administrator troubleshoots, repairs, or takes other measures to restore it to a state where it can rejoin the cluster.

[0048] Cluster upgrade in progress: All storage controllers in the cluster simultaneously enter the system upgrade phase, which involves updating the storage controllers' software, firmware, and configuration parameters. During the upgrade process, to prevent system failures caused by conflicting operations or inconsistent data, all external operations on the storage controllers (such as data read and write requests and configuration modifications) are temporarily frozen until all storage controllers complete the upgrade process, pass consistency verification, and are restored to an operational state.

[0049] The current execution stage may be the execution stage of the controller when a fault occurs. For example, assuming that the controller has four execution stages, and assuming that after receiving a fault location request, it can be determined that the controller is in the second execution stage, the second execution stage may be determined as the current execution stage.

[0050] S203: Acquire the initial operating state of the controller.

[0051] During operation, the controller may include multiple operating states, wherein the operating state may be a normal operating state of the controller or an abnormal operating state of the controller.

[0052] Taking the storage controller as an example, common operating states of the storage controller and the corresponding status description of each operating state can be seen in Table 1.

[0053] Table 1

[0054]

[0055] The initial operating state may be the operating state of the controller before a failure occurs. For example, if the operating state of the storage controller before a failure occurs is network interruption, the initial operating state is network interruption.

[0056] S204: Obtain a preset heterogeneous graph.

[0057] A preset heterogeneous graph may be stored in the memory. The preset heterogeneous graph may include a plurality of state transition relationships. The state transition relationships may be used to indicate a transition situation in which a preset fault event affects the operating state of the controller.

[0058] The state transition relationship may include a preset fault event, a pre-operation state, and a post-operation state, wherein the pre-operation state may be the operation state before the preset fault event is injected into the controller, and the post-operation state may be the operation state after the preset fault event is injected into the controller.

[0059] Table 2 is an example of a state transition relationship improved in an embodiment of the present application.

[0060] Table 2

[0061]

[0062] S205 : Determine target fault information in a preset heterogeneous graph according to the current execution stage and the initial operation state.

[0063] The preset heterogeneous graph also includes multiple execution stages and the running status corresponding to each execution stage.

[0064] For ease of understanding, below, combined with Figure 3 , the preset heterogeneous graph provided in the embodiment of the present application is described.

[0065] Figure 3 This is a schematic diagram of a preset heterogeneous graph provided in an embodiment of the present application. Figure 3 The preset heterogeneous graph includes four execution stages, namely, execution stage 301, execution stage 302, execution stage 303, and execution stage 304. Execution stage 301 includes two running states, namely, running state 311 and running state 312. Execution stage 302 includes three running states, namely, running state 321, running state 322, and running state 323. Execution stage 303 includes three running states, namely, running state 331, running state 332, and running state 333. Execution stage 304 includes running state 341.

[0066] The connecting lines between each operating state are used to indicate the preset fault events injected. Figure 3 The preset heterogeneous graph shown includes 7 preset fault events and the state transition relationship corresponding to each preset fault event, wherein the preset fault events are fault events 351-357 and the state transition relationships are state transition relationships 361-367. Figure 3 In the example, each state transition relationship includes a pre-operation state, a preset fault event, and a post-operation state, as shown in Table 3.

[0067] Table 3

[0068]

[0069] Specifically, the target fault event corresponding to the controller can be determined in the preset heterogeneous graph according to the current execution stage and the initial operation state; and the target fault information can be determined from the fault information corresponding to multiple preset fault events according to the target fault event.

[0070] For example, assuming that the current execution stage is the execution stage 302, the initial operation state is the operation state 312, and assuming that the preset heterogeneous graph can be seen in Figure 3 As shown, it can be determined that the target fault event corresponding to the controller is fault event 354. The target fault information corresponding to fault event 354 can be determined.

[0071] The fault location method provided in the embodiment of the present application can, after determining the current execution stage and the initial operating state, determine the fault event and fault type corresponding to the controller through a preset heterogeneous graph, and then determine the fault information corresponding to the controller through the fault event and fault type. There is no need to locate the log corresponding to the fault in a large amount of log information, which can improve the efficiency of fault location.

[0072] Before locating the fault, a test phase is also included. Through the test phase, the state transition relationship corresponding to each preset fault event is determined, and then a preset heterogeneous graph is constructed. Figure 4 , a specific description is provided of the execution process of constructing a preset heterogeneous graph in the embodiment of the present application.

[0073] Figure 4 This is a flow chart of a method for constructing a preset heterogeneous graph provided in an embodiment of the present application. Figure 4 , the method may include:

[0074] S401: Acquire multiple pre-operation states and multiple preset fault events of a controller.

[0075] The following sets can be used to indicate multiple front-end operating states of the controller:

[0076]

[0077] in, Represents a collection of multiple operating states of the controller. Indicates multiple operating states of the controller. Indicates the total number of controller operating states, is an integer greater than or equal to 1.

[0078] Multiple preset fault events of the controller can be indicated by the following sets:

[0079]

[0080] in, Represents a collection of multiple preset fault events. Indicates multiple preset fault events, represents the total number of injection events, is an integer greater than or equal to 1.

[0081] S402: Perform a fault event test on each preceding operation state through a plurality of preset fault events to determine at least one subsequent operation state corresponding to each preceding operation state.

[0082] For any pre-operating state, the controller inputs the i-th preset fault event in the pre-operating state; if the controller transfers from the pre-operating state to the stable operating state, and the stable operating state is an operating state among multiple pre-operating states, then the stable operating state is determined as the post-operating state corresponding to the i-th preset fault event, and the controller inputs the i+1-th preset fault event in the pre-operating state until at least one post-operating state corresponding to the pre-operating state is obtained; if the controller maintains the pre-operating state or is in an unstable operating state, then the controller inputs the i+1-th preset fault event in the pre-operating state until at least one post-operating state corresponding to the pre-operating state is obtained; wherein i is 1, 2, ..., N-1.

[0083] If the controller still maintains the original pre-operation state or an abnormal situation occurs and the stable operation state cannot be determined, it means that there is no state transition relationship between the pre-operation state and the injected preset fault event.

[0084] S403: For any preceding operating state, determine at least one state transition relationship corresponding to the preceding operating state according to at least one subsequent operating state corresponding to the preceding operating state and at least one preset fault event corresponding to the at least one subsequent operating state.

[0085] That is, you can choose a set Any element in As the previous running state (j is 1, 2, ..., M), select the set Any element in As a preset fault event of the input controller, test it on the actual controller; if the controller stabilizes to a new state , indicating the front-end running status After the preset fault event Transfer to state , the state transition relationship can be marked as ,in, called The pre-operation status of called The post-operation status of called Preset fault events.

[0086] S404: Determine multiple state transition relationships according to at least one state transition relationship corresponding to each preceding running state.

[0087] For example, see Figure 3 , including nine pre-conditional operating states: operating state 311, operating state 312, operating state 321, operating state 322, operating state 323, operating state 331, operating state 332, operating state 333, and operating state 341. Taking operating state 311 as an example, when the controller is in operating state 311 and receives fault event 1, it transitions to operating state 321, resulting in state transition relationship 361 corresponding to operating state 311. No state transition relationship exists between other pre-conditional fault events and operating state 311. Therefore, seven state transition relationships can be determined.

[0088] S405: Construct a state transition diagram based on multiple state transition relationships.

[0089] In multiple state transition relationships, there may be a state transition relationship where the preceding running state is the subsequent running state of another state transition relationship. A state transition diagram can be constructed through multiple state transition relationships.

[0090] Specifically, multiple running nodes corresponding to multiple running states can be determined; for any state transition relationship, among multiple running nodes, the preceding running node corresponding to the preceding running state of the state transition relationship and the following running node corresponding to the following running state are determined, and the preceding running node and the following running node are connected by connecting lines to obtain a state transition diagram.

[0091] The connection line may be a directed edge, and the directed edge points from the preceding running state to the succeeding running state. One state transition relationship may correspond to one directed edge.

[0092] S406: Determine a preset heterogeneous graph according to the execution phase and state transition graph corresponding to each running state.

[0093] Specifically, according to the execution stage corresponding to each running state, multiple execution stages of the controller are determined, and the execution order corresponding to the multiple execution stages is determined; according to the execution order of the multiple execution stages, the multiple running states in the state transition diagram are classified and adjusted to obtain a preset heterogeneous graph.

[0094] In the fault location method provided in the embodiments of the present application, before fault location is performed, a controller can be used to test multiple preset fault events to obtain a preset heterogeneous graph. During fault location, the fault event can be determined in the preset heterogeneous graph based on the current execution stage and initial operating state of the controller, thereby determining the target fault information. In the early testing phase, the state transition relationships of the controller are carefully sorted out to construct a preset heterogeneous graph. When a controller fails during actual application, the initial operating state and current execution stage captured in real time, combined with the operating state transition relationships recorded in the preset heterogeneous graph, can be used to efficiently infer and locate the fault, thereby improving the efficiency of fault location.

[0095] Next, combine Figure 5 , an execution process of determining the current execution stage corresponding to the controller provided in an embodiment of the present application is described.

[0096] Figure 5 This is a flowchart of determining the current execution stage provided in an embodiment of the present application. Figure 5 , the method may include:

[0097] S501: According to a fault detection request, obtain the fan speed of the cooling fan of the controller at the time of the fault.

[0098] The cooling fan can perform heat dissipation treatment on the controller. The information of the cooling fan during the cooling process can be stored in the log information. The fan speed of the cooling fan at the time of the fault can be obtained from the log information according to the fault detection request.

[0099] S502: Determine a current load torque corresponding to the cooling fan according to the fan speed of the cooling fan.

[0100] Load torque is used to indicate the resistance torque that the motor needs to overcome when driving the fan to rotate.

[0101] The load torque of the cooling fan is a dynamic parameter, and the magnitude of the load torque is directly related to the fan speed and cooling requirements.

[0102] In some possible embodiments, the motor angular velocity corresponding to the cooling fan can be determined based on the fan speed; the electromagnetic torque corresponding to the cooling fan can be determined based on the motor data and motor angular velocity corresponding to the cooling fan; and the current load torque can be determined based on the electromagnetic torque, motor angular velocity and motor inertia.

[0103] Specifically, the motor data may include motor voltage, motor current, and motor resistance. The electromagnetic torque can be determined by the following formula:

[0104]

[0105]

[0106] in, represents the electromagnetic torque of the fan motor, Indicates the motor voltage at both ends of the cooling fan motor, Indicates the motor current flowing through the cooling fan motor, Indicates the motor resistance of the cooling fan motor, represents the motor angular velocity of the fan motor, represents pi, Indicates the fan speed.

[0107] Specifically, the current load speed can be determined by the following formula:

[0108]

[0109] in, Indicates the current load torque of the fan motor. Indicates that the motor turns to inertia, It represents the derivative of the angular velocity of the fan motor with respect to time t.

[0110] S503: Determine the current execution phase corresponding to the controller according to the current load torque.

[0111] The controller can include multiple execution stages. Taking the four execution stages of the storage controller, namely, waiting to join the cluster, in the cluster, waiting for abnormal exit from the cluster, and in the cluster upgrade, as an example, as the load of multiple execution stages continues to increase, the heat dissipation of the controller needs to continue to increase, and each execution stage corresponds to a different load torque.

[0112] The current execution stage corresponding to the controller can be determined by different load torques.

[0113] Specifically, a correspondence table between load torque intervals and execution stages can be obtained, which includes multiple load torque intervals and the execution stages corresponding to each load torque interval. The target interval corresponding to the current load torque is determined among the multiple load torque intervals, and the execution stage corresponding to the target interval is determined as the current execution stage.

[0114] The corresponding relationship table can be obtained by testing the controller and the cooling fan during the testing phase.

[0115] The fault location method provided in the embodiment of the present application can determine the current load torque corresponding to the cooling fan through the fan speed of the cooling fan, and determine the current execution stage corresponding to the controller based on the current load torque, which can improve the efficiency of determining the current execution stage.

[0116] Next, combine Figure 6, further explains the execution process of determining the target fault information in the preset heterogeneous graph according to the current execution stage and the initial operation state provided in the embodiment of the present application.

[0117] Figure 6 This is a flow chart of determining target fault information provided by an embodiment of the present application. Figure 6 , the method may include:

[0118] S601. Determine multiple first operating states corresponding to the current execution stage according to multiple operating states corresponding to each execution stage.

[0119] The preset heterogeneous graph includes multiple execution stages and multiple operating states corresponding to each execution stage. Among the multiple execution stages, the target execution stage corresponding to the current execution stage can be determined, and the multiple operating states corresponding to the target execution stage can be determined as multiple first operating states.

[0120] For example, see Figure 3 The preset heterogeneous graph shown includes four execution stages. Assuming that the current execution stage is the second execution stage 302, three first operating states can be determined, namely operating state 321, operating state 322 and operating state 323.

[0121] S602: Determine a target operating state according to the multiple first operating states and the initial operating state.

[0122] Specifically, among multiple first operating states, determine at least one second operating state corresponding to the initial operating state; if the number of at least one second operating state is one, determine the one second operating state as the target operating state; if the number of at least one second operating state is multiple, then according to the preset heterogeneous graph, determine the target operating state among the multiple second operating states.

[0123] For any first operating state, the first operating state is compared with a subsequent operating state corresponding to the initial operating state. If the first operating state is consistent with the subsequent operating state corresponding to the initial operating state, the first operating state is determined as the second operating state.

[0124] The second operating state is used to indicate a possible operating state of the controller when a fault occurs.

[0125] For example, assuming that the multiple first operating states include operating state 321, operating state 322 and operating state 323, if the initial operating state is compared with each first operating state and the second operating state is determined to be operating state 321, then operating state 321 is determined as the target operating state.

[0126] Specifically, if there are multiple second operating states, for any second operating state, the first fault event and the first preceding operating state corresponding to the second operating state are determined in multiple state transition relationships; based on the second operating state and the first fault event, the test operating state is determined; if the test operating state is the same as the second operating state, the second operating state is determined as the target operating state.

[0127] A network topology structure of multiple operating states of the controller under preset fault events is constructed in the preset heterogeneous graph. For any second operating state, the second operating state can be used as a post-state. Among multiple state transition relationships, the first fault event and the first preceding operating state corresponding to the second operating state are determined, that is, the state transition relationship corresponding to the second operating state. Then, through the test operating state corresponding to the initial operating state, it is verified that the second operating state is determined to be the location of the state transition relationship, and the second operating state is determined as the target operating state.

[0128] In some possible embodiments, if the first fault event and the first subsequent operating state corresponding to the second operating state are not stored in the multiple state transition relationships, then the second operating state is not the target operating state.

[0129] In some possible embodiments, if the target operating state does not exist in the preset heterogeneous graph, a fault alarm prompt is generated to prompt the user to perform manual fault location processing; the target fault information corresponding to the alarm prompt is received; based on the target fault information, the target fault event, and the preceding operating state and the following operating state corresponding to the target fault event are processed; and the preset heterogeneous graph is updated based on the target fault event, the preceding operating state and the following operating state corresponding to the target fault event.

[0130] In the subsequent maintenance process, fault events and processing experience can be continuously accumulated, and the inference rules and preset heterogeneous graphs can be further optimized to improve the fault location accuracy of the controller.

[0131] S603: Determine a target fault event in a preset heterogeneous graph according to the target operating state.

[0132] The target operating state is used as the post-operating state, and the target fault event corresponding to the target operating state is determined in the preset heterogeneous graph.

[0133] S604: Determine target fault information from fault information corresponding to multiple preset fault events according to the target fault event.

[0134] The fault information corresponding to multiple preset fault events can be found in Table 4. Table 4 shows the correspondence between some controller faults, preset fault events, and fault types. The actual controller faults, preset fault events, and fault types are not limited to the following table.

[0135] Table 4

[0136]

[0137] The fault location method provided in the embodiments of the present application can proactively establish a transition relationship between a preset fault event and an operating state. It can use a preset heterogeneous graph to reproduce various fault scenarios that a controller may encounter during actual operation. This allows for efficient and accurate fault location after a controller fault occurs, using the preset heterogeneous graph. This improves fault location efficiency. Furthermore, by pre-reproducing the controller's fault scenario using the preset heterogeneous graph, the accuracy of fault location can also be improved.

[0138] Figure 7 This is a schematic diagram of the structure of a fault location device provided in an embodiment of the present application. Figure 7 The fault location device 700 may include a receiving module 701, a first determining module 702, a first acquiring module 703, a second acquiring module 704 and a second determining module 705:

[0139] The receiving module 701 is used to receive a fault location request, where the fault location request is used to request the controller for corresponding fault information;

[0140] The first determining module 702 is configured to determine, according to the fault location request, a current execution phase corresponding to the controller, where the current execution phase is the execution phase of the controller when the fault occurs;

[0141] The first acquisition module 703 is used to acquire the initial operating state of the controller, where the initial operating state is the operating state of the controller before a fault occurs;

[0142] The second acquisition module 704 is used to acquire a preset heterogeneous graph, where the preset heterogeneous graph includes multiple state transition relationships, and the state transition relationship is used to indicate the transition of the operating state of the controller affected by the preset fault event;

[0143] The second determining module 705 is configured to determine target fault information in a preset heterogeneous graph according to the current execution stage and the initial operation state.

[0144] In some possible embodiments, the second determining module 705 is specifically configured to:

[0145] According to the current execution stage and initial operation status, the target fault event corresponding to the controller is determined in the preset heterogeneous graph;

[0146] According to the target fault event, target fault information is determined from the fault information corresponding to a plurality of preset fault events.

[0147] In some possible embodiments, the preset heterogeneous graph includes multiple execution stages and multiple running states corresponding to each execution stage; the second determining module 705 is specifically configured to:

[0148] Determining, according to the multiple running states corresponding to each execution stage, multiple first running states corresponding to the current execution stage;

[0149] determining a target operating state according to the plurality of first operating states and the initial operating state;

[0150] According to the target operating status, the target fault event is determined in the preset heterogeneous graph.

[0151] In some possible embodiments, the second determining module 705 is specifically configured to:

[0152] Determining at least one second operating state corresponding to the initial operating state among the plurality of first operating states;

[0153] If the number of the at least one second operating state is one, determining the one second operating state as the target operating state;

[0154] If the number of the at least one second operating state is plural, a target operating state is determined from the plurality of second operating states according to a preset heterogeneous graph.

[0155] In some possible embodiments, the second determining module 705 is specifically configured to:

[0156] For any second operating state, determining a first fault event and a first subsequent operating state corresponding to the second operating state in a plurality of state transition relationships;

[0157] determining a test operation state according to the second operation state and the first fault event;

[0158] If the test operating state is the same as the first post-operating state, the second operating state is determined as the target operating state.

[0159] In some possible embodiments, the first determining module 702 is specifically configured to:

[0160] According to the fault detection request, the fan speed of the controller's cooling fan at the time of the fault is obtained;

[0161] According to the fan speed of the cooling fan, determine the current load torque corresponding to the cooling fan;

[0162] According to the current load torque, the current execution stage corresponding to the controller is determined.

[0163] In some possible embodiments, the first determining module 702 is specifically configured to:

[0164] Determine the motor angular velocity corresponding to the cooling fan based on the fan speed;

[0165] Determine the electromagnetic torque corresponding to the cooling fan based on the motor data and motor angular velocity corresponding to the cooling fan;

[0166] The current load torque is determined based on the electromagnetic torque, motor angular velocity and motor inertia.

[0167] In some possible embodiments, the first determining module 702 is specifically configured to:

[0168] Obtain multiple load intervals corresponding to the cooling fan and a control execution stage corresponding to each load interval;

[0169] A current execution phase is determined among a plurality of control execution phases according to the current load torque and the control execution phase corresponding to each load interval.

[0170] In some possible embodiments, the second obtaining module 704 is specifically configured to:

[0171] Determine multiple state transition relationships, the state transition relationships including a pre-operation state, a preset fault event, and a post-operation state;

[0172] Construct a state transition diagram based on multiple state transition relationships;

[0173] According to the execution phase and state transition diagram corresponding to each running state, the preset heterogeneous graph is determined.

[0174] In some possible embodiments, the second obtaining module 704 is specifically configured to:

[0175] Obtain multiple pre-operation states and multiple preset fault events of the controller;

[0176] Performing a fault event test on each preceding operating state through a plurality of preset fault events to determine at least one subsequent operating state corresponding to each preceding operating state;

[0177] For any preceding operating state, determining at least one state transition relationship corresponding to the preceding operating state based on at least one subsequent operating state corresponding to the preceding operating state and at least one preset fault event corresponding to the at least one subsequent operating state;

[0178] A plurality of state transition relationships are determined according to at least one state transition relationship corresponding to each preceding operating state.

[0179] In some possible embodiments, the second obtaining module 704 is specifically configured to:

[0180] Determine the running node corresponding to each running state in multiple state transition relationships;

[0181] For any state transition relationship, among multiple running nodes, determine the preceding running node corresponding to the preceding running state of the state transition relationship and the following running node corresponding to the following running state, and connect the preceding running node and the following running node through a connecting line to obtain a state transition diagram.

[0182] In some possible embodiments, the second obtaining module 704 is specifically configured to:

[0183] Determine multiple execution stages of the controller according to the execution stage corresponding to each running state, and determine the execution order corresponding to the multiple execution stages;

[0184] According to the execution order of multiple execution stages, multiple operating states in the state transition diagram are classified and adjusted to obtain a preset heterogeneous graph.

[0185] For the description of the features in the embodiment corresponding to the fault location device, reference can be made to the relevant description of the embodiment corresponding to the fault location method, which will not be repeated here.

[0186] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 800 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the electronic device 800 further includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus.

[0187] During the specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802 , so that the at least one processor 801 executes the above-mentioned fault location method embodiment.

[0188] The specific implementation process of the processor 801 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0189] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0190] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0191] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0192] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned fault location method embodiments when running.

[0193] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0194] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned fault location method embodiments are implemented.

[0195] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault location method embodiments are implemented.

[0196] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0197] The above is a detailed introduction to a fault location method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A fault location method, characterized in that: include: receiving a fault location request, wherein the fault location request is used to request target fault information corresponding to the controller; Determining, according to the fault location request, a current execution phase corresponding to the controller, where the current execution phase is the execution phase of the controller when the fault occurs; Acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a fault occurs; Obtaining a preset heterogeneous graph, the preset heterogeneous graph including a plurality of state transition relationships, the state transition relationships being used to indicate transition situations in which a preset fault event affects an operating state of the controller, the state transition relationships including a pre-operating state, a preset fault event, and a post-operating state; the preset heterogeneous graph including a plurality of execution stages and a plurality of operating states corresponding to each execution stage; Determining, in the preset heterogeneous graph, a target fault event corresponding to the controller according to the current execution stage and the initial operation state; According to the target fault event, determining the target fault information from the fault information corresponding to a plurality of preset fault events; Determining a target fault event corresponding to the controller in the preset heterogeneous graph according to the current execution stage and the initial operation state includes: Determining, according to the multiple running states corresponding to each execution stage, multiple first running states corresponding to the current execution stage; determining a target operating state according to the plurality of first operating states and the initial operating state, wherein the target operating state is a state of the controller after a fault occurs; According to the target operating state, the target fault event is determined in the preset heterogeneous graph.

2. The method according to claim 1, characterized in that Determining the target operating state according to the plurality of first operating states and the initial operating state includes: Determine, among the multiple first operating states, at least one second operating state corresponding to the initial operating state, where the at least one second operating state corresponding to the initial operating state is at least one second operating state having a state transition relationship with the initial operating state; If the number of the at least one second operating state is one, determining the one second operating state as the target operating state; If there are multiple second operating states, the target operating state is determined from the multiple second operating states according to the preset heterogeneous graph.

3. The method according to claim 2, characterized in that Determining the target operating state from among the plurality of second operating states according to the preset heterogeneous graph includes: For any second operating state, determining a first fault event and a first preceding operating state corresponding to the second operating state in the plurality of state transition relationships; Determine a test operating state according to the second operating state and the first fault event, where the test operating state is the operating state of the controller after the first fault event is injected into the initial operating state; If the test operating state is the same as the second operating state, the second operating state is determined as the target operating state.

4. The method according to claim 1, wherein Determining, according to the fault location request, a current execution phase corresponding to the controller, including: acquiring, according to the fault detection request, a fan speed of the cooling fan of the controller at the moment of the fault; Determining a current load torque corresponding to the cooling fan according to a fan speed of the cooling fan; A current execution phase corresponding to the controller is determined according to the current load torque.

5. The method according to claim 4, characterized in that Determining a current load torque corresponding to the cooling fan according to the fan speed of the cooling fan includes: Determining the motor angular velocity corresponding to the cooling fan according to the fan speed; Determining the electromagnetic torque corresponding to the cooling fan according to the motor data corresponding to the cooling fan and the motor angular velocity; The current load torque is determined according to the electromagnetic torque, the motor angular velocity and the motor inertia.

6. The method according to claim 4, characterized in that Determining a current execution stage corresponding to the controller according to the current load torque includes: Obtaining multiple load intervals corresponding to the cooling fan and a control execution stage corresponding to each load interval; The current execution phase is determined among a plurality of control execution phases according to the current load torque and the control execution phase corresponding to each load interval.

7. The method according to claim 1, characterized in that Get the preset heterogeneous graph, including: Determining the plurality of state transition relationships, the state transition relationships including a pre-operation state, a preset fault event, and a post-operation state; Constructing a state transition diagram according to the multiple state transition relationships; The preset heterogeneous graph is determined according to the execution stage corresponding to each running state and the state transition graph.

8. The method according to claim 7, characterized in that Determine multiple state transition relationships, including: Acquiring multiple pre-operation states and multiple preset fault events of the controller; Performing a fault event test on each preceding operating state through the plurality of preset fault events to determine at least one subsequent operating state corresponding to each preceding operating state; For any preceding operating state, determining at least one state transition relationship corresponding to the preceding operating state based on at least one subsequent operating state corresponding to the preceding operating state and at least one preset fault event corresponding to the at least one subsequent operating state; The multiple state transition relationships are determined according to at least one state transition relationship corresponding to each preceding running state.

9. The method according to claim 7, characterized in that Constructing a state transition diagram according to the multiple state transition relationships includes: Determine the running node corresponding to each running state in the multiple state transition relationships; For any state transition relationship, among multiple running nodes, determine the preceding running node corresponding to the preceding running state of the state transition relationship and the following running node corresponding to the following running state, and connect the preceding running node and the following running node through a connecting line to obtain the state transition diagram.

10. The method according to claim 7, characterized in that Determining the preset heterogeneous graph according to the execution phase corresponding to each running state and the state transition graph includes: Determining a plurality of execution stages of the controller according to the execution stage corresponding to each operating state, and determining an execution order corresponding to the plurality of execution stages; According to the execution order of the multiple execution stages, the multiple operating states in the state transition diagram are classified and adjusted to obtain the preset heterogeneous graph.

11. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the fault location method according to any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault location method according to any one of claims 1 to 10.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault location method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Fault positioning method based on finite-state machine and graph neural network

    CN111966076A