Fault positioning method and device, electronic equipment, storage medium and program product
By presetting heterogeneous patterns in the controller, recording the state transition relationship, and positioning faults in the heterogeneous patterns using the current execution stage and initial operating state, the problem of low fault positioning efficiency in the existing technology is solved, and fast and accurate fault positioning is achieved.
Patent Information
- Application Number
- CN202510488833.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the prior art, the efficiency of fault positioning is low, which makes it difficult to quickly and accurately locate the fault source when the controller fails.
By preset heterogeneous diagram, multiple state transition relationships of the controller are recorded, and target fault information is determined in the preset heterogeneous diagram using the current execution stage and initial running state in the fault location request.
It realizes efficient positioning of target faults through preset heterogeneous graphs, improves the efficiency of fault location and reduces dependence on a large amount of log information.
Smart Images

Figure CN120029812A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of controller technology, and in particular to a fault location method, device, electronic device, storage medium and program product. Background Art
[0002] The controller can manage a large number of storage devices such as hard disks and solid-state drives. During continuous and high-intensity work, the controller may experience various failures.
[0003] In the related art, the log information of the controller can be collected and parsed to obtain the fault information corresponding to the controller. However, the amount of data processed by the controller is large, and the number of log information increases accordingly, which makes the efficiency of fault location low. Summary of the invention
[0004] The present application provides a fault location method, device, electronic device, storage medium and program product to at least solve the problem of low efficiency of fault location in related technologies.
[0005] The present application provides a fault location method, including:
[0006] receiving a fault location request, wherein the fault location request is used to request target fault information corresponding to the controller;
[0007] Determine, according to the fault location request, a current execution stage corresponding to the controller, wherein the current execution stage is an execution stage of the controller when the fault occurs;
[0008] Acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a failure occurs;
[0009] Acquire a preset heterogeneous graph, wherein the preset heterogeneous graph includes a plurality of state transition relationships, and the state transition relationship is used to indicate a transition situation in which a preset fault event affects an operating state of the controller;
[0010] According to the current execution stage and the initial operation state, the target fault information is determined in the preset heterogeneous graph.
[0011] The present application also provides a fault location device, including a receiving module, a first determining module, a first acquiring module, a second acquiring module and a second determining module:
[0012] The receiving module is used to receive a fault location request, where the fault location request is used to request fault information corresponding to the controller;
[0013] The first determination module is used to determine, according to the fault location request, a current execution stage corresponding to the controller, where the current execution stage is the execution stage of the controller when a fault occurs;
[0014] The first acquisition module is used to acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a fault occurs;
[0015] The second acquisition module is used to acquire a preset heterogeneous graph, wherein the preset heterogeneous graph includes a plurality of state transition relationships, and the state transition relationship is used to indicate a transition situation in which a preset fault event affects an operating state of the controller;
[0016] The second determination module is used to determine the target fault information in the preset heterogeneous graph according to the current execution stage and the initial operation state.
[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned fault location methods when executing the computer program.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned fault location methods are implemented.
[0019] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned fault location methods when executed by a processor.
[0020] Through the fault location method, device, electronic device, storage medium and program product provided by the present application, a preset heterogeneous graph can be pre-set, and the preset heterogeneous graph includes multiple state transition relationships of the controller, and the state transition relationship can indicate the transition of the operating state of the controller affected by the preset fault event. After a controller fails, the current execution stage of the controller when the fault occurs and the initial operating state of the controller before the fault occur can be determined according to the fault location request. The target fault information corresponding to the target fault can be determined in the preset heterogeneous graph through the current execution stage and the initial operating state. It is possible to locate the target fault through the preset heterogeneous graph and obtain the target fault information, which can improve the efficiency of fault location. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a flow chart of a fault location method provided in an embodiment of the present application;
[0024] Figure 3 A schematic diagram of the structure of a preset heterogeneous graph provided in an embodiment of the present application;
[0025] Figure 4 A schematic diagram of a process for constructing a preset heterogeneous graph provided in an embodiment of the present application;
[0026] Figure 5 A schematic diagram of a process for determining the current execution stage provided in an embodiment of the present application;
[0027] Figure 6 A schematic diagram of a process for determining target fault information provided in an embodiment of the present application;
[0028] Figure 7 A schematic diagram of the structure of a fault location device provided in an embodiment of the present application;
[0029] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0031] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0032] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application. Figure 1 , including a controller 101 and a detection device 102.
[0033] The controller 101 may be a storage controller, and the controller 101 may manage and coordinate data transmission between the storage device and the host system. In the continuous and high-intensity working process of the controller 101, various faults will inevitably occur. If the fault of the controller 101 is not handled in time, a chain reaction of faults of other related devices or systems will be triggered. It is worth noting that the embodiment of the present application only takes the storage controller as an example to illustrate the fault location method. In fact, the controller may also be other storage devices, for example, the controller may also be a server controller, etc.
[0034] After the controller 101 fails, the controller 101 may generate a fault location request and send the fault location request to the detection device 102. The detection device 102 may locate the fault of the controller 101 according to the fault location request and determine the target fault information. The detection device 102 may display the target fault information and process the target fault according to the target fault information to ensure the stable operation of the data storage system.
[0035] In the related art, the log information of the controller can be collected and parsed to obtain the fault information corresponding to the controller. However, the amount of data processed by the controller is large, and the number of log information increases accordingly, making the efficiency of fault location low.
[0036] The fault location method provided in the embodiment of the present application can pre-set a preset heterogeneous graph, which includes multiple state transition relationships of the controller. The state transition relationship can indicate the transition of the operating state of the controller affected by the preset fault event. After a controller fails, the current execution stage of the controller when the fault occurs and the initial operating state of the controller before the fault occur can be determined according to the fault location request. The target fault information corresponding to the target fault can be determined in the preset heterogeneous graph through the current execution stage and the initial operating state. It is possible to locate the target fault through the preset heterogeneous graph and obtain the target fault information, which can improve the efficiency of fault location.
[0037] Figure 2 A flowchart of a fault location method provided in an embodiment of the present application. Figure 2 , the method may include:
[0038] S201: Receive a fault location request.
[0039] The fault location request can be used to request the target fault information corresponding to the controller.
[0040] The target fault information may include the target fault and target fault type corresponding to the controller, etc., wherein the target fault may be a description text of the controller fault, and the fault type may be a hard disk failure, memory failure, network interface failure, operating system failure, database failure, application failure, power supply failure, heat dissipation failure, etc.
[0041] S202: Determine a current execution phase corresponding to the controller according to the fault location request.
[0042] The controller may include multiple execution stages during the operation process, and the multiple execution stages have their corresponding execution orders.
[0043] In some possible embodiments, the multiple execution stages may be respectively waiting to join the cluster, in the cluster, waiting for abnormal exit from the cluster, and upgrading the cluster, wherein the execution order is: waiting to join the cluster, in the cluster, waiting for abnormal exit from the cluster, and upgrading the cluster.
[0044] Waiting to join the cluster: After the storage controller completes the system startup process and passes the self-test mechanism, it can confirm that its hardware and software components are in normal working condition. The storage controller can enter the standby state of waiting to join the cluster.
[0045] Specifically, in the state of waiting to join the cluster, the storage controller can be divided into two situations: one is that after the first startup, it is ready to interact with the cluster and wait for the cluster's acceptance instructions and related configuration information; the other is that the storage controller has actively exited the cluster (for example, after maintenance operations), and after the self-check is normal, it re-enters the state of waiting to join the cluster again. During this period, it continues to listen to the cluster's joining permission signal, waiting to establish a connection and complete the joining process.
[0046] In a cluster: The storage controller successfully establishes a reliable communication connection with other nodes in the cluster, and implements information exchange with the cluster through a series of communication handshake protocols (for example, specific application layer protocols based on TCP / IP). On this basis, the storage controller synchronizes the cluster's configuration information, including but not limited to network parameters, security policies, resource allocation rules, etc., and initializes various services and processes related to the cluster (for example, distributed storage services, data synchronization processes, etc.), officially becoming an effective member of the cluster, participating in the cluster's resource sharing, task scheduling, data processing and other activities, and collaboratively maintaining the normal operation of the cluster.
[0047] Abnormal exit from cluster waiting: The storage controller is identified as an abnormal node by the cluster management module and removed because it violates pre-set cluster rules (e.g., resource usage violations, security policy conflicts, etc.) or no longer meets the necessary conditions for cluster operation (e.g., hardware failure, incompatible software versions, etc.). At this point, the storage controller is in an abnormal exit state, its connection with the cluster has been cut off, and it cannot join the cluster again due to its abnormal state. The node will continue to be in this waiting state until the administrator troubleshoots, repairs, or takes other processing measures to restore it to a state where it can rejoin the cluster.
[0048] Cluster upgrade in progress: All storage controllers in the cluster simultaneously enter the system upgrade phase, which involves updating the storage controller's software, firmware, or configuration parameters. During the upgrade process, to avoid system failures caused by operation conflicts or data inconsistencies, all external operations on the storage controller (for example, data read and write requests, configuration modifications) are temporarily frozen until all storage controllers complete the upgrade process and pass consistency verification, and are restored to an operational state.
[0049] The current execution stage may be the execution stage of the controller when a fault occurs. For example, assuming that the controller has four execution stages in total, and assuming that after receiving a fault location request, it can be determined that the controller is in the second execution stage, the second execution stage may be determined as the current execution stage.
[0050] S203: Obtain the initial operating state of the controller.
[0051] During operation, the controller may include multiple operating states, wherein the operating state may be a normal operating state of the controller or an abnormal operating state of the controller.
[0052] Taking a storage controller as an example, common operating states of a storage controller and a state description corresponding to each operating state can be seen in Table 1.
[0053] Table 1
[0054]
[0055] The initial operating state may be an operating state of the controller before a failure occurs. For example, if the operating state of the storage controller before a failure occurs is network interruption, the initial operating state is network interruption.
[0056] S204: Obtain a preset heterogeneous graph.
[0057] A preset heterogeneous graph may be stored in a memory, and the preset heterogeneous graph may include a plurality of state transition relationships, and the state transition relationship may be used to indicate a transition situation in which a preset fault event affects an operating state of the controller.
[0058] The state transition relationship may include a preset fault event, a pre-operation state, and a post-operation state, wherein the pre-operation state may be the operation state before the preset fault event is injected into the controller, and the post-operation state may be the operation state after the preset fault event is injected into the controller.
[0059] Table 2 is an example of a state transition relationship improved in an embodiment of the present application.
[0060] Table 2
[0061]
[0062] S205: Determine target fault information in a preset heterogeneous graph according to the current execution stage and the initial operation state.
[0063] The preset heterogeneous graph also includes multiple execution stages and the running status corresponding to each execution stage.
[0064] For ease of understanding, below, combined Figure 3 , the preset heterogeneous graph provided in the embodiment of the present application is explained.
[0065] Figure 3 A schematic diagram of a preset heterogeneous graph provided in an embodiment of the present application. Figure 3 The preset heterogeneous graph includes four execution stages, namely, execution stage 301, execution stage 302, execution stage 303 and execution stage 304. Execution stage 301 includes two running states, namely, running state 311 and running state 312. Execution stage 302 includes three running states, namely, running state 321, running state 322 and running state 323. Execution stage 303 includes three running states, namely, running state 331, running state 332 and running state 333. Execution stage 304 includes running state 341.
[0066] The connecting lines between each operating state are used to indicate the preset fault events injected. Figure 3 The preset heterogeneous graph shown includes 7 preset fault events and the state transition relationship corresponding to each preset fault event, wherein the preset fault events are fault events 351-357 and the state transition relationships are state transition relationships 361-367. Figure 3 In the embodiment, each state transition relationship includes a pre-operation state, a preset fault event, and a post-operation state, as shown in Table 3.
[0067] Table 3
[0068]
[0069] Specifically, the target fault event corresponding to the controller can be determined in the preset heterogeneous graph according to the current execution stage and the initial operation state; and the target fault information can be determined from the fault information corresponding to multiple preset fault events according to the target fault event.
[0070] For example, assuming that the current execution stage is execution stage 301, the initial operation state is operation state 12, and assuming that the preset heterogeneous graph can be seen in Figure 3 As shown, it can be determined that the target fault event corresponding to the controller is fault event 4. The target fault information corresponding to fault event 4 can be determined.
[0071] The fault location method provided in the embodiment of the present application can, after determining the current execution stage and the initial operating state, determine the fault event and fault type corresponding to the controller through a preset heterogeneous graph, and then determine the fault information corresponding to the controller through the fault event and the fault type. There is no need to locate the log corresponding to the fault in more log information, which can improve the efficiency of fault location.
[0072] Before locating the fault, a test phase is also included. Through the test phase, the state transition relationship corresponding to each preset fault event is determined, and then a preset heterogeneous graph is constructed. Figure 4 , a specific description is given of the execution process of constructing a preset heterogeneous graph provided in the embodiment of the present application.
[0073] Figure 4 A schematic diagram of a process for constructing a preset heterogeneous graph provided in an embodiment of the present application. Figure 4 , the method may include:
[0074] S401. Acquire multiple pre-operation states and multiple preset fault events of a controller.
[0075] The following sets can be used to indicate multiple front-end operating states of the controller:
[0076]
[0077] in, Represents a collection of multiple operating states of the controller. Indicates multiple operating states of the controller. Indicates the total number of controller operating states. is an integer greater than or equal to 1.
[0078] Multiple preset fault events of the controller can be indicated by the following set:
[0079]
[0080] in, Represents a collection of multiple preset fault events. Indicates multiple preset fault events, represents the total number of injection events, is an integer greater than or equal to 1.
[0081] S402: Perform a fault event test on each preceding operation state through a plurality of preset fault events, and determine at least one subsequent operation state corresponding to each preceding operation state.
[0082] For any pre-operating state, the controller inputs the i-th preset fault event in the pre-operating state; if the controller is transferred from the pre-operating state to the stable operating state, and the stable operating state is an operating state among multiple pre-operating states, the stable operating state is determined as the post-operating state corresponding to the i-th preset fault event, and the controller inputs the i+1-th preset fault event in the pre-operating state until at least one post-operating state corresponding to the pre-operating state is obtained; if the controller maintains the pre-operating state or is in an unstable operating state, the controller inputs the i+1-th preset fault event in the pre-operating state until at least one post-operating state corresponding to the pre-operating state is obtained; wherein i is 1, 2, ..., N-1.
[0083] If the controller still maintains the original pre-operation state or an abnormal situation occurs and a stable operation state cannot be determined, it means that there is no state transfer relationship between the pre-operation state and the injected preset fault event.
[0084] S403: for any preceding operation state, determine at least one state transition relationship corresponding to the preceding operation state according to at least one succeeding operation state corresponding to the preceding operation state and at least one preset fault event corresponding to at least one succeeding operation state.
[0085] That is, you can choose a collection Any element in As the previous running state (j is 1, 2, ..., M), select the set Any element in As a preset fault event of the input controller, test it on the actual controller; if the controller stabilizes to a new state , indicating the front-end running status After the preset fault event Transfer to state , the state transition relationship can be marked as ,in, called The pre-operation status of called The post-operation status of called Preset fault events.
[0086] S404: Determine a plurality of state transition relationships according to at least one state transition relationship corresponding to each preceding running state.
[0087] For example, see Figure 3 , including 9 pre-operation states, namely operation state 311, operation state 312, operation state 321, operation state 322, operation state 323, operation state 331, operation state 332, operation state 333 and operation state 341. Taking operation state 311 as an example, when the controller is in operation state 311, after inputting fault event 1, it transfers to operation state 321 and obtains state transition relationship 361 corresponding to operation state 311. There is no state transition relationship between other preset fault events and operation state 311. In this way, 7 state transition relationships can be determined.
[0088] S405: Construct a state transition diagram according to multiple state transition relationships.
[0089] Among multiple state transition relationships, there may be a state transition relationship whose preceding running state is the subsequent running state of another state transition relationship. Through multiple state transition relationships, a state transition diagram can be constructed.
[0090] Specifically, multiple running nodes corresponding to multiple running states can be determined; for any state transfer relationship, among the multiple running nodes, the preceding running node corresponding to the preceding running state of the state transfer relationship and the following running node corresponding to the following running state are determined, and the preceding running node and the following running node are connected through connecting lines to obtain a state transfer diagram.
[0091] Among them, the connecting line can be a directed edge, and the directed edge points from the previous running state to the subsequent running state. One state transfer relationship can correspond to one directed edge.
[0092] S406: Determine a preset heterogeneous graph according to the execution phase and state transition diagram corresponding to each running state.
[0093] Specifically, according to the execution stage corresponding to each operating state, multiple execution stages of the controller are determined, and the execution order corresponding to the multiple execution stages is determined; according to the execution order of the multiple execution stages, the multiple operating states in the state transition diagram are classified and adjusted to obtain a preset heterogeneous graph.
[0094] In the fault location method provided in the embodiment of the present application, before fault location is performed, multiple preset fault events can be tested by the controller to obtain a preset heterogeneous graph. When fault location is performed, the fault event can be determined in the preset heterogeneous graph according to the current execution stage and initial operation state of the controller, and then the target fault information is determined. In the early testing stage, the state transition relationship of the controller is carefully sorted out to construct a preset heterogeneous graph; when a controller fails during actual application, based on the initial operation state and current execution stage captured in real time, combined with the operation state transition relationship recorded in the preset heterogeneous graph, the fault can be efficiently inferred and located, thereby improving the efficiency of fault location.
[0095] Next, combine Figure 5 , an execution process of determining the current execution stage corresponding to the controller provided in an embodiment of the present application is described.
[0096] Figure 5 A flowchart of determining the current execution stage provided in an embodiment of the present application. Figure 5 , the method may include:
[0097] S501. According to a fault detection request, obtain a fan speed of a cooling fan of a controller at a fault moment.
[0098] The cooling fan can perform heat treatment on the controller. The information of the cooling fan during the cooling process can be stored in the log information. The fan speed of the cooling fan at the time of the fault can be obtained from the log information according to the fault detection request.
[0099] S502: Determine a current load torque corresponding to the cooling fan according to the fan speed of the cooling fan.
[0100] Load torque is used to indicate the resistance torque that the motor needs to overcome when driving the fan to rotate.
[0101] The load torque of the cooling fan is a dynamic parameter, and the size of the load torque is directly related to the fan speed and cooling requirements.
[0102] In some possible embodiments, the motor angular velocity corresponding to the cooling fan can be determined based on the fan speed; the electromagnetic torque corresponding to the cooling fan can be determined based on the motor data and the motor angular velocity corresponding to the cooling fan; and the current load torque can be determined based on the electromagnetic torque, the motor angular velocity and the motor rotation inertia.
[0103] Specifically, the motor data may include motor voltage, motor current and motor resistance, and the electromagnetic torque may be determined by the following formula:
[0104]
[0105]
[0106] in, represents the electromagnetic torque of the fan motor, Indicates the motor voltage at both ends of the cooling fan motor, Indicates the motor current flowing through the cooling fan motor, represents the motor resistance of the cooling fan motor, represents the motor angular velocity of the fan motor, represents pi, Indicates the fan speed.
[0107] Specifically, the current load speed can be determined by the following formula:
[0108]
[0109] in, Indicates the current load torque of the fan motor. Indicates that the motor turns to inertia, It means to find the derivative of the angular velocity of the fan motor with respect to time t.
[0110] S503: Determine the current execution phase corresponding to the controller according to the current load torque.
[0111] The controller may include multiple execution stages. For example, the storage controller may be in the four execution stages of waiting to join the cluster, in the cluster, waiting to exit the cluster abnormally, and in cluster upgrade. As the load of multiple execution stages increases, the heat dissipation of the controller needs to increase continuously, and each execution stage corresponds to a different load torque.
[0112] The current execution stage corresponding to the controller can be determined by different load torques.
[0113] Specifically, a correspondence table between load torque intervals and execution stages can be obtained, the correspondence table including multiple load torque intervals and the execution stages corresponding to each load torque interval, the target interval corresponding to the current load torque is determined among the multiple load torque intervals, and the execution stage corresponding to the target interval is determined as the current execution stage.
[0114] The corresponding relationship table can be obtained by testing the controller and the cooling fan in the testing phase.
[0115] The fault location method provided in the embodiment of the present application can determine the current load torque corresponding to the cooling fan through the fan speed of the cooling fan, and determine the current execution stage corresponding to the controller based on the current load torque, thereby improving the efficiency of determining the current execution stage.
[0116] Next, combine Figure 6, further illustrating the execution process of determining the target fault information in a preset heterogeneous graph according to the current execution stage and the initial operation state provided in the embodiment of the present application.
[0117] Figure 6 A schematic diagram of a process for determining target fault information provided by an embodiment of the present application. Figure 6 , the method may include:
[0118] S601. Determine multiple first operating states corresponding to the current execution stage according to multiple operating states corresponding to each execution stage.
[0119] The preset heterogeneous graph includes multiple execution stages and multiple operating states corresponding to each execution stage. Among the multiple execution stages, the target execution stage corresponding to the current execution stage can be determined, and the multiple operating states corresponding to the target execution stage can be determined as multiple first operating states.
[0120] For example, see Figure 3 The preset heterogeneous graph shown includes four execution stages. Assuming that the current execution stage is the second execution stage 302, three first operating states can be determined, namely operating state 321, operating state 322 and operating state 323.
[0121] S602: Determine a target operating state according to a plurality of first operating states and an initial operating state.
[0122] Specifically, among multiple first operating states, determine at least one second operating state corresponding to the initial operating state; if the number of at least one second operating state is one, determine the one second operating state as the target operating state; if the number of at least one second operating state is multiple, then according to a preset heterogeneous graph, determine the target operating state among the multiple second operating states.
[0123] For any first operating state, the first operating state is compared with the initial operating state. If the first operating state is consistent with the initial operating state, the first operating state is determined as the second operating state.
[0124] The second operating state is used to indicate a possible operating state of the controller when a fault occurs.
[0125] For example, assuming that the multiple first operating states include operating state 321, operating state 322 and operating state 323, if the initial operating state is compared with each first operating state and the second operating state is determined to be operating state 321, then operating state 321 is determined as the target operating state.
[0126] Specifically, if there are multiple second operating states, for any second operating state, among multiple state transition relationships, determine the first fault event and the first post-operating state corresponding to the second operating state; determine the test operating state based on the second operating state and the first fault event; if the test operating state is the same as the first post-operating state, determine the second operating state as the target operating state.
[0127] A network topology structure of multiple operating states of the controller under preset fault events is constructed in the preset heterogeneous graph. For any second operating state, the second operating state can be used as a preceding state, and among multiple state transition relationships, the first fault event and the first subsequent operating state corresponding to the second operating state are determined, that is, the state transition relationship corresponding to the second operating state, and then through the test operating state corresponding to the second operating state, it is verified that the second operating state is determined as the location of the state transition relationship, and the second operating state is determined as the target operating state.
[0128] For example, assuming that the second operating state is operating state 322 and operating state 323, see Figure 3 The preset heterogeneous graph shown can determine that the first fault event corresponding to the running state 322 is fault event 5, and the first post-running state is the running state 332. Assuming that the controller is in the running state 322, the fault event 5 is injected to obtain the test running state. Assuming that the test running state is the same as the running state 332, the running state 322 is determined as the target running state.
[0129] In some possible embodiments, if the first fault event and the first post-operation state corresponding to the second operation state are not stored in the multiple state transition relationships, the second operation state is not the target operation state.
[0130] In some possible embodiments, if the target operating state does not exist in the preset heterogeneous graph, a fault alarm prompt is generated to prompt the user to perform manual fault location processing; the target fault information corresponding to the alarm prompt is received; according to the target fault information, the target fault event, and the pre-operating state and post-operating state corresponding to the target fault event are processed; according to the target fault event, the pre-operating state and post-operating state corresponding to the target fault event, the preset heterogeneous graph is updated.
[0131] In the subsequent maintenance process, fault events and processing experience can be continuously accumulated, and the inference rules and preset heterogeneous graphs can be further optimized to improve the fault location accuracy of the controller.
[0132] S603: Determine a target fault event in a preset heterogeneous graph according to the target operating state.
[0133] The target operating state is used as the post-operating state, and the target fault event corresponding to the target operating state is determined in the preset heterogeneous graph.
[0134] S604: According to the target fault event, determine target fault information from fault information corresponding to a plurality of preset fault events.
[0135] The fault information corresponding to multiple preset fault events can be found in Table 4. Table 4 shows the correspondence between some controller faults, preset fault events, and fault types. The actual controller faults, preset fault events, and fault types are not limited to the following table.
[0136] Table 4
[0137]
[0138] The fault location method provided in the embodiment of the present application can actively create a transfer relationship between a preset fault event and an operating state, and can reproduce various fault scenarios that the controller may encounter in actual operation through a preset heterogeneous graph, so as to achieve efficient and accurate fault location through a preset heterogeneous graph after a controller fails, thereby improving the efficiency of fault location. At the same time, by reproducing the fault scenario of the controller in advance through a preset heterogeneous graph, the accuracy of fault location can also be improved.
[0139] Figure 7 This is a schematic diagram of the structure of a fault location device provided in an embodiment of the present application. Figure 7 The fault location device 700 may include a receiving module 701, a first determining module 702, a first acquiring module 703, a second acquiring module 704 and a second determining module 705:
[0140] The receiving module 701 is used to receive a fault location request, where the fault location request is used to request fault information corresponding to the controller;
[0141] The first determination module 702 is used to determine the current execution stage corresponding to the controller according to the fault location request, where the current execution stage is the execution stage of the controller when the fault occurs;
[0142] The first acquisition module 703 is used to acquire the initial operation state of the controller, where the initial operation state is the operation state of the controller before a fault occurs;
[0143] The second acquisition module 704 is used to acquire a preset heterogeneous graph, where the preset heterogeneous graph includes a plurality of state transition relationships, where the state transition relationship is used to indicate the transition of the operation state of the controller affected by the preset fault event;
[0144] The second determination module 705 is used to determine target fault information in a preset heterogeneous graph according to the current execution stage and the initial operation state.
[0145] In some possible embodiments, the second determining module 705 is specifically configured to:
[0146] According to the current execution stage and the initial operation state, the target fault event corresponding to the controller is determined in the preset heterogeneous graph;
[0147] According to the target fault event, target fault information is determined from fault information corresponding to a plurality of preset fault events.
[0148] In some possible embodiments, the preset heterogeneous graph includes multiple execution stages and multiple running states corresponding to each execution stage; the second determination module 705 is specifically used to:
[0149] Determine, according to the multiple running states corresponding to each execution stage, multiple first running states corresponding to the current execution stage;
[0150] Determining a target operating state according to the plurality of first operating states and the initial operating state;
[0151] According to the target operating status, the target fault event is determined in the preset heterogeneous graph.
[0152] In some possible embodiments, the second determining module 705 is specifically configured to:
[0153] Determine at least one second operating state corresponding to the initial operating state among the plurality of first operating states;
[0154] If the number of the at least one second operating state is one, determining the one second operating state as the target operating state;
[0155] If the number of at least one second operating state is plural, a target operating state is determined from the plurality of second operating states according to a preset heterogeneous graph.
[0156] In some possible embodiments, the second determining module 705 is specifically configured to:
[0157] For any second operating state, determining a first fault event and a first post-operating state corresponding to the second operating state in a plurality of state transition relationships;
[0158] Determining a test operation state according to the second operation state and the first fault event;
[0159] If the test operating state is the same as the first post-operating state, the second operating state is determined as the target operating state.
[0160] In some possible embodiments, the first determining module 702 is specifically configured to:
[0161] According to the fault detection request, the fan speed of the cooling fan of the controller at the time of the fault is obtained;
[0162] According to the fan speed of the cooling fan, determine the current load torque corresponding to the cooling fan;
[0163] According to the current load torque, the current execution stage corresponding to the controller is determined.
[0164] In some possible embodiments, the first determining module 702 is specifically configured to:
[0165] According to the fan speed, determine the motor angular velocity corresponding to the cooling fan;
[0166] Determine the electromagnetic torque corresponding to the cooling fan according to the motor data and the motor angular velocity corresponding to the cooling fan;
[0167] The current load torque is determined based on the electromagnetic torque, motor angular velocity and motor inertia.
[0168] In some possible embodiments, the first determining module 702 is specifically configured to:
[0169] Obtain multiple load intervals corresponding to the cooling fan and a control execution stage corresponding to each load interval;
[0170] According to the current load torque and the control execution stage corresponding to each load interval, a current execution stage is determined among multiple control execution stages.
[0171] In some possible embodiments, the second acquisition module 704 is specifically used for:
[0172] Determine a plurality of state transition relationships, the state transition relationships including a pre-operation state, a preset fault event, and a post-operation state;
[0173] Construct a state transition diagram based on multiple state transition relationships;
[0174] According to the execution phase and state transition diagram corresponding to each running state, a preset heterogeneous graph is determined.
[0175] In some possible embodiments, the second acquisition module 704 is specifically used for:
[0176] Obtain multiple pre-operation states and multiple preset fault events of the controller;
[0177] Through a plurality of preset fault events, a fault event test is performed on each preceding operation state to determine at least one subsequent operation state corresponding to each preceding operation state;
[0178] For any preceding operation state, according to at least one subsequent operation state corresponding to the preceding operation state and at least one preset fault event corresponding to at least one subsequent operation state, determining at least one state transition relationship corresponding to the preceding operation state;
[0179] According to at least one state transition relationship corresponding to each preceding running state, a plurality of state transition relationships are determined.
[0180] In some possible embodiments, the second acquisition module 704 is specifically used for:
[0181] Determine the running node corresponding to each running state in the multiple state transition relationships;
[0182] For any state transition relationship, among multiple running nodes, determine the preceding running node corresponding to the preceding running state of the state transition relationship and the following running node corresponding to the following running state, and connect the preceding running node and the following running node through a connecting line to obtain a state transition diagram.
[0183] In some possible embodiments, the second acquisition module 704 is specifically used for:
[0184] According to the execution stage corresponding to each running state, multiple execution stages of the controller are determined, and the execution order corresponding to the multiple execution stages is determined;
[0185] According to the execution order of multiple execution stages, multiple operating states in the state transfer diagram are classified and adjusted to obtain a preset heterogeneous graph.
[0186] For the description of the features in the embodiment corresponding to the fault location device, reference may be made to the relevant description of the embodiment corresponding to the fault location method, which will not be described in detail here.
[0187] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 800 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the electronic device 800 further includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus.
[0188] In a specific implementation process, at least one processor 801 executes the computer execution instructions stored in the memory 802, so that the at least one processor 801 executes the above-mentioned fault location method embodiment.
[0189] The specific implementation process of the processor 801 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.
[0190] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the application may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0191] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.
[0192] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0193] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned fault location method embodiments when running.
[0194] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0195] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned fault location method embodiments are implemented.
[0196] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault location method embodiments are implemented.
[0197] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0198] The above is a detailed introduction to a fault location method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A fault location method, characterized in that: include: receiving a fault location request, wherein the fault location request is used to request target fault information corresponding to the controller; Determine, according to the fault location request, a current execution stage corresponding to the controller, wherein the current execution stage is an execution stage of the controller when the fault occurs; Acquire an initial operating state of the controller, where the initial operating state is an operating state of the controller before a failure occurs; Acquire a preset heterogeneous graph, wherein the preset heterogeneous graph includes a plurality of state transition relationships, and the state transition relationship is used to indicate a transition situation in which a preset fault event affects an operating state of the controller; According to the current execution stage and the initial operation state, the target fault information is determined in the preset heterogeneous graph.
2. The method according to claim 1, characterized in that According to the current execution stage and the initial operation state, determining the target fault information in the preset heterogeneous graph includes: Determining, in the preset heterogeneous graph, a target fault event corresponding to the controller according to the current execution stage and the initial operation state; According to the target fault event, the target fault information is determined from the fault information corresponding to a plurality of preset fault events.
3. The method according to claim 2, characterized in that The preset heterogeneous graph includes multiple execution stages and multiple running states corresponding to each execution stage; According to the current execution stage and the initial operation state, determining a target fault event corresponding to the controller in the preset heterogeneous graph includes: Determine, according to the multiple running states corresponding to each execution stage, multiple first running states corresponding to the current execution stage; Determining a target operating state according to the plurality of first operating states and the initial operating state; According to the target operating state, the target fault event is determined in the preset heterogeneous graph.
4. The method according to claim 3, characterized in that Determining the target operating state according to the plurality of first operating states and the initial operating state includes: Determine, among the plurality of first operating states, at least one second operating state corresponding to the initial operating state; If the number of the at least one second operating state is one, determining the one second operating state as the target operating state; If the number of the at least one second operating state is plural, the target operating state is determined from among the plurality of second operating states according to the preset heterogeneous graph.
5. The method according to claim 4, characterized in that Determining the target operating state among the plurality of second operating states according to the preset heterogeneous graph includes: For any second operating state, determining a first fault event and a first post-operating state corresponding to the second operating state in the plurality of state transition relationships; Determining a test operation state according to the second operation state and the first fault event; If the test operating state is the same as the first post-operating state, the second operating state is determined as the target operating state.
6. The method according to claim 1, characterized in that Determining, according to the fault location request, a current execution stage corresponding to the controller, including: According to the fault detection request, obtaining a fan speed of the cooling fan of the controller at the moment of the fault; Determining a current load torque corresponding to the cooling fan according to a fan speed of the cooling fan; According to the current load torque, a current execution phase corresponding to the controller is determined.
7. The method according to claim 6, characterized in that Determining a current load torque corresponding to the cooling fan according to the fan speed of the cooling fan includes: Determining the motor angular velocity corresponding to the cooling fan according to the fan speed; Determining the electromagnetic torque corresponding to the cooling fan according to the motor data corresponding to the cooling fan and the motor angular velocity; The current load torque is determined according to the electromagnetic torque, the motor angular velocity and the motor rotation inertia.
8. The method according to claim 6, characterized in that Determining a current execution stage corresponding to the controller according to the current load torque includes: Acquire multiple load intervals corresponding to the cooling fan and a control execution stage corresponding to each load interval; The current execution stage is determined in a plurality of control execution stages according to the current load torque and the control execution stage corresponding to each load interval.
9. The method according to claim 1, characterized in that: Get the preset heterogeneous graph, including: Determining the plurality of state transition relationships, the state transition relationships comprising a pre-operation state, a preset fault event, and a post-operation state; Constructing a state transition diagram according to the multiple state transition relationships; The preset heterogeneous graph is determined according to the execution stage corresponding to each running state and the state transition diagram.
10. The method according to claim 9, characterized in that Determine multiple state transition relationships, including: Acquire multiple pre-operation states and multiple preset fault events of the controller; Performing a fault event test on each pre-operation state through the plurality of preset fault events to determine at least one post-operation state corresponding to each pre-operation state; For any preceding operation state, according to at least one subsequent operation state corresponding to the preceding operation state and at least one preset fault event corresponding to at least one subsequent operation state, determining at least one state transition relationship corresponding to the preceding operation state; The multiple state transition relationships are determined according to at least one state transition relationship corresponding to each preceding running state.
11. The method according to claim 9, characterized in that According to the multiple state transition relationships, a state transition diagram is constructed, including: Determine an operating node corresponding to each operating state in the multiple state transition relationships; For any state transition relationship, among multiple running nodes, determine the preceding running node corresponding to the preceding running state and the following running node corresponding to the following running state of the state transition relationship, and connect the preceding running node and the following running node through a connecting line to obtain the state transition diagram.
12. The method according to claim 9, characterized in that Determining the preset heterogeneous graph according to the execution stage corresponding to each running state and the state transition graph includes: Determine a plurality of execution stages of the controller according to the execution stage corresponding to each running state, and determine an execution order corresponding to the plurality of execution stages; According to the execution order of the multiple execution stages, the multiple operating states in the state transition diagram are classified and adjusted to obtain the preset heterogeneous graph.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the fault location method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault location method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault location method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
NVME SSD fault positioning method and device, equipment and medium
CN110688268A
Fault positioning method based on finite-state machine and graph neural network
CN111966076A
Intelligent high-voltage switch cabinet fault diagnosis method and system based on heterogeneous graph structure learning
CN116244617A
Fault root cause positioning method and device, equipment and storage medium
CN117201289A
Urban rail transit system risk monitoring and early warning method based on situation awareness
CN118114975A