Fault monitoring method and device, equipment, storage medium and program product
By deploying the central monitoring process in the distributed cloud storage system, it is possible to report the fault information through the central monitoring process of the second child node when the first child node cannot report the fault information on itself, solving the problem of low fault monitoring reliability in traditional technology and improving the reliability of fault monitoring.
Patent Information
- Application Number
- CN202311823549.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
Smart Images

Figure CN120223590A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distributed cloud storage, and in particular, to a fault monitoring method, device, equipment, storage medium, and program product. Background Art
[0002] With the development of computer technology, distributed cloud storage systems have emerged. A distributed cloud storage system is a storage system composed of multiple nodes capable of cloud storage. In a distributed cloud storage system, there are multiple processes on each node. When these processes encounter problems such as deadlocks, slow threads, memory leaks, and handle leaks, it will cause IO blockage (input / output blockage). Therefore, it is necessary to perform real-time fault monitoring on each process on each node.
[0003] In traditional fault monitoring methods, usually a process is set on each node as a management process. This management process is used to monitor the running status of each process on the node in real time. When a certain process has a running fault, the management process will report the fault situation of this process to a fault management node for handling faults.
[0004] However, the above-mentioned fault monitoring method has the problem of low monitoring reliability. Summary of the Invention
[0005] Based on this, it is necessary to provide a fault monitoring method, device, equipment, storage medium, and program product that can improve the reliability of fault monitoring for the above technical problems.
[0006] In a first aspect, this application provides a fault monitoring method. This method is used for a first sub-node, and the first sub-node is any one of the sub-nodes in the cloud storage system. The method includes:
[0007] For each target service process in the first sub-node, through the target service process, receive a target heartbeat request sent by a second sub-node in the cloud storage system through a central monitoring process. The second sub-node is at least one sub-node other than the first sub-node;
[0008] In response to the target heartbeat request, monitor the running status of the target service process, and send a target heartbeat feedback message to the second sub-node according to the running status. The target heartbeat feedback message is used for fault monitoring.
[0009] In the above embodiments, a central monitoring process is deployed on the second child node. The second child node sends a target heartbeat request to the first child node through the central monitoring process deployed thereon. Even when the first child node is unable to report fault information by itself, the first child node can still report the target heartbeat feedback message for fault monitoring through the second child node, avoiding the problem of low reliability of fault monitoring caused by the inability of the child node to report fault information when the management process for reporting fault information in the child node fails in the traditional technology. The fault monitoring method provided by the embodiments of the present application can improve the reliability of fault monitoring.
[0010] In one of the embodiments, the method further includes:
[0011] Sending a heartbeat request from the central monitoring process in the first child node to each service process in each child node, where the heartbeat request is used to instruct each child node to return a heartbeat feedback message.
[0012] In the above embodiments, when a central monitoring process is deployed on the first child node, the first child node can receive the heartbeat feedback messages returned by each child node through the central monitoring process, enabling the first child node to not only report the fault monitoring results of its own service process but also report the fault monitoring results of the service processes of other child nodes. Similarly, for other second child nodes with a deployed central monitoring process, they can not only report the fault monitoring results of their own service processes but also report the fault monitoring results of the service processes of other child nodes including the first child node. Then, when a central monitoring process has a problem, other central monitoring processes can also report the fault monitoring results, improving the reliability of fault monitoring.
[0013] In one of the embodiments, the method further includes:
[0014] If a heartbeat feedback message sent by any child node is received through the central monitoring process, and the heartbeat feedback message carries the fault monitoring result of the service process in the child node, sending fault information for the service process to the fault management node, where the fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process.
[0015] In this way, in the above embodiments, the first child node with a deployed central monitoring process sends a heartbeat request to each process in each child node through the central monitoring process, and reports the fault monitoring results included in the received heartbeat feedback messages of each process to the fault management node, so that the fault management node can perform fault handling on the faulty process and node specifically based on the fault monitoring result, node identifier, and process identifier, making the monitoring granularity and fault handling granularity of this fault monitoring method at the process level with higher accuracy.
[0016] In one embodiment, monitoring the running status of a target service process and sending a target heartbeat feedback message to a second child node according to the running status includes:
[0017] Performing status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result;
[0018] If the monitoring result is a fault monitoring result, carrying the fault monitoring result in the target heartbeat feedback message and sending it to the second child node.
[0019] In the above embodiment, when there is a fault in the running status of the target service process, the obtained monitoring result will be a fault monitoring result. Carrying the fault monitoring result in the target heartbeat feedback message and sending it to the second child node can timely report the faults existing in the target service process and improve the reliability of fault monitoring.
[0020] In one embodiment, performing status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result includes:
[0021] Performing a fault detection task through a local fault monitoring program deployed in the target service process;
[0022] In the case where the fault detection task fails, determining that the monitoring result is a fault monitoring result.
[0023] In the above embodiment, performing the fault detection task through the local fault detection program in the target service process avoids unnecessary data transmission and improves the efficiency of fault monitoring.
[0024] In a second aspect, the present application provides a fault monitoring method, which is used for a fault management node. The method includes:
[0025] Receiving target fault information sent by a second child node. The target fault information includes the node identifier of a first child node, the process identifier of a target service process with a fault in the running status in the first child node, and a fault monitoring result. The fault monitoring result is carried in a target heartbeat feedback message sent by the first child node in response to a target heartbeat request sent by the second child node through a central monitoring process and monitoring the running status of the target service process and sending it to the second child node according to the running status;
[0026] Performing fault handling on the target service process according to the target fault information.
[0027] In the above embodiment, the fault management node performs fault handling on the target service process according to the target fault information, making the fault handling more accurate.
[0028] In one embodiment, performing fault handling on a target service process according to target fault information includes:
[0029] Performing a restart process on the target service process according to the process identifier included in the target fault information.
[0030] In the above embodiment, the fault management node determines the target service process with a fault and the first sub-node to which the target service process belongs according to the target fault information sent by the received second sub-node, and performs fault handling on the target service process and the first sub-node level. First, the target service process is processed first, so that the granularity of fault handling is at the service process level, and the granularity of fault handling is finer.
[0031] In one embodiment, the method further includes:
[0032] If updated fault information for the target service process is received, then performing a restart process or an isolation process on the first sub-node according to the node identifier included in the updated fault information.
[0033] In the above embodiment, fault handling is performed on the first sub-node only when the target service process has not recovered after fault handling, avoiding the expansion of the impact of fault handling.
[0034] In a third aspect, the present application provides a fault monitoring device, which is used for a first sub-node. The first sub-node is any one of the sub-nodes in a cloud storage system, and the first sub-node includes a plurality of target service processes. The device includes:
[0035] A receiving module, configured to, for each target service process in the first sub-node, receive, through the target service process, a target heartbeat request sent by a second sub-node in the cloud storage system through a central monitoring process. The second sub-node is at least one sub-node other than the first sub-node;
[0036] A response module, configured to, in response to the target heartbeat request, monitor the running state of the target service process, and send a target heartbeat feedback message to the second sub-node according to the running state. The target heartbeat feedback message is used for fault monitoring.
[0037] In a fourth aspect, the present application provides a fault monitoring device, which is used for a fault management node. The device includes:
[0038] A receiving module, configured to receive target fault information sent by a second child node, where the target fault information includes a node identifier of a first child node, a process identifier of a target service process with a fault in the running state in the first child node, and a fault monitoring result. The fault monitoring result is carried in a target heartbeat feedback message sent by the first child node to the second child node in response to a target heartbeat request sent by the second child node through a central monitoring process, where the target heartbeat request is used to monitor the running state of the target service process;
[0039] A fault handling module, configured to perform fault handling on the target service process according to the target fault information.
[0040] In a fifth aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the methods in the first aspect and the second aspect are implemented.
[0041] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods in the first aspect and the second aspect are implemented.
[0042] In a seventh aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the methods in the first aspect and the second aspect are implemented.
[0043] For each target service process in the first child node, the above-mentioned fault monitoring method, device, equipment, storage medium, and program product receive a target heartbeat request sent by a second child node in a cloud storage system through a central monitoring process by means of the target service process. The second child node is at least one child node other than the first child node. Then, in response to the target heartbeat request, the running state of the target service process is monitored, and a target heartbeat feedback message is sent to the second child node according to the running state. The target heartbeat feedback message is used for fault monitoring. In this way, by deploying a central monitoring process on the second child node, and the second child node sends a target heartbeat request to the first child node through the central monitoring process deployed thereon, even when the first child node cannot report fault information by itself, the first child node can still report the target heartbeat feedback message for fault monitoring through the second child node, avoiding the problem of low reliability of fault monitoring caused by the inability of a child node to report fault information when a management process for reporting fault information in the child node fails in the traditional technology. The fault monitoring method, device, equipment, storage medium, and program product provided by the embodiments of the present application can improve the reliability of fault monitoring. Description of the Drawings
[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 It is a structural schematic diagram of a traditional fault monitoring method;
[0046] Figure 2 It is an application environment diagram of the fault monitoring method in an embodiment;
[0047] Figure 3 It is a flowchart of the fault monitoring method applied to the first sub-node in an embodiment;
[0048] Figure 4 It is a flowchart of the fault monitoring method executed by the central monitoring process in the case where the first sub-node has a central monitoring process in another embodiment;
[0049] Figure 5 It is a structural schematic diagram of the first sub-node having a central monitoring process in another embodiment;
[0050] Figure 6 It is a flowchart of step 302 in another embodiment;
[0051] Figure 7 It is a flowchart of step 601 in another embodiment;
[0052] Figure 8 It is a flowchart of the fault monitoring method applied to the fault management node in an embodiment;
[0053] Figure 9 It is a flowchart of the exemplary fault monitoring method in an embodiment;
[0054] Figure 10 It is a structural block diagram of the fault monitoring device applied to the first sub-node in an embodiment;
[0055] Figure 11 It is a structural block diagram of the fault monitoring device applied to the fault management node in an embodiment;
[0056] Figure 12 It is an internal structural diagram of a computer device in an embodiment. Detailed implementation manners
[0057] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0058] With the development of computer technology, distributed cloud storage systems have emerged. A distributed cloud storage system is a storage system composed of multiple nodes capable of cloud storage. There are multiple nodes in a distributed cloud storage system, and there are multiple processes on each node. When these processes have problems such as deadlocks, slow threads, memory leaks, and handle leaks, it will cause IO blockages (input / output blockages). IO blockages will cause the nodes to retransmit data with a timeout, seriously wasting system resources and network bandwidth, reducing the performance of the system. When the IO blockage is severe, the nodes cannot transmit data, which will cause the upper-layer services to be interrupted. Therefore, it is necessary to perform real-time fault monitoring on each process on each node and timely handle the nodes with faults to ensure the stability of the system.
[0059] Traditional fault monitoring methods, referring to Figure 1 , usually set a process as a management process on each node. This management process is used to monitor the running status of each process on the node in real time. When a certain process has a running fault, the management process will report the fault situation of this process to the fault management node for handling faults. However, if there is a problem with this management process, the faults of this node cannot be reported to the fault management node, and the fault management node cannot perceive the abnormality of this node, resulting in a false dead phenomenon.
[0060] Therefore, the above-mentioned fault monitoring method has the problem of low monitoring reliability.
[0061] In view of this, embodiments of the present application provide a fault monitoring method, device, equipment, storage medium, and program product. For each target service process in the first sub-node, through the target service process, a target heartbeat request sent by a second sub-node in the cloud storage system through a central monitoring process is received. The second sub-node is at least one sub-node other than the first sub-node. Then, in response to the target heartbeat request, the running state of the target service process is monitored, and a target heartbeat feedback message is sent to the second sub-node according to the running state. The target heartbeat feedback message is used for fault monitoring. In this way, a central monitoring process is deployed on the second sub-node, and the second sub-node sends a target heartbeat request to the first sub-node through the central monitoring process deployed thereon. Even when the first sub-node cannot report fault information by itself, the first sub-node can still report the target heartbeat feedback message for fault monitoring through the second sub-node, avoiding the problem of low reliability of fault monitoring caused by the inability of the sub-node to report fault information when the management process for reporting fault information in the sub-node fails in the traditional technology. The fault monitoring method, device, equipment, storage medium, and program product provided by the embodiments of the present application can improve the reliability of fault monitoring.
[0062] The fault monitoring method provided by the embodiments of the present application can be applied to an application environment as Figure 2 shown. Among them, the fault management node 202 in the cloud storage system communicates with each sub-node 204 in the cloud storage system through a network. Among them, the fault management node 202 and each sub-node 204 can both be servers, and the server can be a single server or a server cluster composed of multiple servers.
[0063] In an exemplary embodiment, as Figure 3 shown, a fault monitoring method is provided. Taking the method applied to the first sub-node ( Figure 2 any one of the sub-nodes 204) as an example for description, where the first sub-node is any one of the sub-nodes in the cloud storage system. This embodiment includes step 301 and step 302. Among them:
[0064] Step 301, for each target service process in the first sub-node, through the target service process, a target heartbeat request sent by a second sub-node in the cloud storage system through a central monitoring process is received.
[0065] There are multiple sub-nodes in the cloud storage system. For one of the sub-nodes, which is called the first sub-node, there are multiple target service processes running on the first sub-node. In order to be able to monitor the running state of each target service process in real time, in a possible implementation manner, a local fault detection program is deployed on each target service process, and the first sub-node can monitor the running state of each target service program through the local fault detection program on each target service process.
[0066] When monitoring the failure of each target service process of the first child node, the first child node receives the target heartbeat requests corresponding to the target service processes through the local failure detection programs on the target service processes and responds to the target heartbeat requests. In a possible implementation, the first child node periodically receives the target heartbeat requests sent by multiple second child nodes through the local failure detection programs on the target service processes, and the target heartbeat requests are sent by the second child nodes through the central monitoring process.
[0067] In the embodiment of the present application, the second child node is at least one child node other than the first child node. Many target service processes also run on the second child node. A central monitoring program is deployed on one of the target service processes. The target service process on which the central monitoring program is deployed is the central monitoring process. The second child node can receive the messages of the target service processes of each child node including itself through the central monitoring process. In a possible implementation, the second child node selects a specific target service process to deploy the central monitoring program, and the local failure detection program is not deployed on the target service process, so that the target service process can receive the messages of the target service processes of each child node other than itself; in another possible implementation, the second child node randomly selects a target service process to deploy the central monitoring program, and both the central monitoring program and the local failure detection program are deployed on the target service process.
[0068] In a possible implementation, when the second child node includes multiple child nodes, the target heartbeat requests sent by each second child node through the central monitoring process also carry the node identifier of the second child node and the process identifier of the central monitoring process.
[0069] Step 302, in response to the target heartbeat request, monitor the running status of the target service process, and send a target heartbeat feedback message to the second child node according to the running status.
[0070] The target heartbeat request is used to enable the first child node that receives the target heartbeat request to return a corresponding heartbeat feedback message to the central monitoring process corresponding to each target heartbeat request through each target service process. Therefore, in a possible implementation, the target heartbeat request also includes the node identifier of the second child node that sends the target heartbeat request and the process identifier of the central monitoring process. Similarly, the heartbeat feedback message also carries the node identifier of the first child node that returns the heartbeat feedback message and the process identifier of the target service process.
[0071] During the process of fault monitoring, the first child node can monitor the running status of the target service process through the local fault detection program on the target service process, and then generate a target heartbeat feedback message sent to the second child node according to the running status, where the target heartbeat feedback message is used for fault monitoring.
[0072] In the case where there is a fault in the running status, the target heartbeat feedback message includes the fault monitoring result of the target service process, and the fault monitoring result is used to characterize that the running status of the target service process is abnormal. In a possible implementation manner, the fault monitoring result further includes the specific reason for the abnormality of the target service process. And in the case where there is a fault in the running status, when the first child node sends the target heartbeat feedback message including the fault monitoring result to the second child node through the target service process, so that the second child node sends the target fault information for the target service process to the management node according to the target heartbeat feedback message.
[0073] The target fault information is generated by the second child node according to the fault monitoring result. The target fault information includes the node identifier of the corresponding first child node, the process identifier of the target service process, and the fault monitoring result. The target fault information is used for the fault management node to perform fault handling on the first child node and the target service process with abnormal running status after receiving the target fault information.
[0074] The fault management node is a node in the cloud storage system used to manage the faults of each child node. In a possible implementation manner, the fault management process running on the fault management node is used to perform fault handling on the target fault information sent by the second child node.
[0075] For each target service process, the above-mentioned fault monitoring method receives, through the target service process, the target heartbeat request sent by the second child node in the cloud storage system through the central monitoring process. The second child node is at least one child node other than the first child node. Then, in response to the target heartbeat request, it monitors the running status of the target service process and sends a target heartbeat feedback message to the second child node according to the running status, where the target heartbeat feedback message is used for fault monitoring. In this way, the central monitoring process is deployed on the second child node, and the second child node sends the target heartbeat request to the first child node through the central monitoring process deployed thereon. Even when the first child node cannot report fault information by itself, the first child node can still report the target heartbeat feedback message for fault monitoring through the second child node, avoiding the problem of low reliability of fault monitoring caused by the inability of the child node to report fault information when the management process for reporting fault information in the child node fails in the traditional technology. The fault monitoring method provided by the embodiments of the present application can improve the reliability of fault monitoring.
[0076] In one embodiment, based on the above Figure 3 illustrated embodiment, refer to Figure 4 . This embodiment relates to the process of executing a fault monitoring method through a central monitoring process when there is a central monitoring process in the first child node. As Figure 4 shown, this process may include step 401.
[0077] Step 401, send a heartbeat request to each business process in each child node through the central monitoring process in the first child node.
[0078] Since the second child node includes at least one child node different from the first child node, optionally, the second child node is a child node other than the first child node, and optionally, the second child node is the first child node itself, that is, when a central monitoring program is deployed on a target business process of the first child node, then this target business process can work as a central monitoring process.
[0079] Refer to Figure 5 . When a central monitoring program is deployed on any target business node of the first child node, then this business process is the central monitoring process. The first child node can send a heartbeat request to the business processes of each child node through this central monitoring process, and will also receive the heartbeat feedback information returned by each business process of each node in response to the heartbeat request.
[0080] And a local monitoring program is deployed on the target business process. Therefore, in a possible implementation manner, both a local fault monitoring program and a central monitoring program can be deployed on a target business process. When there is a central monitoring process in the first child node, when the first child node sends a heartbeat request to each business process in each child node through the central monitoring process, the business processes include this central monitoring process.
[0081] In a possible implementation manner, since the number of second child nodes is multiple, the number of central monitoring processes is also multiple. Therefore, the heartbeat request sent by the central monitoring process should also include the process identifier of the central monitoring process that sends this heartbeat request, and the node identifier of the child node to which this central monitoring process belongs.
[0082] The heartbeat request is used to instruct each child node to return heartbeat feedback information. The first child node with a central monitoring process set can detect whether there are fault monitoring results of each business process of each child node in this heartbeat feedback message through the central monitoring process.
[0083] In a possible implementation, if the central monitoring process receives heartbeat feedback information sent by any child node, and the heartbeat feedback message carries the fault monitoring result of the service process in the child node, the fault information for the service process is sent to the fault management node. The fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process, so that the fault management node can perform fault handling on the faulty process and node specifically based on the fault monitoring result, node identifier, and process identifier.
[0084] For the specific method of fault handling, refer to the embodiments described later.
[0085] In this way, in the above embodiments, the first child node deployed with the central monitoring process sends heartbeat requests to each process of each child node through the central monitoring process, and reports the fault monitoring results included in the heartbeat feedback messages received from each process to the fault management node, so that the fault management node can perform fault handling on the faulty process and node specifically based on the fault monitoring result, node identifier, and process identifier, making the monitoring granularity and fault handling granularity of this fault monitoring method at the process level and having higher accuracy.
[0086] In one embodiment, based on the above Figure 3 illustrated embodiment, refer to Figure 6 , this embodiment relates to the process of monitoring the running state of the target service process and sending a target heartbeat feedback message to the second child node according to the running state. As Figure 6 illustrated, step 302 may include step 601 and step 602.
[0087] Step 601, perform status detection on the target service process through the local fault detection program deployed in the target service process to obtain a monitoring result.
[0088] The local fault detection program is a program deployed in the target service process. The first child node can run the local fault detection program to perform status detection on the target service program, thereby obtaining a monitoring result. Optionally, if the target service process is running normally, the monitoring result is normal running state; optionally, if the running state of the target service process is abnormal, the monitoring result is the fault monitoring result used to characterize the specific abnormal information of the target service process.
[0089] Since the traditional fault monitoring method has a relatively single monitoring scheme when monitoring each node, it cannot detect problems such as thread deadlocks, memory fragmentation, and handle leaks on the node. In a possible implementation manner, the first sub-node can use a local fault detection program to perform status detection on the running status of the target service process from multiple aspects, such as thread deadlock detection, handle leak detection, and running exception detection. The obtained fault monitoring results may include at least one of thread deadlock faults, handle leak faults, and running exception faults.
[0090] Step 602: If the monitoring result is a fault monitoring result, carry the fault monitoring result in the target heartbeat feedback message and send it to the second sub-node.
[0091] When the monitoring result is a fault monitoring result, it indicates that the running status of the target service process is abnormal. Then, when the first sub-node responds to the target heartbeat request of the second sub-node, it will carry the fault monitoring result in the target feedback message and feedback it to the second sub-node.
[0092] In this way, in the above embodiments, the running status of the target service process is detected from multiple aspects, enriching the monitoring scheme and making the fault monitoring results more accurate.
[0093] In one embodiment, based on the above Figure 6 illustrated embodiment, see Figure 7 , this embodiment relates to the process of performing status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result. As Figure 7 shown, step 601 may include step 701 and step 702.
[0094] Step 701: Execute a fault detection task through a local fault monitoring program deployed in the target service process.
[0095] Step 702: If the fault detection task fails, determine that the monitoring result is a fault monitoring result.
[0096] In the embodiments of the present application, the first sub-node can execute different fault detection tasks by running a local fault detection program, thereby completing multi-aspect detection of the target service process.
[0097] Optionally, the fault detection task may include a thread deadlock test task, which may be a preset task for testing the target business process, such as opening a file, reading test data, etc. The local fault monitoring program may periodically deliver test tasks to the target business process. For example, it delivers test tasks to the target business process every 2s, and at preset time intervals, it detects whether the target business process has completed each test task delivered within the preset time interval. If none of the test tasks are completed, it determines that the monitoring result is a fault monitoring result, and the fault monitoring result includes a thread deadlock fault.
[0098] Exemplarily, the first child node delivers test tasks to the target business process every 2s through the local fault monitoring program, and the preset time interval is 10s. Then, within 10s, the first child node delivers 5 test tasks to the target business process through the local fault monitoring program. If none of the 5 test tasks are completed within these 10s, it can be determined that there is a thread deadlock problem in the target business process.
[0099] Optionally, the fault detection task may include a handle leak test task. The specific content of the handle leak test task may be to instruct the target business process to open a preset handle file through the local fault monitoring program. The handle file may be a preset test file. When the target business process can successfully allocate a handle and open the preset handle file, it indicates that there is no handle leak fault in the target business process; otherwise, it indicates that there is a handle leak fault in the target business process, and it determines that the monitoring result is a fault monitoring result, and the fault monitoring result includes a handle leak fault. In a possible implementation, to avoid the first child node misjudging that the target business process has a handle leak fault due to a single failed handle leak test task by chance, the first child node periodically executes the handle leak test task through the local fault monitoring program deployed in the target business process. For example, every 10s, it instructs the target business process to open the preset handle file through the local fault monitoring program, sets a preset failure times threshold. If the consecutive failure times of opening the preset handle file reach the preset failure times threshold, it determines that the monitoring result is a fault monitoring result, and the fault monitoring result includes a handle leak fault.
[0100] Exemplarily, the preset failure times threshold is 3 times. Then, when the target business process fails to open the preset handle file normally for 3 consecutive times, it can be determined that there is a handle leak fault in the target business process, the monitoring result is a fault monitoring result, and the fault monitoring result includes a handle leak fault.
[0101] Optionally, the fault detection task may include an anomaly detection task.
[0102] In a possible implementation, the target business process may have running exception faults. For example, abnormal CPU occupancy, abnormal memory occupancy, abnormal handle usage, and abnormal business execution, etc. Therefore, the exception detection tasks that the first child node needs to execute through the local fault monitoring program may include at least one of an abnormal CPU occupancy detection task, an abnormal memory occupancy detection task, an abnormal handle usage detection task, and an abnormal business execution detection task.
[0103] Regarding the process of the local fault monitoring program executing the exception detection task, optionally, when the first child node receives an exception detection instruction, it may execute the exception detection task through the local fault detection program according to the process identifier of the target business process carried in the exception detection instruction; optionally, the first child node periodically executes the exception detection task through the local fault monitoring program, so as to perform exception detection on the target business process.
[0104] Regarding the execution processes of different exception detection tasks, the following gives an exemplary introduction:
[0105] Abnormal CPU occupancy detection task: The first child node monitors the CPU occupancy of the target business process through the local fault monitoring program, and judges whether there is an abnormality in the CPU occupancy of the target business process according to the preset CPU occupancy threshold. If the CPU occupied by the target business process exceeds the preset CPU occupancy threshold, it is determined that the target business process has an abnormal CPU occupancy.
[0106] Abnormal memory occupancy detection task: The first child node monitors the memory occupancy of the target business process through the local fault monitoring program, and judges whether there is an abnormality in the memory occupancy of the target business process according to the preset memory occupancy threshold. If the memory occupied by the target business process exceeds the preset memory occupancy threshold, it is determined that the target business process has an abnormal memory occupancy.
[0107] Abnormal handle usage detection task: The first child node instructs the target business process to open a preset handle file through the local fault monitoring program, and monitors the time required for the target business process to open the preset handle file. A preset handle opening time threshold. If the time required for the target business process to open the preset handle file exceeds the preset handle opening time threshold, it is determined that the target business process has an abnormal handle usage.
[0108] Abnormal business execution detection task: The first child node delivers a test task to the target business process through the local fault monitoring program, and monitors the time required for the target business process to execute the test task. A preset task execution time threshold. If the time required for the target business process to execute the test task exceeds the preset task execution time threshold, it is determined that the target business process has an abnormal business execution.
[0109] In the above embodiments, the CPU occupancy anomaly detection task, the memory occupancy anomaly detection task, the handle usage anomaly detection task, and the service execution anomaly detection task can be executed simultaneously or step by step, or one of them can be selected for execution.
[0110] When the execution result of the above anomaly detection task indicates an execution anomaly, it means that the anomaly detection task fails, and the monitoring result is determined as a fault monitoring result.
[0111] In the above embodiments, the local fault detection program in the target service process is used to execute the fault detection task, avoiding unnecessary data transmission and improving the fault monitoring efficiency.
[0112] In one embodiment, as Figure 8 shown, a fault monitoring method is provided. Taking the case where this method is applied to the Figure 2 fault management node 202 as an example, it includes the following steps:
[0113] Step 801, receive the target fault information sent by the second sub-node.
[0114] The second sub-node is a sub-node provided with a central monitoring process. When the second sub-node detects that the heartbeat feedback information returned by each service process of each sub-node carries a fault monitoring result, the fault management node 202 will receive the target fault information sent by the second sub-node.
[0115] Among them, the target fault information includes the node identifier of the first sub-node, the process identifier of the target service process with a fault in the running state in the first sub-node, and the fault monitoring result. The fault monitoring result is carried in the target heartbeat feedback message sent by the first sub-node to the second sub-node in response to the target heartbeat request sent by the second sub-node through the central monitoring process and monitoring the running state of the target service process.
[0116] Regarding the method for how the second sub-node obtains the target fault information, refer to the relevant description in the above embodiments and will not be elaborated here.
[0117] The fault management node is a node different from other sub-nodes in the cloud storage system. In a possible implementation manner, a fault management process is set on the fault management node, and the fault management node can manage the faults of each process on each sub-node through this fault management process and perform fault handling on each process.
[0118] Step 802, perform fault handling on the target service process according to the target fault information.
[0119] When the fault management node receives the target fault information, it can determine the first sub-node with a fault according to the node identifier included in the target fault information, then determine the target service process with a fault on the first sub-node according to the process identifier, and can determine the specific fault existing in the target service process according to the fault monitoring result.
[0120] In a possible implementation manner, after receiving the target fault information, the fault management node displays the target fault information through a display device to prompt the user that there is a specific fault problem in the target service process of the first sub-node corresponding to the target fault information. In another possible implementation manner, after determining the target service process with a fault and the first sub-node to which the target service process belongs, the fault management node performs fault handling on the target service process.
[0121] Regarding the process of fault handling, the following gives an exemplary introduction:
[0122] When the fault management node receives the target fault information, it can determine the target service process corresponding to the target fault information and the first sub-node to which the target service process belongs. The fault management node performs a restart process on the target service process according to the process identifier included in the target fault information.
[0123] In a possible implementation manner, after the fault management node performs a restart process on the target service process, it generates a process fault to be resolved identifier, which also includes the process identifier of the target service process and the node identifier of the first sub-node, and a preset time threshold. If within the preset time threshold, the node identifier and process identifier included in the subsequent fault information received by the fault management node do not match the node identifier and process identifier included in the process fault to be resolved identifier, it is determined that the fault existing in the target service process has been resolved, and the process fault to be resolved identifier is eliminated. Otherwise, it is determined that the fault existing in the target service process has not been resolved, the first sub-node to which the target service process belongs is restarted, and according to the preset time threshold, it is detected whether the third-round target fault information for the target service process reported by the second sub-node can be received within the preset time threshold. If received, the first sub-node to which the target service process belongs is isolated.
[0124] In another possible implementation, if there are still problems after the target business process is restarted, the fault management node will receive the updated fault information of the target business process uploaded by the second child node. If the updated fault information of the target business process is received, the first child node will be restarted according to the node identifier included in the updated fault information. If the fault management node can still receive three rounds of updated fault information for the target business process after restarting the first child node, the fault management node will isolate the first child node.
[0125] Regarding the isolation process, in one possible implementation, the fault management node generates an isolation instruction according to the node identifier of the first child node, and sends the isolation instruction to each child node to instruct each child node not to communicate with the first child node anymore, and displays the node identifier of the first child node on the display interface to prompt the user that there is a fault that cannot be processed on the first child node, and prompt the user to perform subsequent manual fault handling on the first child node.
[0126] In this way, in the above embodiment, the fault management node determines the target business process with a fault and the first child node to which the target business process belongs according to the target fault information sent by the received second child node, and performs fault handling on the target business process and the first child node level. First, the target business process is processed. If the target business process has not recovered after the fault handling, the first child node is then fault handled, so that the granularity of the fault handling is at the business process level, avoiding the expansion of the impact of the fault handling.
[0127] In one embodiment, please refer to Figure 9 , which shows a flowchart of an exemplary fault monitoring method provided by an embodiment of the present application. This method can be applied to Figure 2 the implementation environment shown.
[0128] Step 901, the first child node sends a heartbeat request to each business process in each child node through the central monitoring process in the first child node.
[0129] Among them, the heartbeat request is used to instruct each child node to return heartbeat feedback information.
[0130] Step 902, for each target business process in the first child node, the first child node receives the target heartbeat request sent by the second child node in the cloud storage system through the central monitoring process through the target business process.
[0131] Among them, the second child node includes at least one child node different from the first child node.
[0132] Step 903, the first child node responds to the target heartbeat request.
[0133] Step 904: The first child node periodically delivers test tasks to the target service process through the local fault monitoring program deployed in the target service process.
[0134] Step 905: The first child node detects, at a preset time interval, whether the target service process has completed each test task delivered within the preset time interval. If none of the test tasks are completed, it determines that the monitoring result is a fault monitoring result.
[0135] Among them, the fault monitoring result includes a thread deadlock fault.
[0136] Step 906: The first child node periodically opens a preset handle file through the local fault monitoring program deployed in the target service process.
[0137] Step 907: If the consecutive failure count of opening the preset handle file by the first child node reaches the preset failure count threshold, it determines that the monitoring result is a fault monitoring result.
[0138] Among them, the fault monitoring result includes a handle leak fault.
[0139] Step 908: The first child node performs anomaly detection on the target service process through the local fault monitoring program deployed in the target service process.
[0140] Among them, the anomaly detection includes at least one of CPU occupancy anomaly detection, memory occupancy anomaly detection, handle usage anomaly detection, and service execution anomaly detection.
[0141] Step 909: When the anomaly detection result indicates that there is an anomaly in the target service process, the first child node determines that the monitoring result is a fault monitoring result.
[0142] Among them, the fault monitoring result includes a running anomaly fault.
[0143] Step 910: If the monitoring result is a fault monitoring result, the first child node sends the fault monitoring result in the target heartbeat feedback message to the second child node.
[0144] Among them, the fault monitoring result includes at least one of a thread deadlock fault, a handle leak fault, and a running anomaly fault.
[0145] Among them, when there is a fault in the running state, the target heartbeat feedback message includes the fault monitoring result of the target service process, so that the second child node sends target fault information for the target service process to the fault management node according to the target heartbeat feedback message.
[0146] Step 911: If the first child node receives the heartbeat feedback information sent by any child node through the central monitoring process, and the heartbeat feedback message carries the fault monitoring result of the service process in the child node, send the fault information about the service process to the fault management node.
[0147] Among them, the fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process.
[0148] Step 912: The fault management node receives the target fault information sent by the second child node.
[0149] Among them, the target fault information includes the node identifier of the first child node, the process identifier of the target service process with a fault in the running state in the first child node, and the fault monitoring result.
[0150] Among them, the fault monitoring result is carried in the target heartbeat feedback message sent by the first child node to the second child node in response to the target heartbeat request sent by the second child node through the central monitoring process, by monitoring the running state of the target service process.
[0151] Step 913: The fault management node performs a restart process on the target service process according to the process identifier included in the target fault information.
[0152] Step 914: If the fault management node receives the updated fault information about the target service process, it performs a restart process or an isolation process on the first child node according to the node identifier included in the updated fault information.
[0153] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0154] Based on the same inventive concept, an embodiment of the present application further provides a fault monitoring device for implementing the above-mentioned fault monitoring method. Among them, the fault monitoring device is used for a first sub-node, and the first sub-node is any one of the sub-nodes in the cloud storage system. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the following fault monitoring devices can refer to the limitations on the fault monitoring method in the foregoing, and will not be elaborated herein.
[0155] In an exemplary embodiment, as Figure 10 shown, a fault monitoring device 1000 is provided, including: a receiving module 1001 and a response module 1002, where:
[0156] The receiving module 1001 is configured to, for each target service process in the first sub-node, receive a target heartbeat request sent by a second sub-node in the cloud storage system through the central monitoring process through the target service process, where the second sub-node is at least one sub-node other than the first sub-node;
[0157] The response module 1002 is configured to, in response to the target heartbeat request, monitor the running state of the target service process and send a target heartbeat feedback message to the second sub-node according to the running state;
[0158] Among them, the target heartbeat feedback message is used for fault monitoring.
[0159] In one embodiment, the fault monitoring device 1000 further includes:
[0160] A heartbeat sending module, configured to send a heartbeat request to each service process in each sub-node through the central monitoring process in the first sub-node, where the heartbeat request is used to instruct each sub-node to return heartbeat feedback information.
[0161] In one embodiment, the heartbeat sending module further includes:
[0162] A fault reporting unit, configured to, if receiving heartbeat feedback information sent by any sub-node through the central monitoring process, when the heartbeat feedback message carries the fault monitoring result of the service process in the sub-node, send fault information about the service process to the fault management node, where the fault information includes the fault monitoring result of the service process, the node identifier of the sub-node, and the process identifier of the service process.
[0163] In one embodiment, the response module 1002 includes:
[0164] The detection unit is configured to perform status detection on the target service process through a local fault detection program deployed in the target service process, and obtain a monitoring result;
[0165] The fault result sending unit is configured to, if the monitoring result is a fault monitoring result, send the fault monitoring result to the second sub-node in the target heartbeat feedback message.
[0166] In one embodiment, the detection unit is further configured to execute:
[0167] Execute a fault detection task through a local fault monitoring program deployed in the target service process;
[0168] In the case where the fault detection task fails, determine that the monitoring result is a fault monitoring result.
[0169] Based on the same inventive concept, an embodiment of the present application further provides a fault monitoring device for implementing the above-mentioned fault monitoring method, wherein the fault monitoring device is used for a fault management node. The implementation solutions provided by the device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the fault monitoring device provided below can refer to the limitations on the fault monitoring method in the above text, and will not be repeated here.
[0170] In an exemplary embodiment, as Figure 11 shown, a fault monitoring device 1100 is provided, including: a fault information receiving module 1101 and a response module 1102, wherein:
[0171] The fault information receiving module 1101 is configured to receive target fault information sent by the second sub-node, where the target fault information includes the node identifier of the first sub-node, the process identifier of the target service process with a fault in the running state in the first sub-node, and a fault monitoring result, where the fault monitoring result is carried in a target heartbeat feedback message sent by the first sub-node in response to a target heartbeat request sent by the second sub-node through a central monitoring process to monitor the running state of the target service process;
[0172] The fault handling module 1102 is configured to perform fault handling on the target service process according to the target fault information.
[0173] In one embodiment, the fault handling module 1102 includes:
[0174] The process fault handling unit is configured to perform a restart process on the target service process according to the process identifier included in the target fault information.
[0175] In one embodiment, the fault handling module 1102 includes:
[0176] A node fault handling unit, configured to, if receiving updated fault information for the target service process, restart or isolate the first sub-node according to the node identifier included in the updated fault information.
[0177] Each module in the above-mentioned fault monitoring device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0178] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 12 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store fault monitoring data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a fault monitoring method.
[0179] Those skilled in the art can understand that Figure 12 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0180] In an exemplary embodiment, a computer device is provided. In a possible implementation manner, the computer device is a first sub-node, and the first sub-node is any one of the sub-nodes in the cloud storage system, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0181] For each of the target service processes in the first child node, receive, through the target service process, a target heartbeat request sent by a second child node in the cloud storage system through a central monitoring process, where the second child node is at least one child node other than the first child node;
[0182] In response to the target heartbeat request, monitor the running status of the target service process, and send a target heartbeat feedback message to the second child node according to the running status;
[0183] The target heartbeat feedback message is used for fault monitoring.
[0184] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0185] Send a heartbeat request to each service process in each child node through the central monitoring process in the first child node, where the heartbeat request is used to instruct each child node to return heartbeat feedback information.
[0186] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0187] If heartbeat feedback information sent by any child node is received through the central monitoring process, and in the case where the heartbeat feedback message carries the fault monitoring result of the service process in the child node, send fault information about the service process to the fault management node, where the fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process.
[0188] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0189] Perform status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result;
[0190] If the monitoring result is a fault monitoring result, carry the fault monitoring result in the target heartbeat feedback message and send it to the second child node.
[0191] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0192] Execute a fault detection task through a local fault monitoring program deployed in the target service process;
[0193] In the case where the fault detection task fails, determine that the monitoring result is a fault monitoring result.
[0194] In an exemplary embodiment, a computer device is provided. In a possible implementation, the computer device is a fault management node, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0195] Receive target fault information sent by a second sub-node. The target fault information includes the node identifier of a first sub-node, the process identifier of a target service process with a fault in the running state in the first sub-node, and a fault monitoring result. The fault monitoring result is carried in a target heartbeat feedback message sent by the first sub-node to the second sub-node according to the running state after monitoring the running state of the target service process in response to a target heartbeat request sent by the second sub-node through a central monitoring process.
[0196] Perform fault handling on the target service process according to the target fault information.
[0197] In an embodiment, when the processor executes the computer program, the following steps are further implemented:
[0198] Perform a restart process on the target service process according to the process identifier included in the target fault information.
[0199] In an embodiment, when the processor executes the computer program, the following steps are further implemented:
[0200] If updated fault information for the target service process is received, perform a restart process or an isolation process on the first sub-node according to the node identifier included in the updated fault information.
[0201] In an embodiment, a computer-readable storage medium is provided. In a possible implementation, the computer-readable storage medium is applied to a first sub-node, where the first sub-node is any one of the sub-nodes in a cloud storage system. A computer program is stored thereon. When the computer program is executed by a processor, the following steps are implemented:
[0202] For each target service process in the first sub-node, receive a target heartbeat request sent by a second sub-node in the cloud storage system through a central monitoring process through the target service process. The second sub-node is at least one sub-node other than the first sub-node.
[0203] In response to the target heartbeat request, monitor the running state of the target service process and send a target heartbeat feedback message to the second sub-node according to the running state.
[0204] The target heartbeat feedback message is used for fault monitoring.
[0205] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0206] Send a heartbeat request from the central monitoring process in the first sub-node to each business process in each sub-node, where the heartbeat request is used to instruct each sub-node to return heartbeat feedback information.
[0207] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0208] If the central monitoring process receives the heartbeat feedback information sent by any one of the sub-nodes, and when the heartbeat feedback message carries the fault monitoring result of the business process in the sub-node, send the fault information for the business process to the fault management node, where the fault information includes the fault monitoring result of the business process, the node identifier of the sub-node, and the process identifier of the business process.
[0209] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: Perform a status detection on the target business process through the local fault detection program deployed in the target business process to obtain a monitoring result;
[0210] If the monitoring result is a fault monitoring result, then carry the fault monitoring result in the target heartbeat feedback message and send it to the second sub-node.
[0211] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0212] Execute a fault detection task through the local fault monitoring program deployed in the target business process;
[0213] If the fault detection task fails, determine that the monitoring result is a fault monitoring result.
[0214] In one embodiment, a computer-readable storage medium is provided. In a possible implementation, this computer-readable storage medium is applied to a fault management node, and a computer program is stored thereon. When the computer program is executed by a processor, the following steps are implemented:
[0215] Receive the target fault information sent by the second sub-node, where the target fault information includes the node identifier of the first sub-node, the process identifier of the target business process with a fault in the running state in the first sub-node, and the fault monitoring result, where the fault monitoring result is carried in the target heartbeat feedback message sent by the first sub-node in response to the target heartbeat request sent by the second sub-node through the central monitoring process and monitoring the running state of the target business process;
[0216] Perform fault handling on the target service process according to the target fault information.
[0217] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0218] Perform restart processing on the target service process according to the process identifier included in the target fault information.
[0219] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0220] If updated fault information for the target service process is received, perform restart processing or isolation processing on the first child node according to the node identifier included in the updated fault information.
[0221] In one embodiment, a computer program product is provided. In a possible implementation manner, the computer program product is applied to a first child node, and the first child node is any one of the child nodes in a cloud storage system. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0222] For each target service process in the first child node, receive a target heartbeat request sent by a second child node in the cloud storage system through a central monitoring process through the target service process. The second child node is at least one child node other than the first child node;
[0223] In response to the target heartbeat request, monitor the running state of the target service process, and send a target heartbeat feedback message to the second child node according to the running state;
[0224] The target heartbeat feedback message is used for fault monitoring.
[0225] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0226] Send a heartbeat request to each service process in each child node through the central monitoring process in the first child node. The heartbeat request is used to instruct each child node to return heartbeat feedback information.
[0227] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0228] If heartbeat feedback information sent by any of the child nodes is received through the central monitoring process, and the heartbeat feedback message carries the fault monitoring result of the service process in the child node, fault information for the service process is sent to the fault management node, where the fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process.
[0229] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: The status of the target service process is detected through a local fault detection program deployed in the target service process to obtain a monitoring result;
[0230] If the monitoring result is a fault monitoring result, the fault monitoring result is carried in the target heartbeat feedback message and sent to the second child node.
[0231] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0232] A fault detection task is executed through a local fault monitoring program deployed in the target service process;
[0233] If the fault detection task fails, the monitoring result is determined to be a fault monitoring result.
[0234] In one embodiment, a computer program product is provided. In a possible implementation, the computer program product is applied to a fault management node and includes a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0235] Receive target fault information sent by a second child node, where the target fault information includes the node identifier of a first child node, the process identifier of a target service process with a fault in the running state in the first child node, and a fault monitoring result, where the fault monitoring result is carried in a target heartbeat feedback message sent by the first child node to the second child node in response to a target heartbeat request sent by the second child node through a central monitoring process and monitoring the running state of the target service process;
[0236] Perform fault handling on the target service process according to the target fault information.
[0237] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0238] Perform a restart process on the target service process according to the process identifier included in the target fault information.
[0239] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0240] If updated fault information for the target service process is received, the first sub-node is restarted or isolated according to the node identifier included in the updated fault information.
[0241] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0242] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0243] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0244] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A fault monitoring method, characterized in that, For a first child node, where the first child node is any one of the child nodes in the storage system, the method includes: For each target service process in the first child node, receive, through the target service process, a target heartbeat request sent by a second child node in the storage system through a central monitoring process, where the second child node is at least one child node other than the first child node; In response to the target heartbeat request, monitor the running status of the target service process, and send a target heartbeat feedback message to the second child node according to the running status, where the target heartbeat feedback message is used for fault monitoring.
2. The method according to claim 1, wherein The method further includes: Send a heartbeat request to each service process in each of the child nodes through the central monitoring process in the first child node, where the heartbeat request is used to instruct each of the child nodes to return heartbeat feedback information.
3. The method according to claim 2, wherein The method further includes: If receiving heartbeat feedback information sent by any one of the child nodes through the central monitoring process, and when the heartbeat feedback message carries a fault monitoring result of a service process in the child node, send fault information about the service process to the fault management node; Wherein, the fault information includes the fault monitoring result of the service process, the node identifier of the child node, and the process identifier of the service process.
4. The method according to claim 1, wherein The monitoring the running status of the target service process and sending a target heartbeat feedback message to the second child node according to the running status includes: Perform status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result; If the monitoring result is a fault monitoring result, carry the fault monitoring result in the target heartbeat feedback message and send it to the second child node.
5. The method according to claim 4, wherein The performing status detection on the target service process through a local fault detection program deployed in the target service process to obtain a monitoring result includes: Execute a fault detection task through a local fault monitoring program deployed in the target service process; When the fault detection task fails, determine that the monitoring result is a fault monitoring result.
6. A fault monitoring method, characterized in that, For a fault management node, the method includes: Receive target fault information sent by a second child node, where the target fault information includes the node identifier of a first child node, the process identifier of a target service process with a faulty running status in the first child node, and a fault monitoring result, where the fault monitoring result is carried in a target heartbeat feedback message sent by the first child node in response to a target heartbeat request sent by the second child node through a central monitoring process, monitoring the running status of the target service process, and sending it to the second child node according to the running status; Perform fault handling on the target service process according to the target fault information.
7. The method according to claim 6, wherein The performing fault handling on the target service process according to the target fault information includes: Perform a restart process on the target service process according to the process identifier included in the target fault information.
8. The method according to claim 6, characterized in that The method further includes: If updated fault information for the target service process is received, restart or isolate the first child node according to the node identifier included in the updated fault information.
9. A fault monitoring device, characterized in that, For a first child node, where the first child node is any one of the child nodes in a storage system, the apparatus includes: A receiving module, configured to, for each target service process in the first child node, receive, through the target service process, a target heartbeat request sent by a second child node in the storage system through a central monitoring process, where the second child node is at least one child node other than the first child node; A response module, configured to, in response to the target heartbeat request, monitor the running state of the target service process, and send a target heartbeat feedback message to the second child node according to the running state, where the target heartbeat feedback message is used for fault monitoring.
10. A fault monitoring device, characterized in that, For a fault management node, the apparatus includes: A receiving module, configured to receive target fault information sent by a second child node, where the target fault information includes the node identifier of a first child node, the process identifier of a target service process with a fault in the running state in the first child node, and a fault monitoring result, where the fault monitoring result is carried in a target heartbeat feedback message sent by the first child node to the second child node after monitoring the running state of the target service process in response to a target heartbeat request sent by the second child node through a central monitoring process; A fault handling module, configured to perform fault handling on the target service process according to the target fault information.
11. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.