Isolation method and apparatus for distributed nodes
By monitoring nodes to detect decision-making opportunities and generating isolation instructions, the isolation operations of distributed nodes are optimized, solving the problem of low efficiency in handling sub-healthy node states and improving the stability and reliability of the system.
Patent Information
- Application Number
- CN202410805678.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-06-20
AI Technical Summary
In existing technologies, the efficiency of handling sub-healthy states of nodes is low, which affects system performance and reliability, and the long-term presence of sub-healthy states affects the normal operation of the system.
By monitoring nodes to detect the decision-making timing of target distributed nodes, isolation instructions are generated and isolation operations are executed. By combining the operational information of multiple distributed nodes, isolation decisions are optimized to avoid the system impact caused by blind isolation.
It improves the efficiency of handling sub-healthy states of nodes, reduces the impact of sub-healthy states on system operation, and ensures system stability and reliability.
Smart Images

Figure CN118631545B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for isolating distributed nodes. Background Technology
[0002] A node's sub-health state typically refers to a situation in a network or distributed system where the performance or functionality of a node fails to meet expected standards. While a node in a sub-healthy state can continue to operate, this state can negatively impact the performance and reliability of the entire system. To prevent the spread of sub-healthy states and their impact on system operation, a node isolation mechanism has been designed. However, isolating sub-healthy nodes can disrupt normal system operation. Therefore, when dealing with sub-healthy nodes, a common approach is to issue alerts upwards without isolation. This approach allows the sub-healthy state of nodes to persist in the system for extended periods, affecting overall system operation, and the efficiency of handling node sub-health states is low. Summary of the Invention
[0003] This application provides a method and apparatus for isolating distributed nodes, which at least solves the problem of low processing efficiency for sub-healthy states of nodes in related technologies.
[0004] According to one embodiment of this application, a method for isolating distributed nodes is provided. The distributed system includes a monitoring node and multiple distributed nodes, the monitoring node being connected to the multiple distributed nodes. The method is applied to a target distributed node among the multiple distributed nodes, and the method includes:
[0005] If the current operating state of the target distributed node is detected to be sub-healthy, the decision timing corresponding to the target distributed node is detected according to the configuration information on the target distributed node, wherein the decision timing is used to indicate when the monitoring node decides whether the target distributed node should be isolated.
[0006] If the decision-making time has been detected, the process continues and an isolation request is reported to the monitoring node, wherein the isolation request is used to request the isolation of the target distributed node;
[0007] The monitoring node receives the isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node is used to generate the isolation instruction based on the running information of the multiple distributed nodes in response to the isolation request.
[0008] The target distributed node is isolated according to the isolation instructions.
[0009] As an optional implementation, detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node includes:
[0010] Extract a decision identifier from the configuration information, wherein the decision identifier is used to indicate the timing of the decision;
[0011] Monitor whether the decision timing indicated by the decision identifier has been reached on the target distributed node.
[0012] As an optional implementation, monitoring whether the decision timing indicated by the decision identifier has been reached on the target distributed node includes:
[0013] When the decision-making timing is the first decision-making timing, a recovery operation is performed, wherein the recovery operation is used to restore the operating state from the sub-healthy state to the healthy state, and the first decision-making timing is used to indicate that if the distributed node fails to recover from the sub-healthy state to the healthy state after being in the sub-healthy state, the monitoring node shall decide on the isolation of the distributed node; the state change information of the operating state is detected, wherein the state change information is used to indicate whether the operating state has changed from the sub-healthy state to the healthy state; if the state change information indicates that the operating state has not yet changed from the sub-healthy state to the healthy state, it is determined that the first decision-making timing has been detected;
[0014] If the decision timing is a second decision timing, it is determined that the second decision timing has been detected, wherein the second decision timing is used to indicate that the monitoring node decides to isolate the distributed node when the distributed node is in a sub-healthy state.
[0015] As an optional implementation, extracting the decision identifier from the configuration information includes:
[0016] The automatic recovery identifier is extracted from the configuration information, wherein the decision identifier includes the automatic recovery identifier, which is used to indicate the automatic recovery capability of the target distributed node from the sub-healthy state to the healthy state;
[0017] When the automatic recovery identifier is a first identifier value, the decision timing is determined to be the first decision timing, wherein the first identifier value is used to indicate that the target distributed node has the ability to automatically recover from the sub-healthy state to the healthy state;
[0018] When the automatic recovery identifier is the second identifier value, the decision timing is determined to be the second decision timing, wherein the second identifier value is used to indicate that the target distributed node does not have the automatic recovery capability from the sub-healthy state to the healthy state.
[0019] As an optional implementation, the step of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected to be in a sub-healthy state includes: detecting the current operating state of the target distributed node through a first process; and when the first process detects that the current operating state of the target distributed node is in the sub-healthy state, transferring the isolation state machine on the target distributed node from a normal state to a fault state through the first process, wherein the second process on the target distributed node is used to perform the operation of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the isolation state machine transfers from the normal state to the fault state.
[0020] The step of continuing to run and reporting an isolation request to the monitoring node when the decision-making time has been detected includes: when the second process detects that the decision-making time has been detected, controlling the target distributed node to continue running through the first process and transferring the isolation state machine from the fault state to the isolation state, wherein the second process is used to perform the operation of reporting an isolation request to the monitoring node when the isolation state machine transfers from the fault state to the isolation state;
[0021] Receiving the isolation instruction returned by the monitoring node in response to the isolation request includes: receiving the isolation instruction returned by the monitoring node in response to the isolation request through the second process after the operation of reporting the isolation request to the monitoring node;
[0022] The isolation of the target distributed node according to the isolation instruction includes: isolating the target distributed node by the first process according to the isolation instruction received by the second process.
[0023] As an optional implementation, isolating the target distributed node according to the isolation instruction includes:
[0024] Even if the isolation instruction indicates that isolation of the target distributed node is prohibited, the services on the target distributed node may continue to run.
[0025] When the isolation instruction indicates that the isolation of the target distributed node is permitted, a reference distributed node is extracted from the isolation instruction, wherein the reference distributed node is a distributed node selected by the monitoring node from the plurality of distributed nodes based on the operation information of the plurality of distributed nodes, which is used to take over the services on the target distributed node; the services on the target distributed node are transferred to the reference distributed node; and the target distributed node is isolated.
[0026] As an optional implementation, before detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node, the method further includes:
[0027] Detect the current operating parameters of the target distributed node;
[0028] If the operating parameters exceed the target parameter range, a repair identifier is extracted from the configuration information on the target distributed node, wherein the repair identifier is used to indicate the target distributed node's ability to repair the sub-healthy state;
[0029] When the repair identifier indicates that the target distributed node has the ability to repair the sub-healthy state, a repair operation is performed, wherein the repair operation is used to repair the sub-healthy state;
[0030] If the repair operation fails to repair the sub-healthy state, or if the repair identifier is used to indicate that the target distributed node does not have the ability to repair the sub-healthy state, then the current operating state of the target distributed node is determined to be sub-healthy.
[0031] According to another embodiment of this application, an isolation device for distributed nodes is provided. The distributed system includes a monitoring node and multiple distributed nodes, the monitoring node being connected to the multiple distributed nodes. The device is applied to a target distributed node among the multiple distributed nodes, and the device includes:
[0032] The first detection module is used to detect the decision timing corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected to be in a sub-healthy state. The decision timing is used to indicate the timing when the monitoring node decides whether to isolate the target distributed node.
[0033] The reporting module is used to continue running and report an isolation request to the monitoring node when it is detected that the decision-making time has arrived, wherein the isolation request is used to request the isolation of the target distributed node;
[0034] A receiving module is configured to receive an isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node is configured to generate the isolation instruction based on the operating information of the plurality of distributed nodes in response to the isolation request;
[0035] An isolation module is used to isolate the target distributed node according to the isolation instructions.
[0036] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0037] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0038] This application, upon detecting that the target distributed node is currently in a sub-healthy state, detects the corresponding decision-making opportunity based on the configuration information of the target distributed node. If the decision-making opportunity has arrived—that is, the time for the monitoring node to decide whether to isolate the target distributed node has arrived—the system continues operation and reports an isolation request to the monitoring node. It then receives the isolation instruction returned by the monitoring node in response to the isolation request and isolates the target distributed node according to the isolation instruction. In other words, it first detects whether the time for making an isolation decision has arrived; if so, it requests... The monitoring node generates isolation instructions based on the operational information of multiple distributed nodes, instructing the isolation of the target distributed node. This method, which combines the operational information of multiple distributed nodes to generate isolation instructions for the target distributed node, can fully measure the impact of node isolation measures on the overall system operation. It provides isolation instructions for the target distributed node adapted to the operational status of each node in the system, ensuring that nodes in a sub-healthy state can be isolated reasonably when isolation is required. This reduces the impact of sub-healthy nodes on the system operation, thus solving the problem of low processing efficiency for sub-healthy nodes and improving the efficiency of handling sub-healthy nodes. Attached Figure Description
[0039] Figure 1 This is a hardware structure outline of a server device for a distributed node isolation method according to an embodiment of this application. Figure 2
[0040] Figure 2This is a flowchart of a distributed node isolation method according to an embodiment of this application. Figure 2
[0041] Figure 3 This is the overall architecture of a method for handling node sub-health based on finite state machines and multidimensional data analysis, according to an embodiment of this application. Figure 2
[0042] Figure 4 This is a state transition method for handling node sub-health based on finite state machines and multidimensional data analysis, according to an embodiment of this application. Figure 2
[0043] Figure 5 This is a specific flow chart of a method for handling node sub-health based on finite state machines and multidimensional data analysis according to an embodiment of this application. Figure 2
[0044] Figure 6 This is a structural block diagram of an isolation device for distributed nodes according to an embodiment of this application. Detailed Implementation
[0045] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0046] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0047] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a server device for a distributed node isolation method according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown. A processor 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0048] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the distributed node isolation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0049] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0050] This embodiment provides a method for isolating distributed nodes. Figure 2 This is a flowchart of a distributed node isolation method according to an embodiment of this application, such as... Figure 2 As shown, the distributed system includes a monitoring node and multiple distributed nodes. The monitoring node is connected to the multiple distributed nodes. The method is applied to a target distributed node among the multiple distributed nodes, and the process includes the following steps:
[0051] Step S202: When it is detected that the current operating state of the target distributed node is in a sub-healthy state, the decision timing corresponding to the target distributed node is detected according to the configuration information on the target distributed node, wherein the decision timing is used to indicate the timing when the monitoring node decides whether the target distributed node should be isolated.
[0052] Step S204: If the decision-making time has been detected, continue running and report an isolation request to the monitoring node, wherein the isolation request is used to request isolation of the target distributed node;
[0053] Step S206: Receive the isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node is used to generate the isolation instruction based on the running information of the multiple distributed nodes in response to the isolation request;
[0054] Step S208: Isolate the target distributed node according to the isolation instruction.
[0055] Through the above steps, since the target distributed node is detected to be in a sub-healthy state, the decision-making timing corresponding to the target distributed node is detected based on the configuration information on the target distributed node. If the decision-making timing has been reached—that is, the timing for the monitoring node to decide whether to isolate the target distributed node has arrived—the process continues, and an isolation request is reported to the monitoring node. The isolation instruction returned by the monitoring node in response to the isolation request is received, and the target distributed node is isolated according to the isolation instruction. In other words, the process first detects whether the timing for making an isolation decision has arrived; if the timing for making an isolation decision has arrived, the process proceeds. The monitoring node generates isolation instructions based on the operational information of multiple distributed nodes to instruct the isolation of the target distributed node. This method, which combines the operational information of multiple distributed nodes to generate isolation instructions, can fully assess the impact of node isolation measures on the overall system operation. It provides isolation instructions for the target distributed node adapted to the operational status of each node within the system, ensuring that nodes in a sub-healthy state can be rationally isolated when isolation is required. This reduces the impact of sub-healthy nodes on system operation, thus solving the problem of low efficiency in handling sub-healthy nodes and improving the efficiency of handling such nodes.
[0056] Optionally, in this embodiment, the distributed system includes a monitoring node and multiple distributed nodes. The multiple distributed nodes include, but are not limited to, those used to process different parts of the same business, or those used to process different businesses. This includes, but is not limited to, requiring at least P nodes to be operational to ensure the operation of the distributed system, or isolating the business run by a node so that the business is run by a designated node. In the face of the above situations, it is necessary to comprehensively judge the isolation status of the target distributed node by combining the operation status of each node in the distributed system, so as to avoid causing more adverse effects on the operation of the distributed system.
[0057] Optionally, in this embodiment, the monitoring node includes, but is not limited to, receiving operational information reported by multiple distributed nodes in the distributed system at a set frequency.
[0058] In the embodiment provided in step S202, the sub-health state of a node refers to a node exhibiting some abnormal symptoms or behaviors in certain aspects, but not yet reaching the point of complete functional loss. This state may be due to the node suffering a certain degree of damage or stress, causing its functions to be affected in some aspects, but it can still operate normally. The sub-health state of a node may lead to performance degradation, prolonged response time, or even failure. Therefore, timely detection and handling of the sub-health state of nodes is crucial to ensure the normal operation and stability of the entire system.
[0059] Optionally, in this embodiment of the application, the configuration information on the target distributed node includes, but is not limited to, recording an automatic recovery identifier, a repair identifier, and an isolation identifier. The automatic recovery identifier is used to indicate the target distributed node's ability to automatically recover from a sub-healthy state to a healthy state. The repair identifier is used to indicate the target distributed node's ability to repair a sub-healthy state. The isolation identifier is used to indicate whether the target distributed node has the ability to be isolated.
[0060] Optionally, in this embodiment of the application, the decision-making timing corresponding to the target distributed node is detected based on the configuration information on the target distributed node, including but not limited to the determination that the isolation decision-making timing of the target distributed node has arrived when the current operating state of the target distributed node is detected to be in a sub-healthy state.
[0061] In the embodiment provided in step S204, the isolation request includes, but is not limited to, carrying configuration information on the target distributed node, and reporting the isolation request to the monitoring node includes, but is not limited to, reporting the isolation request carrying the configuration information on the target distributed node to the monitoring node.
[0062] Optionally, in this embodiment of the application, when it is detected that the decision-making time has arrived, continuing to run and reporting an isolation request to the monitoring node includes, but is not limited to, on the one hand, in order to avoid the impact on the overall operation of the distributed system caused by the target distributed node in a sub-healthy state blindly stopping operation, requiring the target distributed node to continue running without obtaining an isolation instruction, and on the other hand, reporting an isolation request to the monitoring node to handle the existing sub-healthy state of the node, further reducing the impact of the sub-healthy state of the node on the system operation.
[0063] In the embodiment provided in step S206, the isolation instruction includes, but is not limited to, instructions to isolate a portion of the services of the target distributed node, or to isolate a component of the target distributed node, or to isolate the node itself.
[0064] Optionally, in this embodiment, the operating information of multiple distributed nodes includes, but is not limited to, the operating status of multiple distributed nodes, wherein the operating status includes a healthy state, a sub-healthy state, and a fault state, and the fault state includes an isolation state.
[0065] Optionally, in this embodiment, the monitoring node responds to the isolation request by generating the isolation instruction based on the operating information of the plurality of distributed nodes, including but not limited to receiving the isolation request. The isolation request carries configuration information and service information of the target distributed node. The service information indicates the requirements of the service running on the target distributed node for the running node. An isolation identifier is extracted from the configuration information, indicating whether the target distributed node has the capability to be isolated. If the isolation identifier indicates that the target distributed node has the capability to be isolated, the service information is detected. If the service information indicates that the service can only run on the target distributed node, the monitoring node responds to the isolation request by generating the isolation instruction based on the operating information of the target distributed node. In the case of a target distributed node, a first isolation instruction is generated to indicate that isolation of the target distributed node is prohibited. If the service information indicates that the service can run on distributed nodes other than the target distributed node, the operation information of multiple distributed nodes is detected. This operation information indicates the operating status of the multiple distributed nodes. If there are Q non-faulty nodes and at least one node in a healthy state among the multiple distributed nodes, a second isolation instruction is generated to indicate that isolation of the target distributed node is allowed. Here, Q is an integer greater than 2, the non-faulty states include the healthy state and the sub-healthy state, and the isolation instruction includes the first isolation instruction and the second isolation instruction. Through these steps, the potential impact of isolating the target distributed node on system operation is fully considered, and an isolation instruction that conforms to the actual operation of a distributed system is provided to indicate the isolation of the target distributed node.
[0066] In the embodiment provided in step S208, isolating the target distributed node according to the isolation instruction includes, but is not limited to, isolating the target distributed node immediately after receiving the isolation instruction, or, upon receiving the isolation instruction, selecting a reference distributed node from among the other distributed nodes in the distributed system besides the target distributed node to receive the services originally processed by the target distributed node, and isolating the target distributed node after the service handover is completed.
[0067] Optionally, in this embodiment, the method includes, but is not limited to, issuing an alarm message after isolating the target distributed node, prompting staff to analyze and process the isolated node so that the node's state is restored to a healthy state.
[0068] As an optional implementation, detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node includes:
[0069] Extract a decision identifier from the configuration information, wherein the decision identifier is used to indicate the timing of the decision;
[0070] Monitor whether the decision timing indicated by the decision identifier has been reached on the target distributed node.
[0071] Optionally, in this embodiment, the decision identifier includes, but is not limited to, the automatic recovery identifier.
[0072] As an optional implementation, monitoring whether the decision timing indicated by the decision identifier has been reached on the target distributed node includes:
[0073] When the decision-making timing is the first decision-making timing, a recovery operation is performed, wherein the recovery operation is used to restore the operating state from the sub-healthy state to the healthy state, and the first decision-making timing is used to indicate that if the distributed node fails to recover from the sub-healthy state to the healthy state after being in the sub-healthy state, the monitoring node shall decide on the isolation of the distributed node; the state change information of the operating state is detected, wherein the state change information is used to indicate whether the operating state has changed from the sub-healthy state to the healthy state; if the state change information indicates that the operating state has not yet changed from the sub-healthy state to the healthy state, it is determined that the first decision-making timing has been detected;
[0074] If the decision timing is a second decision timing, it is determined that the second decision timing has been detected, wherein the second decision timing is used to indicate that the monitoring node decides to isolate the distributed node when the distributed node is in a sub-healthy state.
[0075] Optionally, in the embodiments of this application, there are including but not limited to the first decision timing corresponding to the presence of an automatic recovery flag, and the second decision timing corresponding to the absence of an automatic recovery flag. The first decision timing represents that the target distributed node can perform a recovery operation, and isolation is decided only if the recovery fails. The second decision timing represents that the target distributed node cannot perform a recovery operation, and isolation is decided directly after the sub-healthy state is determined.
[0076] By taking the above steps, a recovery operation is attempted before implementing an isolation decision, minimizing the use of isolation in the distributed system. This reduces the need for manual intervention after isolation and improves the efficiency of handling sub-healthy states.
[0077] As an optional implementation, extracting the decision identifier from the configuration information includes:
[0078] The automatic recovery identifier is extracted from the configuration information, wherein the decision identifier includes the automatic recovery identifier, which is used to indicate the automatic recovery capability of the target distributed node from the sub-healthy state to the healthy state;
[0079] When the automatic recovery identifier is a first identifier value, the decision timing is determined to be the first decision timing, wherein the first identifier value is used to indicate that the target distributed node has the ability to automatically recover from the sub-healthy state to the healthy state;
[0080] When the automatic recovery identifier is the second identifier value, the decision timing is determined to be the second decision timing, wherein the second identifier value is used to indicate that the target distributed node does not have the automatic recovery capability from the sub-healthy state to the healthy state.
[0081] Optionally, in the embodiments of this application, the first identifier value is determined to be 1 and the second identifier value is determined to be 0.
[0082] By using the above steps, the value of the automatic recovery flag indicates whether the node has the ability to automatically recover from a sub-healthy state to a healthy state, which matches the timing of isolation decisions. This ensures that nodes that can automatically recover do so as much as possible, thus improving the efficiency of handling sub-healthy states of nodes.
[0083] As an optional implementation, the step of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected to be in a sub-healthy state includes: detecting the current operating state of the target distributed node through a first process; and when the first process detects that the current operating state of the target distributed node is in the sub-healthy state, transferring the isolation state machine on the target distributed node from a normal state to a fault state through the first process, wherein the second process on the target distributed node is used to perform the operation of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the isolation state machine transfers from the normal state to the fault state.
[0084] The step of continuing to run and reporting an isolation request to the monitoring node when the decision-making time has been detected includes: when the second process detects that the decision-making time has been detected, controlling the target distributed node to continue running through the first process and transferring the isolation state machine from the fault state to the isolation state, wherein the second process is used to perform the operation of reporting an isolation request to the monitoring node when the isolation state machine transfers from the fault state to the isolation state;
[0085] Receiving the isolation instruction returned by the monitoring node in response to the isolation request includes: receiving the isolation instruction returned by the monitoring node in response to the isolation request through the second process after the operation of reporting the isolation request to the monitoring node;
[0086] The isolation of the target distributed node according to the isolation instruction includes: isolating the target distributed node by the first process according to the isolation instruction received by the second process.
[0087] Through the above steps, by having the first process perform decision-making timing detection and state switching, and the second process perform isolation instruction generation, the process of handling the sub-healthy state of a node is divided into a two-process cooperation process. This allows the handling of the sub-healthy state of multiple distributed nodes to be carried out simultaneously, improving the efficiency of handling the sub-healthy state of nodes.
[0088] As an optional implementation, isolating the target distributed node according to the isolation instruction includes:
[0089] Even if the isolation instruction indicates that isolation of the target distributed node is prohibited, the services on the target distributed node may continue to run.
[0090] When the isolation instruction indicates that the isolation of the target distributed node is permitted, a reference distributed node is extracted from the isolation instruction, wherein the reference distributed node is a distributed node selected by the monitoring node from the plurality of distributed nodes based on the operation information of the plurality of distributed nodes, which is used to take over the services on the target distributed node; the services on the target distributed node are transferred to the reference distributed node; and the target distributed node is isolated.
[0091] Optionally, in this embodiment, extracting a reference distributed node from the isolation instruction includes, but is not limited to, selecting the distributed node with the highest operating efficiency from other distributed nodes in the distributed system besides the target distributed node as the reference distributed node.
[0092] By following the steps above, if it is confirmed that the target distributed node needs to be isolated, a reference distributed node is selected to take over the business on the target distributed node. After the business transfer is completed, the target distributed node is then isolated, which can avoid the impact of the isolation of the target distributed node on the processing and operation of the distributed system business.
[0093] As an optional implementation, before detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node, the method further includes:
[0094] Detect the current operating parameters of the target distributed node;
[0095] If the operating parameters exceed the target parameter range, a repair identifier is extracted from the configuration information on the target distributed node, wherein the repair identifier is used to indicate the target distributed node's ability to repair the sub-healthy state;
[0096] When the repair identifier indicates that the target distributed node has the ability to repair the sub-healthy state, a repair operation is performed, wherein the repair operation is used to repair the sub-healthy state;
[0097] If the repair operation fails to repair the sub-healthy state, or if the repair identifier is used to indicate that the target distributed node does not have the ability to repair the sub-healthy state, then the current operating state of the target distributed node is determined to be sub-healthy.
[0098] Optionally, in this embodiment, detecting whether the current operating state of the target distributed node is in a sub-healthy state includes, but is not limited to, collecting hardware information, system resource information, and network information of the target distributed node. The hardware information indicates the operating status of the hardware on the target distributed node, the system resource information indicates the operating status of the system resources on the target distributed node, and the network information indicates the network operating status of the target distributed node. Hardware detection, system resource detection, and network status detection are performed based on the collected hardware information, system resource information, and network information, respectively. If the hardware detection fails, the system resource detection fails, and / or the network status detection fails, the current operating state of the target distributed node is determined to be in the sub-healthy state.
[0099] Optionally, in this embodiment of the application, detecting the current operating parameters of the target distributed node includes, but is not limited to, detecting multiple current operating parameters of the target distributed node. If any of the multiple operating parameters exceeds the corresponding target parameter range, it is considered that the current operating state of the target distributed node needs further analysis and processing.
[0100] By following the above steps, the operating status of the target distributed node is detected. If the operating status of the target distributed node does not meet the requirements, an attempt is made to repair the target distributed node. Only if the repair fails or cannot be repaired is the node confirmed to have entered a sub-healthy state. This is equivalent to trying to change the sub-healthy state as soon as the sub-healthy state is discovered, which reduces the duration of the sub-healthy state of the nodes in the system and improves the efficiency of handling sub-healthy states.
[0101] As an optional implementation, this application proposes a method and apparatus for handling node sub-health based on finite state machines and multidimensional data analysis. This apparatus sets five node states: normal, automatic recovery, fault, automatic repair, and isolation. By collecting various information from each node (including resource information, network information, hardware information, etc.), it first performs single-node analysis based on this information. If a problem is detected, it triggers a node state transition, thereby triggering different processing logics. When isolation of a component of a node or the node itself is required, the Monitor (MON, i.e., the monitoring node) (because the MON is at the top level in the storage system and has information from all nodes, it can make unified decisions) determines whether isolation can be performed. Specifically:
[0102] Figure 3 This is an overall architecture diagram of a method for handling node sub-health based on finite state machines and multidimensional data analysis, according to an embodiment of this application. Figure 3As shown, the program is mainly divided into three modules: the SUBHEALTH Client module (hereinafter simplified to Client), the SUBHEALTH Server module (hereinafter simplified to Server), and the synchronization command module. The Client module collects information on hardware, system resources, and network status; the Server module processes and analyzes this information, and ultimately issues alarms and takes actions. After collecting information on hardware, system resources, and network status, the Client sends specified information packets to the SUBHEALTH Server on its local node, based on different detection time requirements. Upon receiving the information packets, the Server first parses the packets and then enters different detection processing logic based on the packet content, performing hardware detection, system resource detection, and network status detection respectively. Alarms requiring decision-making are reported to the MON service for decision-making and a response is received; alarms not requiring decision-making are directly reported to the management software. The collected component information is also reported to the management software. The synchronization command module is mainly used for automatic and manual synchronization of alarm thresholds and alarm switches. The Client and Server are on the same node (i.e., the target distributed node).
[0103] Figure 4 This is a state transition diagram of a method for handling node sub-health based on finite state machines and multidimensional data analysis, according to an embodiment of this application. Figure 4 As shown, the node sub-health detection program is divided into five states. Since some states are optional, the program will determine whether to enter an optional state based on the configuration.
[0104] 1) Normal state: Required state, indicating that the current node is normal;
[0105] a) If there is a repairable indicator, then when the fault conditions are met, the state changes from normal to repairable.
[0106] b) If there is no repair indicator, then when the fault conditions are met, the normal state becomes the fault state.
[0107] 2) Repair Status: Optional, indicating that the current node is attempting automatic repair;
[0108] a) If the repair is successful, the repair status will change to normal status.
[0109] b) If the repair fails, the repair status will be changed to the fault status.
[0110] 3) Fault Status: Required, indicates that the current node is faulty;
[0111] a) If there is an isolable identifier, then when the isolation conditions are met, the fault state becomes the isolation state.
[0112] b) If there is no isolation marker but there is an automatic recovery marker, then when the conditions are met (usually time requirements), the fault state -> automatic recovery state.
[0113] c) If there is no isolation marker or automatic recovery marker, then after manual intervention and recovery, if the recovery conditions are met, the fault status will change to normal status.
[0114] 4) Isolation status: Optional, indicating that the current node is in storage service isolation status;
[0115] a) If there is an automatic recovery flag, then when the conditions are met (usually time), the isolation state -> automatic recovery state.
[0116] b) If there is no automatic recovery indicator, manual intervention is required to restore the status. If the recovery conditions are met, the status will change from isolation to normal.
[0117] 5) Automatic recovery status: Optional, indicating that the current node is attempting automatic recovery;
[0118] a) If automatic recovery is successful, the automatic recovery status will change to normal.
[0119] b) If automatic recovery fails, the automatic recovery status will be changed to fault status / isolation status.
[0120] Figure 5 This is a flowchart illustrating a method for handling node sub-health based on finite state machines and multidimensional data analysis, according to an embodiment of this application. Figure 5 As shown, the Client primarily uses configuration to call the corresponding client.script script based on different call cycles. The script collects data and sends the raw data to the Server. The Server primarily uses configuration to call the corresponding server.script script to analyze the raw data. If a threshold is exceeded, an alarm is triggered and sent to the management software; if the threshold is not exceeded, an alarm recovery is triggered and sent to the management software. If isolation decisions are required, the Server interacts with the MON (Monitoring Node). The MON considers various status information from each node to make a decision. If the MON allows isolation, isolation is performed; if the MON does not allow isolation, only an alarm is reported, awaiting processing by operations personnel.
[0121] By following the steps above, and employing finite state machines and multidimensional data analysis of each node, it is possible to efficiently determine whether a node can perform isolation actions, and how to do so. This improves system reliability and stability while reducing labor costs.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0123] This embodiment also provides an isolation device for distributed nodes, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0124] Figure 6 This is a structural block diagram of a distributed node isolation device according to an embodiment of this application, such as... Figure 6 As shown, the distributed system includes a monitoring node and multiple distributed nodes, the monitoring node being connected to the multiple distributed nodes, and the device being applied to a target distributed node among the multiple distributed nodes. The device includes:
[0125] The first detection module 602 is used to detect the decision timing corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected to be in a sub-healthy state. The decision timing is used to indicate the timing when the monitoring node decides whether to isolate the target distributed node.
[0126] The reporting module 604 is used to continue running and report an isolation request to the monitoring node when it is detected that the decision-making time has arrived, wherein the isolation request is used to request the isolation of the target distributed node;
[0127] The receiving module 606 is used to receive the isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node is used to generate the isolation instruction based on the running information of the multiple distributed nodes in response to the isolation request;
[0128] The isolation module 608 is used to isolate the target distributed node according to the isolation instruction.
[0129] Through the above steps, since the target distributed node is detected to be in a sub-healthy state, the decision-making timing corresponding to the target distributed node is detected based on the configuration information on the target distributed node. If the decision-making timing has been reached—that is, the timing for the monitoring node to decide whether to isolate the target distributed node has arrived—the process continues, and an isolation request is reported to the monitoring node. The isolation instruction returned by the monitoring node in response to the isolation request is received, and the target distributed node is isolated according to the isolation instruction. In other words, the process first detects whether the timing for an isolation decision has arrived; if the timing for an isolation decision has arrived, the process proceeds... The monitoring node generates isolation instructions based on the operational information of multiple distributed nodes to instruct the isolation of the target distributed node. This method, which combines the operational information of multiple distributed nodes to generate isolation instructions, can fully assess the impact of node isolation measures on the overall system operation. It provides isolation instructions for the target distributed node adapted to the operational status of each node within the system, ensuring that nodes in a sub-healthy state can be rationally isolated when isolation is required. This reduces the impact of sub-healthy nodes on system operation, thus solving the problem of low efficiency in handling sub-healthy nodes and improving the efficiency of handling such nodes.
[0130] As an optional implementation, the first detection module includes:
[0131] The first extraction unit is used to extract a decision identifier from the configuration information, wherein the decision identifier is used to indicate the decision timing;
[0132] A monitoring unit is used to monitor whether the decision timing indicated by the decision identifier has been reached on the target distributed node.
[0133] As an optional implementation, the monitoring unit is further configured to:
[0134] When the decision-making timing is the first decision-making timing, a recovery operation is performed, wherein the recovery operation is used to restore the operating state from the sub-healthy state to the healthy state, and the first decision-making timing is used to indicate that if the distributed node fails to recover from the sub-healthy state to the healthy state after being in the sub-healthy state, the monitoring node shall decide on the isolation of the distributed node; the state change information of the operating state is detected, wherein the state change information is used to indicate whether the operating state has changed from the sub-healthy state to the healthy state; if the state change information indicates that the operating state has not yet changed from the sub-healthy state to the healthy state, it is determined that the first decision-making timing has been detected;
[0135] If the decision timing is a second decision timing, it is determined that the second decision timing has been detected, wherein the second decision timing is used to indicate that the monitoring node decides to isolate the distributed node when the distributed node is in a sub-healthy state.
[0136] As an optional implementation, the first extraction unit is further configured to:
[0137] The automatic recovery identifier is extracted from the configuration information, wherein the decision identifier includes the automatic recovery identifier, which is used to indicate the automatic recovery capability of the target distributed node from the sub-healthy state to the healthy state;
[0138] When the automatic recovery identifier is a first identifier value, the decision timing is determined to be the first decision timing, wherein the first identifier value is used to indicate that the target distributed node has the ability to automatically recover from the sub-healthy state to the healthy state;
[0139] When the automatic recovery identifier is the second identifier value, the decision timing is determined to be the second decision timing, wherein the second identifier value is used to indicate that the target distributed node does not have the automatic recovery capability from the sub-healthy state to the healthy state.
[0140] As an optional implementation, the first detection module further includes: a detection unit, configured to detect the current operating state of the target distributed node through a first process; and a first transfer unit, configured to transfer the isolation state machine on the target distributed node from a normal state to a fault state through the first process when the first process detects that the current operating state of the target distributed node is the sub-healthy state, wherein the second process on the target distributed node is configured to perform an operation to detect the decision timing corresponding to the target distributed node based on the configuration information on the target distributed node when the isolation state machine transfers from the normal state to the fault state.
[0141] As an optional implementation, the reporting module includes: a second transfer unit, configured to, when the second process detects that the decision timing has arrived, control the target distributed node to continue running through the first process and transfer the isolation state machine from a fault state to an isolation state, wherein the second process is configured to, when the isolation state machine transfers from the fault state to the isolation state, perform the operation of reporting an isolation request to the monitoring node;
[0142] As an optional implementation, the receiving module includes: a receiving unit, configured to receive an isolation instruction returned by the monitoring node in response to the isolation request after the second process reports the isolation request to the monitoring node;
[0143] As an optional implementation, the isolation module includes: a first isolation unit, used to isolate the target distributed node by the first process according to the isolation instruction received by the second process.
[0144] As an optional implementation, the isolation module further includes:
[0145] The operating unit is configured to continue running the services on the target distributed node even when the isolation instruction indicates that isolation of the target distributed node is prohibited;
[0146] The second extraction unit is used to extract a reference distributed node from the isolation instruction when the isolation instruction indicates that the isolation of the target distributed node is permitted. The reference distributed node is a distributed node selected by the monitoring node from the plurality of distributed nodes based on the operation information of the plurality of distributed nodes, which is used to take over the business on the target distributed node. The third transfer unit is used to transfer the business on the target distributed node to the reference distributed node. The second isolation unit is used to isolate the target distributed node.
[0147] As an optional implementation, the isolation device further includes:
[0148] The second detection module is used to detect the current operating parameters of the target distributed node;
[0149] An extraction module is used to extract a repair identifier from the configuration information on the target distributed node when the running parameters exceed the target parameter range, wherein the repair identifier is used to indicate the target distributed node's ability to repair the sub-healthy state;
[0150] An execution module is configured to perform a repair operation when the repair identifier indicates that the target distributed node has the ability to repair the sub-healthy state, wherein the repair operation is used to repair the sub-healthy state;
[0151] The determination module is used to determine that the current operating state of the target distributed node is in a sub-healthy state when the repair operation fails to repair the sub-healthy state, or when the repair identifier is used to indicate that the target distributed node does not have the ability to repair the sub-healthy state.
[0152] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0153] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0154] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0155] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0156] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0157] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0158] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0159] Embodiments of this application also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0160] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0161] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0162] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for isolating distributed nodes, characterized in that, The distributed system includes a monitoring node and multiple distributed nodes, wherein the monitoring node is connected to the multiple distributed nodes, and the method is applied to a target distributed node among the multiple distributed nodes. The method includes: If the current operating state of the target distributed node is detected to be sub-healthy, the decision timing corresponding to the target distributed node is detected according to the configuration information on the target distributed node, wherein the decision timing is used to indicate when the monitoring node decides whether to isolate the target distributed node; If the decision-making time has been detected, the process continues and an isolation request is reported to the monitoring node, wherein the isolation request is used to request the isolation of the target distributed node; The monitoring node receives the isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node generates the isolation instruction based on the running information of the multiple distributed nodes in response to the isolation request. The target distributed node is isolated according to the isolation instructions.
2. The method according to claim 1, characterized in that, The step of detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node includes: Extract a decision identifier from the configuration information, wherein the decision identifier is used to indicate the timing of the decision; Monitor whether the decision timing indicated by the decision identifier has been reached on the target distributed node.
3. The method according to claim 2, characterized in that, The monitoring of whether the decision timing indicated by the decision identifier has been reached on the target distributed node includes: When the decision-making timing is the first decision-making timing, a recovery operation is performed, wherein the recovery operation is used to restore the operating state from the sub-healthy state to the healthy state, and the first decision-making timing is used to indicate that if the distributed node fails to recover from the sub-healthy state to the healthy state after being in the sub-healthy state, the monitoring node shall decide on the isolation of the distributed node; the state change information of the operating state is detected, wherein the state change information is used to indicate whether the operating state has changed from the sub-healthy state to the healthy state; if the state change information indicates that the operating state has not yet changed from the sub-healthy state to the healthy state, it is determined that the first decision-making timing has been detected; If the decision timing is a second decision timing, it is determined that the second decision timing has been detected, wherein the second decision timing is used to indicate that the monitoring node decides to isolate the distributed node when the distributed node is in a sub-healthy state.
4. The method according to claim 3, characterized in that, Extracting the decision identifier from the configuration information includes: The automatic recovery identifier is extracted from the configuration information, wherein the decision identifier includes the automatic recovery identifier, which is used to indicate the automatic recovery capability of the target distributed node from the sub-healthy state to the healthy state; When the automatic recovery identifier is a first identifier value, the decision timing is determined to be the first decision timing, wherein the first identifier value is used to indicate that the target distributed node has the ability to automatically recover from the sub-healthy state to the healthy state; When the automatic recovery identifier is the second identifier value, the decision timing is determined to be the second decision timing, wherein the second identifier value is used to indicate that the target distributed node does not have the automatic recovery capability from the sub-healthy state to the healthy state.
5. The method according to claim 1, characterized in that, The step of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected as sub-healthy includes: detecting the current operating state of the target distributed node through a first process; when the first process detects that the current operating state of the target distributed node is sub-healthy, transferring the isolation state machine on the target distributed node from a normal state to a fault state through the first process, wherein the second process on the target distributed node is used to perform the operation of detecting the decision-making opportunity corresponding to the target distributed node based on the configuration information on the target distributed node when the isolation state machine transfers from the normal state to the fault state; The step of continuing to run and reporting an isolation request to the monitoring node when the decision-making time has been detected includes: when the second process detects that the decision-making time has been detected, controlling the target distributed node to continue running through the first process and transferring the isolation state machine from the fault state to the isolation state, wherein the second process is used to perform the operation of reporting an isolation request to the monitoring node when the isolation state machine transfers from the fault state to the isolation state; Receiving the isolation instruction returned by the monitoring node in response to the isolation request includes: receiving the isolation instruction returned by the monitoring node in response to the isolation request through the second process after the operation of reporting the isolation request to the monitoring node; The isolation of the target distributed node according to the isolation instruction includes: isolating the target distributed node by the first process according to the isolation instruction received by the second process.
6. The method according to claim 1, characterized in that, The isolation of the target distributed node according to the isolation instruction includes: Even if the isolation instruction indicates that isolation of the target distributed node is prohibited, the services on the target distributed node may continue to run. When the isolation instruction indicates that the isolation of the target distributed node is permitted, a reference distributed node is extracted from the isolation instruction, wherein the reference distributed node is a distributed node selected by the monitoring node from the plurality of distributed nodes based on the operation information of the plurality of distributed nodes, which is used to take over the services on the target distributed node; the services on the target distributed node are transferred to the reference distributed node; and the target distributed node is isolated.
7. The method according to claim 1, characterized in that, Before detecting the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node, the method further includes: Detect the current operating parameters of the target distributed node; If the operating parameters exceed the target parameter range, a repair identifier is extracted from the configuration information on the target distributed node, wherein the repair identifier is used to indicate the target distributed node's ability to repair the sub-healthy state; When the repair identifier indicates that the target distributed node has the ability to repair the sub-healthy state, a repair operation is performed, wherein the repair operation is used to repair the sub-healthy state; If the repair operation fails to repair the sub-healthy state, or if the repair identifier is used to indicate that the target distributed node does not have the ability to repair the sub-healthy state, then the current operating state of the target distributed node is determined to be sub-healthy.
8. An isolation device for distributed nodes, characterized in that, The distributed system includes a monitoring node and multiple distributed nodes, the monitoring node being connected to the multiple distributed nodes, and the device being applied to a target distributed node among the multiple distributed nodes. The device includes: The first detection module is used to detect the decision-making timing corresponding to the target distributed node based on the configuration information on the target distributed node when the current operating state of the target distributed node is detected to be sub-healthy. The decision-making timing is used to indicate when the monitoring node decides whether to isolate the target distributed node. The reporting module is used to continue running and report an isolation request to the monitoring node when the decision-making timing is detected to have arrived. The isolation request is used to request the isolation of the target distributed node. A receiving module is configured to receive an isolation instruction returned by the monitoring node in response to the isolation request, wherein the monitoring node is configured to generate the isolation instruction based on the operating information of the plurality of distributed nodes in response to the isolation request; and an isolation module is configured to isolate the target distributed node according to the isolation instruction.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Distributed storage environment network sub-health detection and fault automatic processing method
CN115499294A
Node device control method, node device, Bluetooth Mesh system and storage medium
CN115665754A