Cluster management method and device, storage medium, and electronic device

By setting up primary and standby management nodes in the cluster, rapid switching and service recovery are achieved when management node abnormalities are achieved, the problem of cluster service interruption is solved, and the reliability and stability of the cluster is improved.

CN119322706BActive Publication Date: 2025-07-04ZHEJIANG DAHUA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411870229.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-04
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

When existing cluster services are abnormal in management nodes, it is difficult to provide continuous business services, resulting in unstable services and the management nodes cannot be re-elected normally.

Method used

The primary management node and the standby management node are set up in the first working group. When the primary management node is abnormal, the standby management node takes over its responsibilities and coordinates the nodes in the target cluster to provide services. Through periodic communication and state synchronization between the primary management node and the standby management node, the service continuity is ensured.

Benefits of technology

It simplifies the management node switching process, reduces the time and complexity of resource reacquisition, improves the reliability and continuity of cluster services, avoids service interruptions and network splits, and ensures the stable operation of the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119322706B_ABST
    Figure CN119322706B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for managing a cluster, a storage medium, and an electronic device. Among them, the method includes: setting a primary management node in a first working group, where the primary management node is set to coordinate nodes in a target cluster to provide services, the target cluster includes the first working group and the second working group, and the nodes in the first working group are set as standby nodes for the nodes in the second working group; setting a standby management node in the first working group, where the standby management node represents a standby management node of the primary management node; in the case where the primary management node fails, controlling the standby management node to replace the primary management node and coordinate the nodes in the target cluster to provide services. The present application solves the technical problem that it is difficult to provide normal services during the process of re-electing a management node when the management node of the cluster fails.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more particularly, to a method and apparatus for managing a cluster, a storage medium, and an electronic device. Background Art

[0002] In existing cluster services, when the management node responsible for scheduling and controlling the management service fails, the entire cluster service will be interrupted. At this time, even if a re-election is carried out, the service will still be interrupted, and continuous business services cannot be provided to users. Specifically, during the process of the cluster re-negotiating a new management node, since it is necessary to reorganize each node and obtain the business resources corresponding to the nodes, the service is unstable.

[0003] In summary, in the related art, there is a technical problem that it is difficult to provide normal services during the process of re-electing a management node when the management node of the cluster fails.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a method and apparatus for managing a cluster, a storage medium, and an electronic device, so as to at least solve the technical problem of difficulty in providing normal services during the process of re-electing a management node when the management node of the cluster fails.

[0006] According to one aspect of the embodiments of this application, a method for managing a cluster is provided, including: setting a primary management node in a first working group, where the primary management node is configured to coordinate nodes in a target cluster to provide services, the target cluster includes the first working group and a second working group, and the nodes in the first working group are configured as standby nodes for the nodes in the second working group; setting a standby management node in the first working group, where the standby management node represents a standby management node of the primary management node; and in the case where the primary management node fails, controlling the standby management node to replace the primary management node and coordinate the nodes in the target cluster to provide services.

[0007] According to another aspect of the embodiments of the present application, a management device for a cluster is further provided, including: a primary setting module, configured to set a primary management node in a first working group, where the primary management node is configured to coordinate nodes in a target cluster to provide services, the target cluster includes the first working group and a second working group, and the nodes in the first working group are set as standby nodes for the nodes in the second working group; a standby setting module, configured to set a standby management node in the first working group, where the standby management node represents a standby management node of the primary management node; a control module, configured to control the standby management node to replace the primary management node and coordinate the nodes in the target cluster to provide services when the primary management node fails.

[0008] Optionally, the device is further configured to: establish communication between the primary management node and the nodes in the target cluster, where the primary management node is configured to periodically obtain node service information sent by the nodes in the target cluster, the node service information includes the service status and service configuration of the nodes, and the node service information corresponds to the services provided by the nodes; control the primary management node to periodically send the node service information to the standby management node; when the primary management node fails, control the standby management node to establish communication with the nodes in the target cluster to replace the primary management node to periodically obtain the node service information, and coordinate the nodes in the target cluster to provide services based on the node service information.

[0009] Optionally, the device is further configured to: control the primary management node to periodically send its own working status to the standby management node; when the sending of the working status fails, detect the connection status between the primary management node and the standby management node; when the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node, and when the connection status indicates that both the primary management node and the standby management node are normal, control the primary management node to re-send the working status to the standby management node; obtain the reporting time corresponding to the standby management node receiving the working status; when the reporting time exceeds a preset time threshold, detect the connection status and the service resource information stored by the standby management node; when the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the node service information corresponding to each node in the target cluster, send a node notification message to each node and update the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0010] Optionally, the device is used to set the primary management node in the first working group in the following manner, including: initiating a cluster configuration request to divide the nodes in the target cluster into the first working group and the second working group; determining the primary management node that meets the target node conditions from the first working group, where the target node conditions include being able to communicate normally with any node in the target cluster and being recognized by the nodes in the first working group; and configuring the primary management node to manage the nodes in the target cluster.

[0011] Optionally, the device is used to configure the primary management node in the following manner to manage the nodes in the target cluster: establishing communication between the primary management node and the nodes in the target cluster, where the primary management node is used to periodically obtain the service status and service configuration sent by the nodes in the target cluster; when the primary management node receives the first service status sent by the first service node, sending the first service status to the standby management node and updating the working status of the target cluster, where the first service node represents the node with a changed service status; and when the primary management node receives the second service configuration sent by the second service node, sending the second service configuration to the standby management node and updating the working status of the target cluster, where the second service node represents the node with a changed service configuration.

[0012] Optionally, the device is further used to: when the first working group includes multiple nodes, obtain the other nodes in the first working group except the primary management node; and determine the node with the smallest load among the other nodes in the idle state as the standby management node.

[0013] Optionally, the device is further used to: when the first working group only includes the primary management node, obtain the other nodes in the second working group; and determine the node with the smallest load among the other nodes in the idle state as the standby management node.

[0014] Optionally, the device is further used to: control the primary management node to periodically send its own working status to the standby management node; when the sending of the working status fails, detect the connection status between the primary management node and the standby management node; and when the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node from the first working group according to the primary management node.

[0015] Optionally, the device is used to set a standby management node in the first working group in the following manner: obtain the reporting time corresponding to the working status sent by the primary management node received by the standby management node; in the case where the reporting time exceeds a preset time threshold, detect the connection status between the standby management node and the primary management node, and the service resource information stored by the standby management node; in the case where the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the service status and service configuration of each node in the target cluster, send a node notification message to each node, and update the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0016] Optionally, the device is further used to: after setting the standby management node in the first working group, in the case where the primary management node determines that the primary service node has an abnormality, control the primary management node to initiate a primary-standby switch operation to instruct the standby service node to replace the primary service node, and send the service configuration of the primary service node to the standby service node, where the primary service node is a node in the second working group, and the standby service node is a node in the first working group; in the case where the primary service node has repaired the abnormality, control the primary management node to initiate a primary-standby recovery operation to instruct the primary service node to replace the standby service node, and send the service configuration of the standby service node to the primary service node.

[0017] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is set to execute the above-mentioned cluster management method when running.

[0018] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the cluster management method as described above.

[0019] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is set to execute the above-mentioned cluster management method through the computer program.

[0020] In the embodiments of the present application, a primary management node is set in the first working group, where the primary management node is configured to coordinate the nodes in the target cluster to provide services. The target cluster includes the first working group and the second working group, and the nodes in the first working group are set as standby nodes for the nodes in the second working group; a standby management node is set in the first working group, where the standby management node represents the standby management node of the primary management node; in the case where the primary management node fails, the standby management node is controlled to take over the primary management node and coordinate the nodes in the target cluster to provide services. This method simplifies the management node switching process, reduces the time and complexity of resource re-acquisition, achieves the purpose of effectively improving the reliability and continuity of cluster services, and avoids problems such as service interruption and network splitting during the process of re-electing the management node. Thus, the cluster can provide stable services, and therefore, effectively solves the technical problem of difficulty in providing normal services during the process of re-electing the management node when the management node of the cluster fails. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0022] Figure 1 is a schematic diagram of the application environment of an optional cluster management method according to an embodiment of the present application;

[0023] Figure 2 is a schematic flowchart of an optional cluster management method according to an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of a cluster of an optional cluster management method according to an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of an optional cluster management method according to an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of another optional cluster management method according to an embodiment of the present application;

[0027] Figure 6 is a schematic diagram of an application scenario of an optional cluster management method according to an embodiment of the present application;

[0028] Figure 7 is a schematic structural diagram of an optional cluster management device according to an embodiment of the present application;

[0029] Figure 8It is a schematic structural diagram of an optional cluster management product according to an embodiment of the present application;

[0030] Figure 9 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] The present application will be described below with reference to the embodiments:

[0034] According to one aspect of the embodiments of the present application, a method for managing a cluster is provided. Optionally, in this embodiment, the above-mentioned method for managing a cluster can be applied to a hardware environment composed of a server 101 and a terminal device 103 as shown in Figure 1 the figure. As shown in Figure 1As shown in the figure, the server 101 is connected to the terminal 103 through a network and can be used to provide services for the terminal device or the application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. The database 105 can be set on the server or independently of the server to provide data storage services for the server 101. For example, a game data storage server. The above network can include, but is not limited to: a wired network, a wireless network. Among them, the wired network includes: a local area network, a metropolitan area network, and a wide area network. The wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The terminal device 103 can be a terminal configured with an application and can include, but is not limited to, at least one of the following: a mobile phone (such as an Android mobile phone, an iOS mobile phone, etc.), a laptop computer, a tablet computer, a handheld computer, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a mixed reality (MR) terminal, and other computer devices. The above server can be a single server, a server cluster composed of multiple servers, or a cloud server.

[0035] Combined with Figure 1 As shown in the figure, the above cluster management method can be executed by an electronic device, which can be a terminal device or a server. The above cluster management method can be implemented separately by the terminal device or the server, or jointly implemented by the terminal device and the server.

[0036] The above is only an example, and this embodiment does not make specific limitations.

[0037] Optionally, as an alternative implementation, as Figure 2 shown in the figure, the above cluster management method includes:

[0038] S202, set a primary management node in the first working group, where the primary management node is set to coordinate the nodes in the target cluster to provide services. The target cluster includes the first working group and the second working group. The nodes in the first working group are set as standby nodes for the nodes in the second working group;

[0039] Optionally, in the embodiments of the present application, the primary management node refers to a device responsible for daily management and service control within the first working group in the target cluster, including but not limited to video storage servers, network devices, computing nodes, etc., and is used to monitor and maintain the overall operating status of the cluster; the target cluster refers to a system composed of multiple nodes, including but not limited to video storage server clusters, cloud computing platforms, distributed database systems, etc.

[0040] Optionally, in the embodiments of the present application, the first working group and the second working group can be understood as different node sets divided according to working roles in the target cluster, including but not limited to primary management nodes, backup management nodes, and ordinary working nodes, etc.

[0041] It should be noted that the setting of the primary management node can be adjusted according to different requirements and environments, and the present application does not limit this. For example, the primary management node can be a physical device or a virtual machine; the nodes in the target cluster can be devices of the same type or devices of different types or configurations; the service can be a computing service, a storage service, or any other type of service.

[0042] Exemplarily, first configure the target cluster, including the first working group and the second working group, ensure that the nodes in the first working group can be used as spares for the nodes in the second working group, set the above-mentioned primary management node from the first working group, and this node will be responsible for coordinating the services in the target cluster. Implement a coordination mechanism on the primary management node to ensure the continuity and high availability of the services. Specifically, this primary management node can be used to monitor the node status of the nodes in the target cluster. Once a node failure or unavailability in the second working group is detected, it will automatically switch to the corresponding spare node in the first working group.

[0043] It should also be noted that the above-mentioned target cluster can include multiple working groups, and the nodes in each working group can be dynamically adjusted and allocated according to needs; the nodes in the first working group can also be used as spare nodes for other working groups to achieve the maximum utilization of resources.

[0044] S204, set a backup management node in the first working group, where the backup management node represents the backup management node of the primary management node;

[0045] Optionally, in the embodiments of the present application, the backup management node refers to a device in the first working group that is in a standby state and can quickly take over the functions of the primary management node when the primary management node fails, including but not limited to video storage servers, network devices, computing nodes, etc.

[0046] It should be noted that the setting of the above standby management node can vary according to the requirements and configurations of different workgroups, and the present application does not limit this. For example, the standby management node can be other nodes within the same workgroup, or nodes across workgroups, or virtual nodes in the cloud. In addition, the activation conditions of the standby management node can also be customized according to actual needs, such as automatic activation, manual activation, or triggering based on specific events.

[0047] Exemplarily, the process of setting the above standby management node can be summarized into the following steps:

[0048] First, determine the primary management node, and select one or more standby management nodes according to the requirements and policies of the workgroup.

[0049] Next, configure the standby management node to ensure that it can quickly take over the functions when the primary node is unavailable, which can include but is not limited to synchronizing the status information, configuration information, and any necessary security credentials of the primary management node; for example, periodically synchronize the service status, service configuration, and its own working status of each node obtained by the primary management node to the standby management node.

[0050] Furthermore, when an exception occurs in the above primary management node, promptly trigger the standby management node, and the standby management node realizes the functions of the above primary management node.

[0051] S206, in the case where an exception occurs in the primary management node, control the standby management node to take over the primary management node and coordinate the nodes in the target cluster to provide services.

[0052] It should be noted that the present application does not limit the reasons for the above exception in the primary management node, such as hardware failure, software crash, or network problem, etc.

[0053] Exemplarily, when the primary management node detects an exception, the standby management node will be activated; first, the standby management node will perform a fault detection to confirm the status of the primary management node; once it is confirmed that the primary management node is unavailable, the standby management node will start to take over its work, including but not limited to establishing communication with other nodes in the target cluster, synchronizing status information, and starting to coordinate the nodes in the cluster to ensure service continuity, and, the standby management node can also reassign tasks to each node, update configurations, and maintain the stability of the cluster until the primary node returns to normal or is permanently replaced.

[0054] In an exemplary embodiment, Figure 3 is a cluster schematic diagram of an optional cluster management method according to an embodiment of the present application, as Figure 3As shown, the target cluster may be an N+M cluster (N represents N active devices in the cluster working group, and M represents M standby devices in the cluster working group). There are four types of status nodes, namely: active devices (nodes in the second working group), standby devices (nodes in the first working group), DCS working machines (the active management nodes), and DCS backup machines (the standby management nodes). Specifically:

[0055] The main equipment provides business services to the outside world, and the DCS working machine manages all DCS services. After the DCS working machines have negotiated and determined, the DCS working machine initiates a request to designate a DCS backup machine. The DCS working machine monitors the business configurations of all main services, sends changes to the DCS working machine in a timely manner, and the DCS working machine promptly sets the business configurations to the DCS backup machine, so that the DCS backup machine and the DCS working machine have the same working capabilities.

[0056] In addition, the DCS working machine monitors the working status of all main services provided by the main equipment. If an abnormality occurs in the main service, the DCS working machine will initiate a master-slave switch and synchronize the corresponding host configuration to the selected destination backup equipment; and, when the DCS working machine monitors the recovery of the abnormal main equipment, it will actively initiate master-slave repair and business migration; and the DCS backup machine can be used to monitor the status of the DCS working machine and take over immediately if the DCS working machine is abnormal.

[0057] In another exemplary embodiment, taking the application scenario of a data center as an example, the main management node in the backup device is responsible for monitoring the operating status and service configuration of each server in the data center (the nodes in the above-mentioned target cluster), and some servers are main devices, and some servers are backup devices. Each server regularly sends its corresponding service information to the main management node, and the above-mentioned main management node stores the service information separately and periodically sends it to the backup management node in the backup device. When the main management node fails to work due to a fault, the backup management node immediately takes over, establishes communication with the servers in the data center, coordinates their services according to the service information of the servers, and ensures the continuous operation of the data center.

[0058] Through the embodiments of the present application, a primary management node is set in the first working group. The primary management node is configured to coordinate the nodes in the target cluster to provide services. The target cluster includes the first working group and the second working group. The nodes in the first working group are configured as standby nodes for the nodes in the second working group. A standby management node is set in the first working group. The standby management node represents the standby management node of the primary management node. In the case where the primary management node fails, the standby management node is controlled to take over the primary management node and coordinate the nodes in the target cluster to provide services. This method simplifies the management node switching process, reduces the time and complexity of resource re-acquisition, achieves the purpose of effectively improving the reliability and continuity of the cluster service, and avoids problems such as service interruption and network splitting during the process of re-electing the management node. Thus, the cluster can provide stable services. Therefore, the technical problem of difficulty in providing normal services during the process of re-electing the management node when the management node of the cluster fails is effectively solved.

[0059] As an optional solution, the above method further includes: establishing communication between the primary management node and the nodes in the target cluster. The primary management node is used to periodically obtain the node service information sent by the nodes in the target cluster. The node service information includes the service status and service configuration of the nodes, and the node service information corresponds to the services provided by the nodes. Controlling the primary management node to periodically send the node service information to the standby management node. In the case where the primary management node fails, controlling the standby management node to establish communication with the nodes in the target cluster to take over the primary management node to periodically obtain the node service information and coordinate the nodes in the target cluster to provide services based on the node service information.

[0060] Optionally, in the embodiments of the present application, the above node service information refers to the information corresponding to the services provided by the nodes, including but not limited to the service status and service configuration. The service status may include but not limited to whether the node is currently online, whether it is providing services, whether the service is running normally, etc. The service configuration refers to the configuration parameters of the node, such as resource allocation, service parameters, security settings, etc.

[0061] It should be noted that the above primary management node and standby management node may be different types of devices or the same type of devices. The node service information may include various types of data, such as performance indicators, resource usage, service logs, etc. The nodes in the target cluster may be physical devices or virtual devices, and the present application does not make any limitations in this regard.

[0062] Exemplarily, the primary management node periodically obtains node service information from the nodes in the target cluster, and then periodically sends the node service information to the standby management node. When the primary management node fails, the standby management node will take over its functions, establish communication with the nodes in the target cluster, and coordinate the nodes based on the node service information to provide services.

[0063] It should also be noted that the frequencies at which different nodes in the above-mentioned target cluster send their own service information can be the same or different. Moreover, the nodes in the above-mentioned target cluster can choose to send their own service information, or choose not to send their own service information. This application also does not limit this.

[0064] Through the embodiments of this application, by adopting the cooperative working mechanism of the primary management node and the standby management node, the technical effects of high availability and failover are achieved, and the purpose of improving service stability and reliability is achieved.

[0065] As an optional solution, the above method further includes: controlling the primary management node to periodically send its own working status to the standby management node; detecting the connection status between the primary management node and the standby management node when the sending of the working status fails; re-determining the standby management node when the connection status indicates that the primary management node is normal and the standby management node is abnormal; controlling the primary management node to re-send the working status to the standby management node when the connection status indicates that both the primary management node and the standby management node are normal; obtaining the reporting time corresponding to the standby management node receiving the working status; detecting the connection status and the service resource information stored by the standby management node when the reporting time exceeds a preset time threshold; sending a node notification message to each of the nodes and updating the working status of the target cluster when the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the node service information corresponding to each node in the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0066] Optionally, in the embodiments of the present application, the above service resource information may include, but is not limited to, configuration parameters of each node in the cluster, resource usage, service load, network status, etc. The above working state refers to the real-time operating status of the primary management node, including but not limited to device health status, system operating parameters, service processing status, error and warning information, etc. The above preset time threshold can be flexibly set manually. The above node notification message may include, but is not limited to, forms such as text, image, voice, video, etc., and is used to convey the information that the standby management node will take over the cluster management service to each node in the cluster, ensuring that all nodes can respond in a timely manner and make corresponding status adjustments to ensure the smooth transition of the cluster service.

[0067] It should be noted that the communication protocol between the above primary management node and each node, the content of the status report, the specific setting of the preset time threshold, etc. can be adjusted according to different cluster types and service requirements, and the present application does not make any limitations in this regard.

[0068] Exemplarily, first, the above primary management node periodically sends the working state to the standby management node, including its own operating conditions, fault information, etc.; when it detects that the sending of the working state fails, it further detects the connection state between the primary management node and the standby management node to determine whether it is caused by a network fault; if the network connection is normal, the primary management node attempts to resend the working state to the standby management node.

[0069] It should be noted that the above standby management node detects the connection state and the abnormality of the primary management node, including but not limited to the following two situations:

[0070] S1. The standby management node determines through the connection state that the data transmission channel between itself and the primary management node is abnormal, while the primary management node itself is not abnormal. For example, the data interaction between the primary management node and other service nodes proceeds normally, and the data interaction between the standby management node and other service nodes can also proceed normally, but only the data interaction between the primary management node and the standby management node cannot be carried out;

[0071] S2. The standby management node detects that the primary management node is abnormal. For example, it cannot normally obtain the working state of the primary management node, while the data transmission channel between the two is not abnormal.

[0072] It can be understood that both of the above situations can trigger the standby management node to manage the state switch of the entire cluster and the corresponding nodes of the cluster.

[0073] Furthermore, the above standby management node records the reporting time of each received working state and sets a preset time threshold. Once the working state is not received after exceeding this threshold, the standby management node will detect its own connection state and the stored service resource information.

[0074] Next, when the standby management node confirms that the primary management node is abnormal and it has complete service resource information itself, the standby management node immediately starts the takeover mechanism, sends a notification message to each node in the cluster, indicating that it will take over the functions of the primary management node, and updates the working state of the cluster to ensure the normal operation of the cluster.

[0075] In an exemplary embodiment, Figure 4 is a schematic diagram of an optional cluster management method according to an embodiment of the present application. As Figure 4 shown, taking the application scenario of a video storage server cluster as an example, it includes but is not limited to:

[0076] After the above-mentioned primary management node ( Figure 4 the DCS working machine shown) detects the communication interruption with the standby management node, it will automatically attempt to re-establish the connection and send the latest working state information to the standby management node to ensure that the standby management node can update its own information library in a timely manner.

[0077] After the above-mentioned standby management node ( Figure 4 the DCS backup machine shown) receives the working state information of the primary management node, it will continuously monitor the health status of the primary management node. Once it is found that the primary management node has not sent the working state information for a long time, the standby management node will immediately start the takeover process, take over the management control right of the cluster, send a notice to each node in the cluster, and update the cluster status to ensure the continuity of the video storage service.

[0078] Through the embodiment of the present application, by adopting the mechanism of cooperative work and automatic detection and switching of states of the primary and standby management nodes, the technical effect of quickly responding to the failure of the primary management node is achieved, and the purpose of improving the stability of the cluster service and reducing the service interruption time is achieved. Especially in fields such as video storage servers that require high availability, the application of this technical solution will greatly improve the reliability of the system and the user experience.

[0079] As an optional solution, setting the primary management node in the above-mentioned first working group includes: initiating a cluster configuration request, dividing the nodes in the above-mentioned target cluster into the above-mentioned first working group and the above-mentioned second working group; determining the above-mentioned primary management node that meets the target node conditions from the above-mentioned first working group, where the above-mentioned target node conditions include being able to communicate normally with any node in the above-mentioned target cluster and being recognized by the nodes in the above-mentioned first working group; configuring the above-mentioned primary management node to manage the nodes in the above-mentioned target cluster.

[0080] Optionally, in the embodiments of the present application, the above-mentioned cluster configuration request refers to an instruction or signal initiated by a system administrator or an automatic configuration mechanism for initializing or reconfiguring a cluster management architecture, which may include, but is not limited to, defining the roles of the primary and standby management nodes, dividing the first working group and the second working group, setting the inter-node communication protocol, defining the node status detection frequency and threshold, etc.

[0081] It should be noted that, in the embodiments of the present application, if the current primary management node is unable to perform cluster configuration, any node other than the non-primary management node can be used to configure the entire cluster, and the configuration information related to the node is sent to the primary management node.

[0082] Exemplarily, first, the above-mentioned cluster configuration request is initiated by the client to reasonably divide the nodes in the target cluster into a first working group and a second working group; then, a node that meets the target node conditions is found in the above-mentioned first working group as the primary management node, and the above-mentioned target node conditions mainly include the ability to communicate normally with any node in the cluster and the recognition by other nodes; finally, the selected primary management node is configured so that the primary management node can manage other nodes.

[0083] In an exemplary embodiment, taking the application scenario of a video storage server cluster as an example, in the initial configuration stage of the cluster nodes, the cluster management system initiates a cluster configuration request to divide the nodes in the cluster into two working groups: the first working group includes service nodes that are currently in a standby state and do not provide business services, while the second working group includes service nodes that are currently in a running state and provide business services. Then, a node that meets the target node conditions is selected from the first working group as the primary management node, and this node needs to have the ability to communicate normally with all other nodes in the cluster and obtain the recognition of all nodes in the first working group through negotiation.

[0084] Through the embodiments of the present application, dynamic role assignment and configuration are performed for cluster nodes, achieving the goal of quickly determining and configuring the primary management node, and achieving the purpose of optimizing the cluster management efficiency, enhancing the system stability and service continuity.

[0085] As an alternative solution, configuring the above-mentioned primary management node to manage the nodes in the above-mentioned target cluster includes: establishing communication between the primary management node and the nodes in the above-mentioned target cluster, where the primary management node is used to periodically obtain the service status and service configuration sent by the nodes in the above-mentioned target cluster; when the primary management node receives the first service status sent by the first service node, sending the first service status to the standby management node and updating the working status of the above-mentioned target cluster, where the first service node represents the node with a changed service status; when the primary management node receives the second service configuration sent by the second service node, sending the second service configuration to the standby management node and updating the working status of the above-mentioned target cluster, where the second service node represents the node with a changed service configuration.

[0086] It should be noted that the establishment method of the communication between the primary management node and the nodes in the above-mentioned target cluster, the acquisition frequency of the service status and service configuration, and the update mechanism of the working status, etc., can all be flexibly adjusted according to actual business requirements and network conditions. For example, the communication can be based on the TCP / IP protocol, UDP, HTTP or other custom protocols.

[0087] It should also be noted that the above-mentioned service status can include but is not limited to the health status, performance metrics, error logs, etc. of the nodes; the above-mentioned service configuration can include but is not limited to network settings, resource allocation policies, security configurations, etc.; and the update of the above-mentioned working status can be completed through methods such as real-time broadcasting, periodic polling or event-driven, and the present application does not make any limitations in this regard.

[0088] In an exemplary embodiment, taking the application scenario of a video storage server cluster as an example, the primary management node first establishes stable network communication with all server nodes.

[0089] During the normal service provision of the target cluster, when the health status of a certain server (the first service node) changes, such as detecting disk read / write errors, the server will immediately notify the primary management node of this status change.

[0090] Then, after receiving this information, the primary management node will not only update its own status information, but also synchronize the first service status to the standby management node and update the working status of the entire cluster, notifying other nodes to adjust the service load to avoid data writing to the faulty node.

[0091] Similarly, during the normal service provision of the target cluster, if the configuration of a certain server (the second service node) changes, such as increasing the read / write rate or adjusting the storage policy, the server will also notify the primary management node.

[0092] Next, the primary management node synchronizes the updated second service configuration to the standby management node and updates the working status of the target cluster to ensure that all nodes follow the latest configuration and optimize the cluster performance.

[0093] Through the embodiments of the present application, a real-time communication and status synchronization mechanism between the primary management node and the nodes in the target cluster is adopted, achieving the technical effect of quickly responding to service status and service configuration changes and achieving the purpose of improving cluster management efficiency.

[0094] As an optional solution, the above method further includes: when the first working group includes multiple nodes, obtaining other nodes in the first working group except the primary management node; determining the node with the lowest load among the other nodes in the idle state as the standby management node.

[0095] As an optional solution, the above method further includes: when the first working group only includes the primary management node, obtaining other nodes in the second working group; determining the node with the lowest load among the other nodes in the idle state as the standby management node.

[0096] Optionally, in the embodiments of the present application, the above idle state refers to the situation where a node does not execute key tasks or services within a certain period of time, which may include but is not limited to that the processing capacity is not fully utilized, there is no ongoing high-priority task, and the resource consumption is lower than a preset threshold.

[0097] It should be noted that the load evaluation criteria of the above nodes, the definition of the idle state, the election algorithm of the standby management node, etc. can all be flexibly set according to factors such as the scale of the cluster, business characteristics, and resource allocation strategies. For example, the lowest load can be evaluated based on indicators such as CPU usage, memory occupancy, network bandwidth consumption, and disk I / O; the determination of the idle state may include but is not limited to being determined through real-time monitoring data of the target cluster, historical statistical information, or a preset service level agreement; furthermore, the election of the above standby management node may adopt a simple principle of the lowest load or a more complex multi-factor comprehensive consideration algorithm, and the present application does not limit this.

[0098] In an exemplary embodiment, taking the application scenario of a video surveillance system as an example, when the first working group includes multiple servers, the primary management node will first obtain the real-time load information of these servers, including but not limited to CPU usage, memory occupancy rate, network traffic, etc.

[0099] Next, the primary management node will filter out the servers in the idle state, that is, the servers with relatively low current load and not executing critical tasks. Finally, based on the load evaluation results, the primary management node selects the server with the lowest load among them as the standby management node and performs configuration and status synchronization.

[0100] In an exemplary embodiment, Figure 5 is a schematic diagram of another optional cluster management method according to an embodiment of the present application. The process of creating a standby management node by the primary management node is as Figure 5 shown.

[0101] In another exemplary embodiment, assume that the current first working group consists only of the primary management node, indicating that the first working group has only one node and this node has been determined as the above-mentioned primary management node. In this case, the primary management node in the embodiment of the present application will actively detect all nodes in the second working group and determine the current load and idle degree of the nodes in the second working group according to indicators such as CPU utilization rate, memory usage, network traffic, and disk I / O load.

[0102] Furthermore, when the primary management node needs to determine the standby management node, the primary management node can filter out the nodes in the idle state from the second working group.

[0103] It should be noted that the determination of the above idle state may include but is not limited to determining based on the resource utilization rate of the node being lower than a preset threshold or not executing high-priority tasks within a certain period of time. Further, after determining the nodes in the idle state, the primary management node can further analyze their load conditions and select the node with the lowest load as the standby management node.

[0104] In yet another exemplary embodiment, if in the first working group, in addition to the target primary node, there are also multiple standby nodes, however, when these multiple standby nodes are all in the state of providing services or all in the high-load state, it is also possible to select a general service node with relatively low load and high stability from the second working group as the standby management node.

[0105] That is to say, in the embodiments of the present application, priority is given to determining a standby management node in the first working group. When there is only one node that has been set as the target primary node in the first working group, or when the nodes other than the target primary node in the first working group cannot be set as the target primary node, the present application can also consider determining a standby management node from the ordinary service nodes that are providing services normally in the second working group. By intelligently selecting the standby management node, the fault recovery time can be effectively reduced, the overall stability and availability of the cluster can be improved. At the same time, resource utilization is optimized, unnecessary performance waste is avoided, and the technical effects of enhancing system management capabilities and high availability are achieved.

[0106] As an alternative solution, the above method further includes: controlling the primary management node to periodically send its own working status to the standby management node; detecting the connection status between the primary management node and the standby management node when the sending of the working status fails; and re-determining the standby management node from the first working group according to the primary management node when the connection status indicates that the primary management node is normal and the standby management node is abnormal.

[0107] Optionally, in the embodiments of the present application, the working status of the primary management node refers to the operating condition and management ability of the primary management node, including but not limited to its health status, resource occupancy, tasks being executed, communication status with other nodes in the cluster, etc.; the connection status refers to the real-time status of the communication link between the primary management node and the standby management node, including but not limited to indicators such as network connectivity, data transmission delay, and packet loss rate.

[0108] It should be noted that the primary management node and the standby management node can not only coordinate and manage each node in the target cluster, but also serve as the standby nodes corresponding to the nodes in the second working group. When the nodes in the second working group are abnormal, they take over the nodes in the second working group and provide services externally.

[0109] In an exemplary embodiment, taking the application scenario of a video surveillance server cluster as an example, the primary management node (DCS working machine) sends its working status, including health check results, resource usage, and service task lists, etc., to the standby management node (DCS backup machine) at a predetermined period, such as every 10 seconds. If the standby management node fails to receive the working status information successfully for several consecutive times, the primary management node will automatically trigger a connection status detection to check the communication link between the two. If the detection result shows that the primary management node is normal while the standby management node cannot respond, the primary management node will immediately select a node with good working status and moderate load from the first working group (i.e., N primary devices in N+M) as the new standby management node and start synchronizing necessary status information and configurations to it, so as to ensure that when the primary management node fails, the new standby management node can quickly take over the cluster management task and avoid service interruption.

[0110] Through the embodiments of the present application, by adopting a periodic status report and a dynamic connection status detection mechanism, the monitoring and management of real-time communication between the primary management node and the standby management node are realized, achieving the purpose of timely detecting communication anomalies and quickly switching the standby node to ensure the high availability of the cluster.

[0111] As an optional solution, setting the standby management node in the above first working group includes: obtaining the reporting time corresponding to the working status sent by the primary management node received by the standby management node; when the reporting time exceeds a preset time threshold, detecting the connection status between the standby management node and the primary management node and the service resource information stored by the standby management node; when the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the service status and service configuration of each node in the target cluster, sending a node notification message to each node and updating the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0112] Optionally, in the embodiments of the present application, the preset time threshold refers to a set value used to evaluate the communication delay or interruption between the primary management node and the standby management node, including but not limited to several seconds to several minutes, and its specific value should be reasonably set according to the network environment, business requirements, and fault tolerance.

[0113] In an exemplary embodiment, taking the application scenario of a video surveillance server cluster as an example, the standby management node (DCS backup machine) continuously records the reporting time of receiving the working status of the primary management node (DCS working machine);

[0114] When it is detected that the reporting time exceeds the preset time threshold, such as 30 seconds, the standby management node will actively detect the network connection status with the primary management node, and at the same time check whether the business resource information stored by the standby management node itself is complete, including the latest service status and service configuration of all nodes in the cluster.

[0115] If the detection result shows that the connection between the standby management node and the primary management node has been disconnected, and the standby management node itself is in a normal state and has stored all necessary business resource information, then the standby management node will send a "node notification message" to each node in the target cluster, indicating that it will take over the cluster management task.

[0116] At the same time, the standby management node can also update the working state of the target cluster to reflect the current management node change, ensuring the business continuity and stability of the cluster.

[0117] Through the embodiments of the present application, by adopting dynamic communication delay detection and business resource information integrity verification, it is realized that when the primary management node has an abnormality, the standby management node can quickly and accurately judge whether it meets the takeover conditions, and when the conditions are met, quickly switch to the management state, realizing the automation of business takeover, achieving the purpose of reducing the fault recovery time, improving the system high availability and user experience.

[0118] As an optional solution, after setting the standby management node in the above first working group, it includes: when the primary management node determines that the primary service node has an abnormality, controlling the primary management node to initiate a primary-standby switch operation to instruct the standby service node to take over the primary service node, and sending the service configuration of the primary service node to the standby service node, where the primary service node is a node in the above second working group, and the standby service node is a node in the above first working group; when the primary service node has repaired the abnormality, controlling the primary management node to initiate a primary-standby recovery operation to instruct the primary service node to take over the standby service node, and sending the service configuration of the standby service node to the primary service node.

[0119] Optionally, in the embodiments of the present application, the above primary-standby switch operation refers to a process initiated by the primary management node when the primary service node has an abnormality, and the purpose is to promote the standby service node to the primary state to take over the responsibilities of the faulty node, including but not limited to updating the role status of each node in the cluster, synchronizing service configuration parameters to the standby service node, adjusting network routing and load balancing policies, etc.

[0120] Optionally, in the embodiments of the present application, the above primary / standby recovery operation refers to a process initiated by the primary management node after the primary service node repairs the anomaly, with the aim of restoring the service state before the fault, repositioning the primary service node as the primary, and demoting the previous standby service node to standby, including but not limited to resynchronizing the service configuration to the primary service node, adjusting the node roles and cluster status, updating the load balancing rules, etc.

[0121] In an exemplary embodiment, taking the application scenario of a data storage server cluster as an example, assume that the primary service node (the node within the second working group) in the target cluster has an anomaly due to CPU overload. The primary management node (DCS working machine) will detect this situation and automatically initiate a primary / standby switch operation, instructing the standby service node located in the standby working group (the node within the first working group) to take over the responsibilities of the faulty node.

[0122] Specifically, during the switch process, the primary management node will send the service configuration of the faulty node, including storage policies, access permissions, load balancing settings, etc., to the standby service node to ensure that it can seamlessly take over and provide the same service.

[0123] When the primary service node repairs the anomaly and the CPU load returns to normal, the primary management node will initiate the primary / standby recovery operation again, migrate the service configuration from the standby service node back to the primary service node, and adjust its status to primary, while the standby service node returns to the standby state to restore the original working mode of the cluster.

[0124] Through the embodiments of the present application, when the primary service node has an anomaly, it is possible to quickly and accurately perform a primary / standby switch, and after the anomaly is repaired, smoothly and efficiently restore to the primary state, realizing high service continuity and the self-healing ability of the cluster, achieving the purpose of improving the overall stability of the system and the user experience.

[0125] In an exemplary embodiment, the cluster management method proposed in the present application can implement a hot standby management system for cluster scheduling management services.

[0126] First, explain the nouns that appear in the embodiments of the present application:

[0127] N+M cluster: A total of N+M service nodes in the system jointly form a cluster to provide services externally. N represents N primary devices in the cluster working group, and M represents M standby devices in the cluster working group. Any number of devices (not more than M) in the working group N fail, and the devices in the standby group M seamlessly replace them and take over the work of the primary devices to ensure that the external service is not interrupted during the anomaly.

[0128] Exemplarily, Figure 6It is a schematic diagram of an application scenario of an optional cluster management method according to an embodiment of the present application. As Figure 6 shown: The DCS working machine (the above-mentioned primary management node) provides a dispatching console service, which is the center of the entire cluster control and is responsible for the management of the entire cluster. The roles played by the DCS working machine in the cluster include but are not limited to: initiating negotiation and forming a cluster; detecting the status of each service node and managing the status set; detecting node anomalies, initiating primary / standby switching, and migrating services; repairing the primary device and migrating back the cluster services, etc.

[0129] In the related art, when the DCS working machine fails, service interruption will occur. Even if the election negotiation management node is recreated, there will be a state where the cluster is not readable or writable for a certain period of time during this period. Moreover, the re-negotiated DCS working machine needs to reorganize the cluster and obtain cluster resources. Specifically, there is a possibility of split-brain during the re-negotiation process, and external intervention in the cluster work is required.

[0130] Based on the above technical problems, the embodiment of the present application can use the DCS working machine to create a DCS backup machine (the above-mentioned standby management node), and the two work together to reduce the pressure on the service nodes. If the DCS working machine is abnormal, the DCS backup machine can take over the management of the cluster in time, and the external services of the cluster will not be interrupted during the switching between the two. There is no need to initiate an election again, and the cluster-related resource information and configuration of the DCS working machine can be synchronized to the DCS backup machine in time. In other words, both the DCS working machine and the DCS backup machine are selected from the M standby working group devices, and both the DCS working machine and the DCS backup machine can take over the work of the primary device. All nodes in the above N+M cluster (the above-mentioned target cluster) have the ability to provide the same service, and the DCS working machine can negotiate to assign different states to different nodes. Specifically, there are four types of status nodes in the N+M cluster: primary device, standby device, standby machine + DCS working machine, and standby machine + DCS backup machine. Specifically:

[0131] The primary device provides external business services, and the DCS working machine manages all DCS services; after the DCS working machine is negotiated and determined, the DCS working machine initiates a request to specify the DCS backup machine; the DCS working machine listens to the business configurations of all primary services and notifies the DCS working machine in time when changes occur. The DCS working machine timely sets the business configurations to the DCS backup machine, so that the DCS backup machine and the DCS working machine have the same working ability.

[0132] In addition, the DCS working machine monitors the working status of all main services provided by the main equipment. If an abnormality occurs in the main service, the DCS working machine will initiate a master-slave switch and synchronize the corresponding host configuration to the selected destination backup equipment; and, when the DCS working machine monitors the recovery of the abnormal main equipment, it will actively initiate master-slave repair and business migration; and the DCS backup machine can be used to monitor the status of the DCS working machine and take over immediately if the DCS working machine is abnormal.

[0133] In an exemplary embodiment, the steps of negotiating a DCS working machine and creating a DCS backup machine may include but are not limited to:

[0134] S1, the client initiates the request and specifies the primary and backup devices;

[0135] S2, negotiate the DCS working machine from the backup equipment. Negotiation rules and requirements:

[0136] S2-1, select the standby machine;

[0137] S2-2, the standby machine initiates negotiation;

[0138] S3, start the required functions of the DCS working machine. Specifically including:

[0139] S3-1, keep-alive detection with the main service;

[0140] S3-2, and other DCS keep-alive detection;

[0141] It can be understood as the keep-alive detection between the DCS working machine and the backup equipment, and the keep-alive detection between the DCS working machine and the main equipment.

[0142] S4, the DCS working machine connects, obtains, and subscribes to the business configuration of the devices in the cluster;

[0143] S5, the DCS working machine sets the working status of each main and backup equipment;

[0144] S6, the DCS working machine creates a DCS backup machine.

[0145] In another exemplary embodiment, the DCS working machine may include but is not limited to implementing:

[0146] S1, the DCS working machine regularly reports the working status to the DCS backup machine;

[0147] S2, DCS will modify the monitored business configuration to DCS backup machine;

[0148] S3, report failure, need to check its own connection status. If normal, reselect DCS backup machine;

[0149] S4, update the cluster status.

[0150] In another exemplary embodiment, the DCS backup machine may include, but is not limited to, implementing:

[0151] S1. Regularly detect the reporting time of the DCS working machine;

[0152] S2. When the reporting time of the DCS working machine exceeds the threshold time, detect its own connection status and whether it has the service resources (i.e., service configuration information) of all cluster nodes;

[0153] S3. Notify each node in the cluster that the DCS backup machine will take over the node management from the DCS working machine;

[0154] S4. Update the cluster status.

[0155] Through the embodiments of the present application, a DCS working machine is set in the cluster, and the DCS working machine selects the DCS backup machine. Moreover, the DCS working machine and the DCS backup machine can work together. The DCS backup machine can timely monitor the status of the DCS working machine and can timely perform a switch, ensuring that even if the central management node fails, there will be no cluster service failure, and there is no need for additional re-negotiation and re-acquisition of cluster service resources. The working mode of timely synchronization of the DCS working machine to the DCS backup machine can ensure that the two have the same working ability and can immediately take over the cluster and manage the cluster, guaranteeing the rapid takeover of the service.

[0156] It can be understood that in the specific implementation manner of the present application, when it comes to data such as user information, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0157] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0158] According to another aspect of the embodiments of the present application, there is also provided a cluster management device for implementing the above cluster management method. As Figure 7 shown, the device includes:

[0159] The primary setting module 702 is used to set a primary management node in the first working group. The primary management node is set to coordinate the nodes in the target cluster to provide services. The target cluster includes the first working group and the second working group. The nodes in the first working group are set as standby nodes for the nodes in the second working group;

[0160] The standby setting module 704 is used to set a standby management node in the first working group. The standby management node represents the standby management node of the primary management node;

[0161] The control module 706 is used to control the standby management node to replace the primary management node and coordinate the nodes in the target cluster to provide services when the primary management node has an exception.

[0162] As an optional solution, the above device is further used to: establish communication between the primary management node and the nodes in the target cluster. The primary management node is used to periodically obtain the node service information sent by the nodes in the target cluster. The node service information includes the service status and service configuration of the nodes, and the node service information corresponds to the services provided by the nodes; control the primary management node to periodically send the node service information to the standby management node; when the primary management node has an exception, control the standby management node to establish communication with the nodes in the target cluster to replace the primary management node to periodically obtain the node service information and coordinate the nodes in the target cluster to provide services based on the node service information.

[0163] As an optional solution, the above device is further used to: control the primary management node to periodically send its own working status to the standby management node; when the sending of the working status fails, detect the connection status between the primary management node and the standby management node; when the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node. When the connection status indicates that both the primary management node and the standby management node are normal, control the primary management node to re-send the working status to the standby management node; obtain the reporting time corresponding to the working status received by the standby management node; when the reporting time exceeds the preset time threshold, detect the connection status and the service resource information stored by the standby management node; when the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the node service information corresponding to each node in the target cluster, send a node notification message to each node and update the working status of the target cluster. The node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0164] As an alternative solution, the above device is used to set the primary management node in the first working group in the following manner, including: initiating a cluster configuration request to divide the nodes in the target cluster into a first working group and a second working group; determining the primary management node that meets the target node conditions from the first working group, where the target node conditions include being able to communicate normally with any node in the target cluster and being recognized by the nodes in the first working group; and configuring the primary management node to manage the nodes in the target cluster.

[0165] As an alternative solution, the above device is used to configure the primary management node to manage the nodes in the target cluster in the following manner: establishing communication between the primary management node and the nodes in the target cluster, where the primary management node is used to periodically obtain the service status and service configuration sent by the nodes in the target cluster; when the primary management node receives the first service status sent by the first service node, sending the first service status to the standby management node and updating the working status of the target cluster, where the first service node represents the node whose service status has changed; when the primary management node receives the second service configuration sent by the second service node, sending the second service configuration to the standby management node and updating the working status of the target cluster, where the second service node represents the node whose service configuration has changed.

[0166] As an alternative solution, the above device is further used to: when the first working group includes multiple nodes, obtain the other nodes in the first working group except the primary management node; and determine the node with the smallest load among the other nodes in the idle state as the standby management node.

[0167] As an alternative solution, the above device is further used to: when the first working group only includes the primary management node, obtain the other nodes in the second working group; and determine the node with the smallest load among the other nodes in the idle state as the standby management node.

[0168] As an alternative solution, the above device is further used to: control the primary management node to periodically send its own working status to the standby management node; when the sending of the working status fails, detect the connection status between the primary management node and the standby management node; and when the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node from the first working group according to the primary management node.

[0169] As an alternative solution, the above-mentioned device is used to set a standby management node in the first working group in the following manner: obtain the reporting time corresponding to the working status sent by the primary management node received by the standby management node; in the case where the reporting time exceeds a preset time threshold, detect the connection status between the standby management node and the primary management node, and the service resource information stored by the standby management node; in the case where the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the service status and service configuration of each node in the target cluster, send a node notification message to each node, and update the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

[0170] As an alternative solution, the above-mentioned device is further used to: after setting a standby management node in the first working group, in the case where the primary management node determines that the primary service node has an abnormality, control the primary management node to initiate a primary-standby switch operation to instruct the standby service node to replace the primary service node, and send the service configuration of the primary service node to the standby service node, where the primary service node is a node in the second working group, and the standby service node is a node in the first working group; in the case where the primary service node has repaired the abnormality, control the primary management node to initiate a primary-standby recovery operation to instruct the primary service node to replace the standby service node, and send the service configuration of the standby service node to the primary service node.

[0171] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0172] Regarding the device in the above-mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here in detail.

[0173] According to one aspect of the present application, a computer program product is provided, and the computer program product includes a computer program.

[0174] The serial numbers of the embodiments of the present application above are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0175] Figure 8 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.

[0176] It should be noted that Figure 8 The computer system 800 of the illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0177] As Figure 8 shown, the computer system 800 includes a central processing unit 801 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 802 (ROM) or the program loaded from the storage section 808 into the random access memory 803 (RAM). In the random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other via a bus 804. The input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.

[0178] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that the computer program read from it can be installed into the storage section 808 as needed.

[0179] Specifically, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 809 and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, various functions defined in the system of the present application are executed.

[0180] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, various functions provided by the embodiments of the present application are executed.

[0181] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned cluster management method is further provided. The electronic device may be Figure 1 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 9 shown, the electronic device includes a memory 902 and a processor 904. A computer program is stored in the memory 902, and the processor 904 is configured to execute the steps in any of the above method embodiments through the computer program.

[0182] Optionally, in this embodiment, the above-mentioned electronic device may be at least one of multiple network devices in a computer network.

[0183] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the methods in the embodiments of the present application through the computer program.

[0184] Optionally, those of ordinary skill in the art can understand that Figure 9 the structure shown is only schematic, Figure 9 and it does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 9 , or have a different configuration from that shown in Figure 9 .

[0185] Among them, the memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the cluster management method and device in the embodiments of the present application. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, implements the above-mentioned cluster management method. The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 902 may further include a memory remotely set relative to the processor 904, and these remote memories can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and their combinations. Among them, the memory 902 can specifically but not limitedly be used to store the above-mentioned node service information and other information. As an example, as Figure 9As shown, the above-mentioned memory 902 may but is not limited to include the primary setting module 702, the standby setting module 704, and the control module 706 in the management device of the above-mentioned cluster. In addition, it may also include but is not limited to other module units in the management device of the above-mentioned cluster, which will not be elaborated in this example.

[0186] Optionally, the above-mentioned transmission device 906 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 906 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 906 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0187] In addition, the above-mentioned electronic device further includes: a display 908, which is used to display the above-mentioned node service information; and a connection bus 910, which is used to connect each module component in the above-mentioned electronic device.

[0188] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer network, and any form of computing device, such as an electronic device such as a server or a terminal, can become a node in the blockchain system by joining the peer-to-peer network.

[0189] According to one aspect of the present application, there is provided a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the cluster management method provided in various optional implementation manners of the above-mentioned cluster management.

[0190] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store the methods for executing the embodiments of the present application.

[0191] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above-mentioned embodiments can be completed by a program instructing the relevant hardware of the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disc, etc.

[0192] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0193] If the integrated units in the above embodiments are implemented in the form of software function units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0194] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0195] In the several embodiments provided by the present application, it should be understood that the disclosed application program can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0196] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0197] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software function units.

[0198] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for managing a cluster, characterized in that, Including: A primary management node is set in the first working group. The primary management node is configured to coordinate nodes in a target cluster to provide services. The target cluster includes the first working group and the second working group. Nodes in the first working group are configured as standby nodes for nodes in the second working group. The first working group includes service nodes that are currently in a standby state and not providing business services. The second working group includes service nodes that are currently in an operating state and providing business services; A standby management node is set in the first working group. Control the primary management node to periodically obtain node service information of nodes in the target cluster, and send its own working state and the node service information to the standby management node. Control the standby management node to record the reporting time of the working state of the primary management node received each time. In the case where the reporting time exceeds a preset time threshold, control the standby management node to detect its own connection state and stored business resource information. The standby management node represents the standby management node of the primary management node; In the case where the primary management node is abnormal, control the standby management node to perform a fault detection to determine whether the primary management node is available. In the case where the primary management node is unavailable and the standby management node includes complete business resource information, control the standby management node to replace the primary management node and coordinate nodes in the target cluster to provide services; The method further includes: controlling the primary management node to periodically send its own working state to the standby management node; in the case where the sending of the working state fails, detecting the connection state between the primary management node and the standby management node; in the case where the connection state indicates that the primary management node is normal and the standby management node is abnormal, re-determining the standby management node. In the case where the connection state indicates that both the primary management node and the standby management node are normal, controlling the primary management node to re-send the working state to the standby management node; obtaining the reporting time corresponding to the standby management node receiving the working state; in the case where the reporting time exceeds the preset time threshold, detecting the connection state and the business resource information stored by the standby management node; in the case where the connection state indicates that the standby management node is normal, the primary management node is abnormal, and the business resource information includes the node service information corresponding to each node in the target cluster, sending a node notification message to each node and updating the working state of the target cluster. The node notification message indicates that the standby management node will establish communication with nodes in the target cluster; The method further includes: initiating a cluster configuration request to divide the nodes in the target cluster into the first working group and the second working group; determining the primary management node that meets the target node conditions from the first working group, where the target node conditions include being able to communicate normally with any node in the target cluster and being recognized by the nodes in the first working group; configuring the primary management node to manage the nodes in the target cluster.

2. The method according to claim 1, wherein The method further includes: Establishing communication between the primary management node and the nodes in the target cluster, where the primary management node is used to periodically obtain the node service information sent by the nodes in the target cluster, and the node service information includes the service status and service configuration of the node, and the node service information corresponds to the service provided by the node; Controlling the primary management node to periodically send the node service information to the standby management node; In the case where the primary management node has an abnormality, controlling the standby management node to establish communication with the nodes in the target cluster to take over the primary management node to periodically obtain the node service information, and coordinating the nodes in the target cluster based on the node service information to provide services.

3. The method according to claim 1, wherein The configuring the primary management node to manage the nodes in the target cluster includes: Establishing communication between the primary management node and the nodes in the target cluster, and the primary management node is used to periodically obtain the service status and service configuration sent by the nodes in the target cluster; In the case where the primary management node receives the first service status sent by the first service node, sending the first service status to the standby management node and updating the working state of the target cluster, where the first service node represents the node whose service status has changed; In the case where the primary management node receives the second service configuration sent by the second service node, sending the second service configuration to the standby management node and updating the working state of the target cluster, where the second service node represents the node whose service configuration has changed.

4. The method according to claim 1, wherein The method further includes: In the case where the first working group includes multiple nodes, obtaining the other nodes in the first working group except the primary management node; Determining the node with the smallest load among the other nodes in the idle state as the standby management node.

5. The method according to claim 1, characterized in that, The method further includes: In the case where the first working group only includes the primary management node, obtaining the other nodes in the second working group; Determining the node with the smallest load among the other nodes in the idle state as the standby management node.

6. The method according to claim 1, wherein The method further includes: Controlling the primary management node to periodically send its own working state to the standby management node; In the case where the sending of the working state fails, detecting the connection state between the primary management node and the standby management node; When the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node from the first working group according to the primary management node.

7. The method according to claim 1, wherein Setting a standby management node in the first working group includes: Obtaining the reporting time corresponding to the working status sent by the primary management node received by the standby management node; When the reporting time exceeds a preset time threshold, detecting the connection status between the standby management node and the primary management node, and the service resource information stored by the standby management node; When the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the service status and service configuration of each node in the target cluster, send a node notification message to each node, and update the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster.

8. The method according to claim 1, characterized in that, After setting the standby management node in the first working group, it includes: When the primary management node determines that the primary service node has an abnormality, control the primary management node to initiate a primary-standby switch operation to instruct the standby service node to replace the primary service node, and send the service configuration of the primary service node to the standby service node, where the primary service node is a node in the second working group, and the standby service node is a node in the first working group; When the primary service node has repaired the abnormality, control the primary management node to initiate a primary-standby recovery operation to instruct the primary service node to replace the standby service node, and send the service configuration of the standby service node to the primary service node.

9. A management device for a cluster, characterized in that, It includes: A primary setting module for setting a primary management node in the first working group, where the primary management node is set to coordinate the nodes in the target cluster to provide services, the target cluster includes the first working group and the second working group, the nodes in the first working group are set as standby nodes for the nodes in the second working group, the first working group includes service nodes that are currently in a standby state and do not provide business services, and the second working group includes service nodes that are currently in a running state and provide business services; A standby setting module for setting a standby management node in the first working group, controlling the primary management node to periodically obtain the node service information of the nodes in the target cluster, and sending its own working status and the node service information to the standby management node, controlling the standby management node to record the reporting time of each time it receives the working status of the primary management node, and when the reporting time exceeds the preset time threshold, controlling the standby management node to detect its own connection status and the stored service resource information, where the standby management node represents the standby management node of the primary management node; A control module, configured to, when an exception occurs in the primary management node, control the standby management node to perform a fault detection to determine whether the primary management node is available, and when the primary management node is unavailable and the standby management node includes complete service resource information, control the standby management node to replace the primary management node and coordinate the nodes in the target cluster to provide services; The apparatus is further configured to: control the primary management node to periodically send its working status to the standby management node; when the sending of the working status fails, detect the connection status between the primary management node and the standby management node; when the connection status indicates that the primary management node is normal and the standby management node is abnormal, re-determine the standby management node, and when the connection status indicates that both the primary management node and the standby management node are normal, control the primary management node to re-send the working status to the standby management node; obtain the reporting time corresponding to the standby management node receiving the working status; when the reporting time exceeds a preset time threshold, detect the connection status and the service resource information stored in the standby management node; when the connection status indicates that the standby management node is normal, the primary management node is abnormal, and the service resource information includes the node service information corresponding to each node in the target cluster, send a node notification message to each node and update the working status of the target cluster, where the node notification message indicates that the standby management node will establish communication with the nodes in the target cluster; The apparatus is further configured to: initiate a cluster configuration request to divide the nodes in the target cluster into the first working group and the second working group; determine the primary management node that meets the target node conditions from the first working group, where the target node conditions include being able to communicate normally with any node in the target cluster and being recognized by the nodes in the first working group; configure the primary management node to manage the nodes in the target cluster.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, where the computer program, when being run by an electronic device, executes the method according to any one of claims 1 to 8.

11. A computer program product, comprising a computer program, characterized in that, When being executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.

12. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.

Citation Information

Patent Citations

  • Server cluster system and load balancing implementation method thereof

    CN104283948A

  • Data processing node management method and device, equipment and storage medium

    CN111818159A

  • Backup method and device and server cluster system

    CN113472556A

  • Database cluster management control method and device, equipment and storage medium

    CN115981919A