Edge cluster management method and device, computing equipment and edge cluster
By self-checking the communication status of edge clusters and clouds at the edge, and monitoring the node status using heartbeat packets and EBPF programs, the problem of misjudgment caused by instability in the cloud is solved, and the accurate management of edge clusters and rational resource utilization is achieved.
Patent Information
- Application Number
- CN202510423435.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-12
AI Technical Summary
In the edge computing scenario, instability in communication between the cloud and the edge cluster leads to inaccurate detection of the communication status of the edge cluster, and it is impossible to reasonably manage the edge cluster.
By self-checking the communication status of the edge cluster and the communication status of the cloud at the edge end, managing the edge cluster, avoiding relying on cloud communication, and using heartbeat packets and EBPF programs to monitor the status of the nodes to achieve accurate communication status detection.
It improves the accuracy of communication status detection of edge clusters, reduces the number of service migrations, rationally utilizes resources, and ensures the stable operation of services of edge clusters.
Smart Images

Figure CN120475024A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computing technology, and in particular to a management method, apparatus, computing device, and edge cluster for an edge cluster. Background Art
[0002] In edge computing scenarios, the cloud provides computing resources, data processing, and storage services to the edge, while also managing and controlling the edge. Edge nodes utilize these resources to perform computing and storage close to end devices, improving response speed and reducing network bandwidth requirements. At the edge, multiple nodes providing services can form an edge cluster, working together to provide high availability and scalability. To ensure the stable operation of edge cluster services, edge cluster management is necessary.
[0003] Typically, the cloud establishes a communication connection with each node in the edge cluster. The cloud then monitors the communication status of each node based on the communication between the cloud and the node, thereby managing the edge cluster. However, the communication connection between the cloud and each node is unstable, resulting in inaccurate detection of the communication status of each node by the cloud, making it difficult to properly manage the edge cluster. Summary of the Invention
[0004] The embodiments of the present application provide an edge cluster management method, apparatus, computing device, and edge cluster, thereby enabling reasonable management of the edge cluster.
[0005] In a first aspect, embodiments of the present application provide an edge cluster management method, applied to a master node in a first edge cluster, the first edge cluster including a master node and slave nodes, each of which including a first service. The method includes: determining a communication status of the first edge cluster, the communication status of the first edge cluster being used to indicate whether the master node and the slave nodes can communicate normally; determining a communication status of a cloud, the communication status of the cloud being used to indicate whether the master node and the cloud can communicate normally; and managing the first service based on the communication status of the first edge cluster and the communication status of the cloud.
[0006] In this way, the first edge cluster is managed based on its communication status and the cloud's communication status. Because the first edge cluster's communication status is self-checked at the edge, without relying on communication with the cloud, this avoids misjudgments of status caused solely by unstable communication between the cloud and the first edge cluster. This makes detection of the first edge cluster's communication status more accurate, leading to more efficient management of the first edge cluster, reduced service migration times, and more efficient resource utilization.
[0007] In a possible implementation manner of the first aspect, the master node may be a computing device, and the slave node may be a computing device.
[0008] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates normal, the first edge cluster does not migrate the first service, and migration is used to indicate migrating the first service from one or more nodes to other nodes.
[0009] In this way, when the communication status indication of the edge cluster is normal and the communication status indication of the cloud is normal, it indicates that most nodes in the edge cluster are communicating normally and the master node is communicating normally. At this time, the edge cluster can provide the first service normally. Therefore, there is no need to migrate the first service in the edge cluster, which can ensure the stable operation of the edge cluster service, reduce the number of service migrations, and make rational use of resources.
[0010] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, the first edge cluster does not migrate the first service, and migration is used to indicate migrating the first service from one or more nodes to other nodes.
[0011] In this way, when the communication status indication of the first edge cluster is normal and the communication status indication of the cloud is abnormal, it indicates that although the communication between the master node and the cloud is abnormal, the communication of most nodes in the first edge cluster is normal. At this time, the first edge cluster can provide the first service normally. Therefore, there is no need to migrate the first service in the first edge cluster, which can ensure the stable operation of the service of the first edge cluster, reduce the number of service migrations, and reasonably use resources.
[0012] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: triggering re-election of the master node when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal.
[0013] In this way, when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, it means that although most nodes in the first edge cluster are communicating normally, the communication between the master node and the cloud is abnormal. Therefore, the re-election of the master node can be triggered and the master node can be replaced.
[0014] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud indicates a normal state, the first edge cluster does not migrate the first service, and migration is used to indicate migrating the first service from one or more nodes to other nodes.
[0015] In this case, if the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normality, this indicates that although the majority of nodes in the first edge cluster are experiencing communication abnormalities, that is, the first service in at least a majority of nodes in the first edge cluster is abnormal, since the master node is communicating normally with the cloud, this further indicates that the node with the abnormal communication status is a slave node. In this case, since at least the master node can normally provide the first service, it is confirmed that the first edge cluster can normally provide the first service. Migration of the first service in the first edge cluster is unnecessary, reducing the number of service migrations and optimizing resource utilization.
[0016] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normal, migrating the first service in the node where the communication status indicates abnormality in the first edge cluster to the node where the communication status indicates normal in the second edge cluster, the second edge node and the first edge node are located in the same local area network.
[0017] In this case, if the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normality, it indicates that the majority of nodes in the first edge cluster have abnormal communication, that is, the first service in at least the majority of nodes in the first edge cluster is abnormal. In this case, although the communication between the master node and the cloud is normal, the stability of the service operation can be improved by migrating the first service in the node with abnormal communication to a node with normal communication in another edge cluster (such as the second edge cluster).
[0018] In a possible implementation of the first aspect, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud indicates an abnormality, migrating the first service in the node where the communication status indicates an abnormality in the first edge cluster to the node where the communication status indicates a normality in the second edge cluster, and the second edge node and the first edge node are located in the same local area network.
[0019] In this case, if the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud is abnormal, it indicates that the majority of nodes in the first edge cluster are experiencing abnormal communication, and the communication between the master node and the cloud is abnormal. In this case, the first service in at least the majority of nodes in the first edge cluster is abnormal, including the master node. Therefore, by migrating the first service in the node with abnormal communication to a node in another edge cluster with normal communication, the normal operation of the service in the first edge cluster is guaranteed.
[0020] In a possible implementation of the first aspect, managing the first service based on the communication status of the first edge cluster and the communication status of the cloud also includes: triggering the re-election of the master node; and when the communication status of the cloud determined each time the master node is re-elected indicates an abnormality, migrating the first service in the node in the first edge cluster where the communication status indicates an abnormality is to the node in the second edge cluster where the communication status indicates a normal state.
[0021] In this way, the current master node is replaced by triggering a reelection. If the cloud communication status determined each time a master node is elected is abnormal, it indicates that the master nodes in the first edge cluster are unable to properly provide the first service. Based on this, the first service in the node with abnormal communication status in the first edge cluster is migrated to a node in the second edge cluster with normal communication status, ensuring the normal operation of the service in the first edge cluster.
[0022] In a possible implementation of the first aspect, the first service is managed according to the communication status of the first edge cluster and the communication status of the cloud, including: when the communication status of the cloud determined each time the master node is re-elected indicates that it is normal, the first edge cluster does not migrate the first service, and migration is used to indicate that the first service is migrated from one or more nodes to other nodes.
[0023] In this way, if the cloud communication status determined each time a master node is elected indicates normal, it indicates that at least one master node in the first edge cluster is able to normally provide the first service. Based on this, it is determined that the first edge cluster can normally provide the first service, and there is no need to migrate the first service in the first edge cluster, thereby reducing the number of service migrations and optimizing resource utilization.
[0024] In a possible implementation of the first aspect, determining the communication status of the first edge cluster includes: when the communication status indication of a first preset number of slave nodes within the first edge cluster is abnormal, determining that the communication status indication of the first edge cluster is abnormal, and the communication status of the slave node is used to indicate whether the slave node can communicate normally with the master node; when the communication status indication of a first preset number of slave nodes within the first edge cluster is normal, determining that the communication status indication of the first edge cluster is normal.
[0025] In this way, the distributed system architecture of the edge cluster (such as the first edge cluster) provides a certain degree of fault tolerance. This is based on the edge cluster's automatic recovery mechanism. When the communication status of a few nodes indicates an abnormality, the cluster can still communicate normally with the outside world through self-repair and adjustment, ensuring that the services provided by the edge cluster are not affected and maintaining service continuity and stability. However, when the communication status of the majority of nodes is abnormal, the edge cluster's communication status is determined to be normal. In the case of abnormal communication of the majority of nodes, it means that the edge cluster cannot communicate normally.
[0026] In a possible implementation of the first aspect, determining the communication status of the first edge cluster further includes: determining whether a heartbeat packet of a slave node is received within a preset time period; if the heartbeat packet of the slave node is received, determining that the communication status indication of the slave node is normal; if the heartbeat packet of the slave node is not received, determining that the communication status indication of the slave node is abnormal.
[0027] In this way, within each edge cluster (such as the first edge cluster), each node will send heartbeat packets to other nodes in the edge cluster at a preset time interval to notify other nodes of its communication status. Here, the master node, as the packet receiver, can determine the communication status of the slave node by determining whether it can receive the heartbeat packet from the slave node.
[0028] In one possible implementation of the first aspect, determining the cloud communication status includes: determining that the cloud communication status indicates normal when communication with the cloud is established; and determining that the cloud communication status indicates abnormal when communication with the cloud is disconnected. In this manner, determining the cloud communication status can reflect whether the master node can communicate normally with the cloud, thereby accurately reflecting the master node's communication status.
[0029] In one possible implementation of the first aspect, the method further includes synchronizing service information to each node in the first edge cluster so that the service information stored by each node is consistent, the service information being used to indicate whether each node in the first edge cluster can normally provide the first service. In this manner, the master node obtains information about the slave nodes and whether the master node itself can normally provide the first service to form service information, and synchronizes the service information to the slave nodes, thereby ensuring that the service information stored by each node is consistent.
[0030] In a possible implementation of the first aspect, the heartbeat packets received and sent by the master node are stored in a data structure ebpf map in the Berkeley Packet Filter EBPF program of the master node, and the heartbeat packets received and sent by the master node represent service information; the service information is synchronized to each node in the first edge cluster so that the service information stored by each node is consistent, including: storing the heartbeat packets in the ebpf map in the memory through the distributed consensus algorithm raft protocol, so that each node obtains the service information by reading the heartbeat packets received and sent by the master node from the memory. In this way, by storing the heartbeat packets in the ebpf map in the memory through the raft protocol, each node can obtain the service information by reading the heartbeat packets received and sent by the master node from the memory, thereby enabling each node to obtain consistent service information.
[0031] In a possible implementation manner of the first aspect, the method further includes: when the communication status indication of the cloud is normal, sending service information to the cloud.
[0032] In this way, when the communication status indication of the cloud is normal, the master node can act as the representative of the first edge cluster to interact with the cloud, thereby sending service information to the cloud, so that the cloud can obtain whether each node in the first edge cluster can provide the first service normally. There is no need to rely on the communication between the cloud and each node. The cloud only needs to communicate with the master node, which reduces the misjudgment of whether each node can provide the first service normally due to unstable communication between the cloud and each node, and improves the accuracy of the cloud in obtaining the service status of each node (that is, whether the first service can be provided normally).
[0033] In a possible implementation of the first aspect, the method further includes: when the communication status of the slave node indicates normal, determining whether the slave node can normally provide the first service based on a heartbeat packet of the slave node, the heartbeat packet indicating whether the node can normally provide the first service; and when the communication status of the slave node indicates abnormal, determining that the slave node cannot normally provide the first service. In this way, whether the slave node can normally provide the first service can be determined.
[0034] In a second aspect, an embodiment of the present application provides a management device for an edge cluster, which is applied to a master node in a first edge cluster, the first edge cluster including a master node and a slave node, both of which include a first service; the device includes: a first determination module for determining the communication status of the first edge cluster, the communication status of the first edge cluster being used to indicate whether the master node and the slave node can communicate normally; a second determination module for determining the communication status of the cloud, the communication status of the cloud being used to indicate whether the master node and the cloud can communicate normally; a management module for managing the first service based on the communication status of the first edge cluster and the communication status of the cloud.
[0035] In a third aspect, a computing device is provided, comprising: a processor, and a memory for storing processor-executable instructions; when the processor is configured to execute instructions, the computing device implements the method executed by the master node as described above.
[0036] In a possible implementation manner of the third aspect, the computing device is configured as a master node.
[0037] In a fourth aspect, an edge cluster is provided, comprising a master node and slave nodes, each of which includes a first service. The master node is configured to determine the communication status of the first edge cluster, where the communication status of the first edge cluster indicates whether the master node and the slave nodes can communicate normally; to determine the communication status of the cloud, where the communication status of the cloud indicates whether the master node and the cloud can communicate normally; and to manage the first service based on the communication status of the first edge cluster and the communication status of the cloud.
[0038] In a fifth aspect, a storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a computing device, the computing device implements the method as described above.
[0039] Among them, the technical effects brought about by any implementation method in the second to fifth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here.
[0040] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic diagram of the structure of a first edge cluster provided in an embodiment of the present application;
[0042] Figure 2 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0043] Figure 3 A schematic diagram of a process flow of an edge cluster management method provided in an embodiment of the present application;
[0044] Figure 4 A schematic diagram of a process for a candidate node to run for a master node according to an embodiment of the present application;
[0045] Figure 5 A flowchart of another edge cluster management method provided in an embodiment of the present application;
[0046] Figure 6 A schematic diagram of the structure of an edge cluster management device provided in an embodiment of the present application;
[0047] Figure 7A schematic diagram of another computing device structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0049] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.
[0050] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0051] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0052] The following is an illustrative introduction to the application scenarios of the embodiments of the present application.
[0053] In edge computing scenarios, the cloud provides computing resources, data processing, and storage services to the edge, while also managing and controlling the edge. Edge nodes utilize these resources to perform computing and storage close to end devices, improving response speed and reducing network bandwidth requirements. At the edge, multiple nodes providing services can form an edge cluster, working together to provide high availability and scalability. To ensure the stable operation of edge cluster services, edge cluster management is necessary.
[0054] In view of this, the present application provides a method for managing an edge cluster. On the one hand, the communication status of the first edge cluster is determined. On the other hand, the communication status of the cloud is determined. In this way, the first edge cluster is managed based on the communication status of the first edge cluster and the communication status of the cloud. Since the communication status of the first edge cluster is self-checked at the edge, it does not need to rely on communication with the cloud, thus avoiding misjudgment of the status caused by unstable communication between the cloud and the first edge cluster alone, making the communication status detection of the first edge cluster more accurate, thereby more reasonably managing the first edge cluster, reducing the number of service migrations, and rationally utilizing resources.
[0055] The following is an exemplary introduction to the system architecture of the embodiment of the present application.
[0056] like Figure 1 As shown, an embodiment of the present application provides a first edge cluster. The first edge cluster includes multiple edge nodes (hereinafter referred to as nodes), the multiple nodes include master nodes and slave nodes, and the master nodes and the slave nodes both include a first service.
[0057] in:
[0058] The master node is configured to determine the communication status of the first edge cluster and the communication status of the cloud, and to manage the first service based on the communication status of the first edge cluster and the communication status of the cloud.
[0059] The communication status of the first edge cluster is used to indicate whether the master node and the slave node can communicate normally. The communication status of the cloud is used to indicate whether the master node and the cloud can communicate normally.
[0060] For example, the number of master nodes is one, such as Figure 1 Node 1 is shown.
[0061] Exemplarily, the number of slave nodes is at least one, such as Figure 1 Node 2 and Node 3 are shown.
[0062] In some embodiments, there are multiple edge clusters, each of which includes multiple nodes (including a master node and at least one slave node) configured to support the same service (eg, the first service).
[0063] Exemplarily, edge clusters are divided according to the services included in the nodes, such as dividing nodes including the same service (such as the first service) into the same edge cluster. For ease of distinction, different edge clusters are referred to as the first edge cluster and the second edge cluster. The first edge cluster and the second edge cluster are located in the same local area network. It is understandable that the same node can include multiple services (such as the first service and the second service), that is, different edge clusters can have the same node.
[0064] Each node has both user space and kernel space. Kernel space is the operating system kernel's runtime space, while user space is the runtime space for user programs. User space and kernel space are isolated to prevent mutual interference.
[0065] In some embodiments, each node runs a monitoring program, which is used to obtain heartbeat packets sent and received by each node.
[0066] Among them, the heartbeat packet is a signal sent periodically. For the same node, it can be both the sender and the receiver.
[0067] In each edge cluster (such as the first edge cluster), each node will send heartbeat packets to other nodes in the edge cluster except for the node itself according to a preset time period. The purpose is to notify other nodes of the status of the node, such as communication status, service status, etc., wherein the service status refers to whether the node can provide the first service normally. If it can, the service status of the node indicates normal; if not, the service status of the node indicates abnormal. Specifically, for any node, the communication status of other nodes can be known by whether the heartbeat packets sent by other nodes can be received within the preset time period; based on the communication status and heartbeat packets of other nodes, the service status of other nodes can be further determined.
[0068] Exemplary monitoring programs include Extended Berkeley Packet Filter (EBPF) programs. EBPF is a kernel technology that allows programs to run in kernel space and obtain packet information directly from the network stack, thereby enabling efficient monitoring of network, storage, system calls, etc.
[0069] In one implementation, an EBPF program is mounted on each node's data transmission and / or reception path to monitor the heartbeat packets sent and received by each node. The EBPF program includes an eBPF map, a key data structure used to transfer data between the EBPF program and user space. It provides a way to share data between kernel space and user space, allowing users to read and write data in the EBPF program and also access and manipulate this data in user space. The eBPF map is used to store the heartbeat packets sent and received by the node.
[0070] Exemplarily, the node monitors whether it receives heartbeat packets from other nodes within a preset period of time through the EBPF program deployed on the node. For this node, the master node and the other nodes are slave nodes.
[0071] Take the example of a master node monitoring whether a heartbeat packet from a slave node is received within a preset period of time through an EBPF program deployed on the master node: the EBPF program of the master node is mounted on the data receiving path of the master node and captures the heartbeat packet sent by the slave node. If the heartbeat packet of the slave node is captured, it is determined that the heartbeat packet from the slave node is received; if the heartbeat packet from the slave node is not captured, it is determined that the heartbeat packet from the slave node is not received. If the heartbeat packet sent by the slave node is captured, the frequency of the heartbeat packet sent by the slave node and the heartbeat packet are saved in the ebpf map. Each time the capture is performed, the corresponding information is saved in the ebpf map. That is, the EBPF program on the data receiving path of the master node monitors the heartbeat packets and the sending frequency sent by the slave node in real time, and saves them in the ebpf map in a timely manner. The saving here neither overwrites the original information in the ebpf map (i.e., the heartbeat packets sent by the slave node and the sending frequency stored before) nor replaces the original information, but adds the heartbeat packets and the sending frequency sent by the slave node. In this way, by continuously recording the heartbeat packets sent by the slave node and the sending frequency, it is possible to determine in real time whether the master node has received the heartbeat packets of the slave node, and then obtain the communication status of the slave node.
[0072] In some embodiments, the master node is determined by election. In the first edge cluster, if there is no master node (perhaps because the master node has resigned or the election is being held for the first time), all nodes in the first edge cluster are candidate nodes. During a given period, only one candidate node is in the election. The candidate node in the election state can be understood as the current candidate node.
[0073] Exemplarily, when the current candidate node is running for election, whether the current candidate node is successful in the election is determined based on the number of nodes that are in normal communication with the current candidate node.
[0074] Exemplarily, the current candidate node monitors whether heartbeat packets from other nodes are received within a preset period of time through an EBPF program mounted on the data receiving path of the current candidate node, thereby determining the number of nodes that normally communicate with the current candidate node.
[0075] In some embodiments, each node deploys a gateway pod to perform network inbound and outbound forwarding and other functions. A pod is the smallest unit for running and deploying applications or services in a cluster and is an encapsulation of one or more containers.
[0076] For example, a gateway pod is deployed on each node using a daemonset to ensure that each node has an available gateway. A daemonset is a controller object that enables each node to run one or more pod replicas.
[0077] In one implementation, the monitoring program running on each node is deployed in the corresponding gateway pod. That is, the gateway pod provides an operating environment for the monitoring program on each node.
[0078] In one implementation, each node deploys a service pod. A service pod includes one or more user containers (UserContainers) for providing a service (e.g., a first service). If multiple service pods are deployed on the same node, it indicates that the node can provide multiple services (e.g., a first service, a second service, etc.).
[0079] Network data packets (including heartbeat packets) generated by user pods need to be sent to other nodes via the node's data transmission path. Data packets sent by other nodes to user pods are received via the node's data reception path. Gateway pods, as the entry and exit points for network traffic, also rely on the node's data reception and transmission paths to forward data packets, enabling communication between user pods and external networks. Based on this, the EBPF program captures heartbeat packets sent by other nodes to user pods and gateway pods, as well as heartbeat packets sent by the node itself, along the node's data reception and transmission paths.
[0080] In some embodiments, the master node is also used to: when the communication status indication of the slave node is normal, determine whether the slave node can provide the first service normally based on the heartbeat packet of the slave node; or, when the communication status indication of the slave node is abnormal, determine that the slave node cannot provide the first service normally.
[0081] The heartbeat packet represents whether the node can provide a service (such as the first service) normally. It can be understood that when the node sends and receives a heartbeat packet, the heartbeat packet includes the heartbeat packet information of the contractor (i.e., the node that sends the heartbeat packet) (including the service status of whether the contractor can provide the first service normally, etc.). Based on this, the heartbeat packet represents whether the node that sends the packet can provide the first service normally. For example, for the heartbeat packet sent from the slave node, the heartbeat packet represents whether the slave node can provide the first service normally. In this way, when the master node receives the heartbeat packet from the slave node, it can obtain the service status of whether the slave node can provide the first service normally. For the heartbeat packet sent from the master node, the heartbeat packet represents whether the master node can provide the service status of the first service normally.
[0082] Specifically, the master node is used to determine whether the slave node can normally provide the first service according to the heartbeat packet of the slave node stored in the ebpf map corresponding to the master node.
[0083] In some embodiments, the master node is further used to determine whether the master node can normally provide the first service.
[0084] Specifically, the master node's EBPF program is mounted on the master node's data transmission path to capture heartbeat packets sent by the master node. For example, when capturing heartbeat packets sent by the master node, the frequency and frequency of heartbeat packets sent by the master node are saved to the eBPF map. It is understood that the EBPF program on the master node's data transmission path monitors the heartbeat packets and frequency of heartbeat packets sent by the master node in real time and saves them in the eBPF map in a timely manner. In this way, by recording the heartbeat packet information in the heartbeat packets sent by the master node, it is determined whether the master node can normally provide the first service.
[0085] In some embodiments, the master node is further configured to synchronize service information with each node in the first edge cluster so that the service information stored by each node is consistent. The service information is used to indicate whether each node in the first edge cluster can normally provide the first service. In this way, each node in the first edge cluster reaches a consensus on the service information.
[0086] In one implementation, the master node synchronizes service information with each node based on the distributed consensus algorithm, the Raft protocol, ensuring that each node maintains consistent service information. Raft is a consensus algorithm used to ensure consistency in service information across nodes. Thus, through the Raft protocol, consistent service information is maintained across nodes.
[0087] Heartbeat packets sent and received by the master node are stored in the eBPF map in the master node's EBPF program. These heartbeat packets represent service information based on the heartbeat packets received by the master node from slave nodes and the heartbeat packets sent by the master node. Therefore, the heartbeat packets sent and received by the master node encompass the heartbeat packets of both slave nodes and the master node, indicating the service status of both nodes (i.e., whether the primary service can be provided normally), thereby representing service information.
[0088] Exemplarily, the master node is used to synchronize service information to each node in the first edge cluster so that the service information stored by each node is consistent, specifically including: the master node is used to store the heartbeat packet in the ebpf map into the memory through the raft protocol, so that each node obtains the service information by reading the heartbeat packet sent and received by the master node from the memory.
[0089] In some embodiments, the master node is further configured to send service information to the cloud when the communication status of the cloud indicates normal.
[0090] In the embodiment of the present application, the edge node may be a computing device, wherein the computing device may be a network device, an electronic device (such as a smart terminal), etc.
[0091] Optionally, the smart terminal may include a mobile phone, a tablet computer, a handheld computer, a personal computer (PC), a cellular phone, a personal digital assistant (PDA), a wearable device (such as a smart watch, a smart bracelet, etc.), a smart home device (such as a television, etc.), a car machine (such as a car computer, etc.), a smart screen, a game console, a headset, an AI speaker, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a laptop computer, a netbook, a desktop computer or an all-in-one computer, etc.
[0092] Optionally, the network device may include a server, etc. The server may be a physical or logical server, or may be two or more physical or logical servers that share different responsibilities and cooperate with each other to implement various functions of the server.
[0093] Illustratively, the server may be a blade server, a high-density server, a rack server, a tower server, or the like.
[0094] In the embodiment of the present application, the computing device is configured as a master node. The hardware part of the computing device includes a processor, a basic input and output system (BIOS) chip, an out-of-band controller, and a memory, and the software part mainly includes the BIOS, an out-of-band management module, and an operating system (OS). Figure 2 As shown, an embodiment of the present application provides a structural schematic diagram of a computing device.
[0095] The processor may include a central processing unit (CPU), which includes one or more CPU cores. The CPU's data processing operations are all performed by the CPU cores. The more CPU cores a CPU includes, the faster it processes data.
[0096] The BIOS chip is a chip installed on the motherboard that initializes and detects various hardware components during the computer's startup process. The BIOS chip includes a flash memory area.
[0097] The out-of-band management module is located in the out-of-band controller, and the operating system is located in the processor.
[0098] An out-of-band management module can be a management unit for non-business modules. For example, an out-of-band management module can remotely maintain and manage a computing device through a dedicated data channel. This out-of-band management module is completely independent of the computing device's operating system and can communicate with the BIOS and operating system through the computing device's out-of-band management interface.
[0099] Exemplarily, the out-of-band management module may include a management unit for the computing device's operating status, a management system in a management chip, a baseboard management controller (BMC) for the computing device, a system management module (SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above description is merely exemplary.
[0100] The operating system (OS) is a computer program that manages and controls the hardware and software resources of a computing device. Any other software must be supported by the OS to run. After a computing device is powered on, the BIOS first performs a series of operations, including self-tests and initialization, and then boots the OS, allowing the user to use the computing device normally.
[0101] BIOS is a set of programs embedded in the BIOS chip on the motherboard of a computing device. The main function of BIOS is to provide the lowest-level and most direct hardware settings and control for the computing device.
[0102] Memory, also known as internal storage or main memory, is installed in memory slots on the motherboard of a computing device.
[0103] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0104] For ease of understanding, the following describes an exemplary method for managing an edge cluster provided in an embodiment of the present application in combination with the above-mentioned first edge cluster and the accompanying drawings.
[0105] like Figure 3 As shown, an embodiment of the present application provides a management method for an edge cluster, which is applied to the master node of a first edge cluster and executed by the master node in the first edge cluster. The method includes:
[0106] S301: Determine the communication status of a first edge cluster.
[0107] The edge cluster includes multiple nodes, including a master node and at least one slave node, and both the master node and the slave node include a first service.
[0108] In some embodiments, there are multiple edge clusters, each edge cluster is used to support the same service (such as the first service). Exemplarily, the edge clusters are divided according to the services included in the nodes, such as dividing the nodes including the same service (such as the first service) into the same edge cluster. For ease of distinction, different edge clusters are referred to as the first edge cluster and the second edge cluster. The first edge cluster and the second edge cluster are located in the same local area network. It can be understood that the same node can include multiple services (such as the first service and the second service), that is, different edge clusters can have the same node.
[0109] The communication status is used to indicate whether the communication is normal or abnormal. In the embodiment of the present application, the communication status is divided into: the communication status of the edge cluster, the communication status of the cloud, the communication status of the master node, and the communication status of the slave node.
[0110] The communication status of the edge cluster is used to indicate whether the master node and the slave node can communicate normally. In one possible implementation, the communication status of the edge cluster can be determined based on whether the master node and the slave node in the edge cluster can communicate normally.
[0111] The cloud communication status indicates whether the master node and the cloud can communicate normally. For example, if the master node and the cloud can communicate normally, the cloud communication status indicates normal. If the master node and the cloud cannot communicate normally, the cloud communication status indicates abnormal.
[0112] In one possible implementation, the communication status of the cloud can be determined based on whether the communication between the master node and the cloud is connected or disconnected.
[0113] The communication status of the master node is used to indicate whether the master node can communicate normally with the cloud during the term. For example, during the term, if the master node can communicate normally with the cloud, the communication status of the master node indicates normal; if the master node cannot communicate normally with the cloud, the communication status of the master node indicates abnormal.
[0114] A term refers to the period of time a node serves as a master. A term begins when a node is elected as a master and ends when a master reelection is triggered. As you can see, since a node's communication status changes dynamically, if a master no longer meets the requirements for normal communication with the cloud during a term, a master reelection can be triggered to replace the master.
[0115] The communication status of a slave node indicates whether the slave node can communicate normally with the master node. For example, in the first edge cluster, if the slave node can communicate normally with the master node, the communication status of the slave node indicates normal; if the slave node cannot communicate normally with the master node, the communication status of the slave node indicates abnormal.
[0116] In one possible implementation, the communication status of the slave node can be determined based on whether the master node can receive the heartbeat packet of the slave node. The heartbeat packet is a signal sent periodically, and for the same node, it can be both the sender and the receiver.
[0117] In some embodiments, determining the communication status of the first edge cluster includes: determining that the communication status of the first edge cluster indicates abnormality when the communication status of a first preset number of slave nodes in the first edge cluster indicates abnormality; and determining that the communication status of the first edge cluster indicates normal when the communication status of a first preset number of slave nodes in the first edge cluster indicates normality.
[0118] The fact that the communication status of a first preset number of slave nodes indicates normal indicates that the communication status of the majority of nodes in the first edge cluster indicates normal. This is because if a slave node can communicate normally with the master node, it implicitly implies that the master node can also communicate normally with the slave node. Therefore, the number of nodes in the first edge cluster that are communicating normally is actually the sum of the number of slave nodes indicating normal communication status and the number of master nodes.
[0119] In an embodiment of the present application, the distributed system architecture of the edge cluster (such as the first edge cluster) enables the edge cluster to have a certain fault tolerance. This is based on the automatic recovery mechanism of the edge cluster. When the communication status indication of a few nodes is abnormal, it can still communicate normally with the outside world through self-repair and adjustment, ensuring that the services provided by the edge cluster are not affected and maintaining the continuity and stability of the service. However, when the communication of most nodes is abnormal, it is determined that the communication status indication of the edge cluster is normal. In the case of abnormal communication of most nodes, it means that the edge cluster cannot communicate normally.
[0120] Exemplarily, the first preset number is greater than or equal to 1 / 2 of the number of the multiple nodes in the first edge cluster, and the first preset number is an integer.
[0121] For example, if the number of the multiple nodes in the first edge cluster is 6, the first preset number is greater than or equal to 3. In the first edge cluster, if the number of slave nodes indicating a normal communication status is greater than or equal to 3 (such as 3, 4, 5, or 6), the communication status of the first edge cluster is normal; if the number of slave nodes indicating an abnormal communication status is greater than or equal to 3, the communication status of the first edge cluster is abnormal.
[0122] In some other embodiments, determining the communication status of the first edge cluster further includes determining the communication status of slave nodes within the first edge cluster. In this way, determining the communication status of the first edge cluster based on the communication status of the slave nodes within the first edge cluster is more accurate and efficient.
[0123] In one implementation, determining the communication status of a slave node in the first edge cluster includes: determining whether a heartbeat packet of the slave node is received within a preset time period; if the heartbeat packet of the slave node is received, determining that the communication status of the slave node is normal; and if the heartbeat packet of the slave node is not received, determining that the communication status of the slave node is abnormal.
[0124] In an embodiment of the present application, within each edge cluster (e.g., the first edge cluster), each node sends heartbeat packets to other nodes in the edge cluster (excluding the node itself) at predetermined intervals to notify the other nodes of the node's communication status. Here, the master node, as the packet receiver, can determine the communication status of a slave node by determining whether it can receive the slave node's heartbeat packet.
[0125] For example, the preset time period can be set to 1s, 10s, 30s or 60s, and can be set according to the actual needs of the user.
[0126] S302: Determine the communication status of the cloud.
[0127] In some embodiments, determining the communication status of the cloud includes: determining that the communication status of the cloud indicates normal when the communication with the cloud is connected, and determining that the communication status of the cloud indicates abnormal when the communication with the cloud is disconnected.
[0128] In the embodiment of the present application, by determining the communication status of the cloud, it can reflect whether the master node and the cloud can communicate normally, thereby accurately reflecting the communication status of the master node.
[0129] S303: Manage the first service according to the communication status of the first edge cluster and the communication status of the cloud.
[0130] In some embodiments, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates normal, the first edge cluster does not migrate the first service.
[0131] Migrate is used to indicate migrating the first service from one or more nodes to other nodes.
[0132] In an embodiment of the present application, when the communication status indication of the first edge cluster is normal and the communication status indication of the cloud is normal, it indicates that most nodes in the first edge cluster are communicating normally, and the communication between the master node and the cloud is also normal. At this time, the first edge cluster can provide the first service normally. Therefore, there is no need to migrate the first service in the first edge cluster, which can ensure the stable operation of the service of the first edge cluster, reduce the number of service migrations, and rationally utilize resources.
[0133] In other embodiments, managing the first edge service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, the first edge cluster does not migrate the first service.
[0134] In an embodiment of the present application, when the communication status indication of the first edge cluster is normal and the communication status indication of the cloud is abnormal, it indicates that although the communication between the master node and the cloud is abnormal, the communication of most nodes in the first edge cluster is normal. At this time, the first edge cluster can provide the first service normally. Therefore, there is no need to migrate the first service in the first edge cluster, which can ensure the stable operation of the service of the first edge cluster, reduce the number of service migrations, and reasonably utilize resources.
[0135] In other embodiments, managing the first edge service according to the communication status of the first edge cluster and the communication status of the cloud includes: triggering re-election of the master node when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal.
[0136] In an embodiment of the present application, when the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, it indicates that although most nodes in the first edge cluster are communicating normally, the communication between the master node and the cloud is abnormal. Therefore, the re-election of the master node can be triggered and the master node can be replaced.
[0137] In other embodiments, managing the first edge service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normal, the first edge cluster does not migrate the first service.
[0138] In an embodiment of the present application, when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normality, this indicates that although the majority of nodes in the first edge cluster have abnormal communication, that is, at least the first service in the majority of nodes in the first edge cluster is abnormal, but since the communication between the master node and the cloud is normal, it is further explained that the node with the abnormal communication status is a slave node. In this case, at least the master node can normally provide the first service, so it is determined that the first edge cluster can normally provide the first service, and there is no need to migrate the first service in the first edge cluster, thereby reducing the number of service migrations and rationally utilizing resources.
[0139] In other embodiments, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normality, migrating the first service in the node where the communication status in the first edge cluster indicates abnormality to the node where the communication status in the second edge cluster indicates normality.
[0140] In an embodiment of the present application, when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates normality, it indicates that the majority of nodes in the first edge cluster have abnormal communication, that is, at least the first service in the majority of nodes in the first edge cluster is abnormal. In this case, although the communication between the master node and the cloud is normal, the stability of the service operation can be improved by migrating the first service in the node with abnormal communication to a node with normal communication in another edge cluster (such as the second edge cluster).
[0141] For example, before migration, the nodes in the first edge cluster include nodes 1-5, which are used to jointly provide service 1, wherein nodes 1-3 have communication anomalies. The nodes in the second edge cluster include nodes 6-12, which are used to jointly provide service 2, wherein nodes 6-12 have normal communication. Then, the services in nodes 1-3 can be migrated to any three nodes in the second edge cluster (such as nodes 8-10). After migration, the edge clusters are re-divided, and one edge cluster includes: nodes 4, 5, 8, 9, 10, which are used to jointly provide service 1. The other edge cluster includes: nodes 6-12, which are used to jointly provide service 2.
[0142] In other embodiments, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates abnormality, migrating the first service in the node where the communication status in the first edge cluster indicates abnormality to the node where the communication status in the second edge cluster indicates normality.
[0143] In an embodiment of the present application, when the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud is abnormal, it indicates that the majority of nodes in the first edge cluster are experiencing communication abnormalities, and that the communication between the master node and the cloud is abnormal. In this case, the first service in at least the majority of nodes in the first edge cluster is abnormal, including the master node. Therefore, by migrating the first service in the node with abnormal communication to a node in another edge cluster with normal communication, the normal operation of the service in the first edge cluster is guaranteed.
[0144] In one implementation, managing the first service based on the communication status of the first edge cluster and the communication status of the cloud also includes: triggering the re-election of the master node; when the communication status of the cloud determined each time the master node is re-elected indicates an abnormality, migrating the first service in the node whose communication status in the first edge cluster indicates an abnormality to the node whose communication status in the second edge cluster indicates a normal state.
[0145] In an embodiment of the present application, the current master node is replaced by triggering a re-election. If the cloud communication status determined each time the master node is elected is abnormal, it indicates that the master nodes in the first edge cluster are unable to provide the first service normally. Based on this, the first service in the node with abnormal communication status indication in the first edge cluster is migrated to the node with normal communication status indication in the second edge cluster, thereby ensuring the normal operation of the service in the first edge cluster.
[0146] In another implementation, managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: when the communication status of the cloud determined each time the master node is re-elected indicates that it is normal, the first edge cluster does not migrate the first service.
[0147] In this embodiment of the present application, if the cloud communication status determined each time a master node is elected indicates normal, then this indicates that at least one master node within the first edge cluster is capable of normally providing the first service. Based on this, it is determined that the first edge cluster is capable of normally providing the first service, and migration of the first service within the first edge cluster is unnecessary, thereby reducing the number of service migrations and optimizing resource utilization.
[0148] In this way, by comprehensively considering the communication status of the first edge cluster and the communication status of the cloud, the first edge cluster can be managed in multiple ways, which can make it possible to manage the first edge cluster more flexibly and more reasonably, reduce the number of service migrations, and make rational use of resources.
[0149] In some embodiments, triggering the re-election of the master node includes: notifying the slave node to cancel the master node to trigger the re-election of the master node.
[0150] In an embodiment of the present application, the condition for re-election of the master node is that the master node notifies the slave node to cancel the master node (that is, the current master node resigns), thereby triggering the re-election of the master node.
[0151] In some embodiments, electing a master node includes: for each edge cluster, a candidate node running for master node.
[0152] In the embodiments of the present application, for each edge cluster, all nodes except the master node are slave nodes during the master node's term. After the master node resigns, it becomes a candidate node. When a slave node learns that the master node has resigned or that there is no master node in the current cluster, each slave node also becomes a candidate node. The candidate nodes compete to determine a new master node.
[0153] In one implementation, the candidate node electing as the master node includes: the current candidate node determining the number of nodes that can normally communicate with the current candidate node within a preset time period. If the number of nodes that can normally communicate with the current candidate node is greater than or equal to a first preset number, the current candidate node is determined to be the master node in the first edge cluster.
[0154] In an embodiment of the present application, the current candidate node is a node among the candidate nodes, and the number of votes obtained by the current candidate node is determined by determining the number of nodes that communicate normally with the current candidate node. That is, if a node communicates normally with the current candidate node, it can be regarded as the current candidate node obtaining the votes of the node, that is, the node supports the current candidate node as the master node. When the number of nodes that communicate normally with the current candidate node is at least the first preset number, the current candidate node will also vote for itself by default. Based on this, it can be seen that the current candidate node has obtained the majority of votes and can serve as the representative of each node in the first edge cluster, that is, as the master node. In this way, the selected master node has good communication and coordination capabilities within the edge cluster, and can serve as the representative of the edge cluster, thereby utilizing the master node to aggregate and distribute data and synchronize information.
[0155] Exemplarily, the current candidate node may be the i-th candidate node among the plurality of candidate nodes, where i is greater than or equal to 1 and less than or equal to N, N is the number of the plurality of nodes in the first edge cluster, and i is an integer.
[0156] For example, the preset duration corresponds to the current candidate node, that is, different current candidate nodes may have different corresponding preset durations, which can be set according to the actual needs of the user.
[0157] In another implementation, within a preset time period, if the number of nodes that communicate normally with the current candidate node is less than a first preset number, the i+1th candidate node among the multiple candidate nodes is used as the current candidate node, and the number of nodes that communicate normally with the current candidate node is obtained until the current candidate node serves as the main node.
[0158] For example, each candidate node is randomly assigned a time slice of varying length. This time slice is used to determine the preset duration and election order for each candidate node. Each time slice can range from 60ms to 600ms, and the start time of the time slice can be either the current time or a preset time, depending on the user's needs.
[0159] For example, the current time is used as the starting time of the time slice corresponding to each candidate node, where the first time slice set for candidate node 1 is 60ms, the second time slice set for candidate node 2 is 120ms, the third time slice set for candidate node 3 is 180ms, etc. The election order of the candidate nodes (that is, the order of the current candidate nodes) is determined according to the end time of the time slice consumption, and the preset duration corresponding to each candidate node is determined.
[0160] Furthermore, taking the candidate nodes 1-3 shown above as an example, the first time slice is the shortest and is consumed first, followed by the second time slice and the third time slice, and the election order is: candidate nodes 1, 2, 3. Based on the election order, the preset duration corresponding to the current candidate node is the time period between the end time of the time slice corresponding to the current candidate node and the end time of the time slice corresponding to the next current candidate node. That is, the preset duration corresponding to candidate node 1 is the time period between the end time of the first time slice and the end time of the second time slice (hereinafter referred to as the first time period), and the preset duration corresponding to candidate node 2 is the time period between the end time of the second time slice and the end time of the third time slice (hereinafter referred to as the second time period).
[0161] Specifically, during a first time period, candidate node 1 is used as the current candidate node, and the number of nodes that can communicate normally with candidate node 1 is obtained. If the number of nodes that can communicate normally with candidate node 1 is greater than or equal to a first preset number, candidate node 1 is used as the master node. If the number of nodes that can communicate normally with candidate node 1 is less than the first preset number, during a second time period, candidate node 2 is used as the current candidate node, and the number of nodes that can communicate normally with candidate node 2 is obtained. If the number of nodes that can communicate normally with candidate node 2 is greater than or equal to the first preset number, candidate node 2 is used as the master node. If the number of nodes that can communicate normally with candidate node 2 is less than the first preset number, candidate node 3 is used as the current candidate node, and a determination is made as to whether candidate node 3 can serve as the master node.
[0162] In this way, according to the order of the current candidate nodes, each current candidate node determines in turn whether it can obtain the majority of votes, that is, whether it can serve as the master node, which is efficient and reasonable.
[0163] like Figure 4 As shown, an embodiment of the present application provides a schematic diagram of a process for a candidate node to run for a master node, which is executed by the current candidate node in the first edge cluster, including:
[0164] S401, obtaining the number of nodes that can communicate normally with the current candidate node.
[0165] S402: Determine whether the number of nodes that can communicate normally with the current candidate node is greater than or equal to a first preset number. If so, proceed to S403; if not, proceed to S404.
[0166] S403, serves as the master node.
[0167] S404, failed to elect the master node.
[0168] Based on this, when the i-th candidate node is the current candidate node and the current candidate node fails in the election, the i+1-th candidate node is replaced as the current candidate node, and the current candidate node continues to run for the master node according to the above process until the current candidate node succeeds in the election.
[0169] In this way, the votes each candidate node receives are determined by whether it can communicate with the node normally. The number of votes each candidate node receives is determined based on the number of nodes that can communicate normally with the node, allowing the candidate node to confirm its eligibility for the master node, improving the accuracy and rationality of the master node election.
[0170] Nodes in the edge cluster send heartbeat packets to each other. For any node (this node), it can send heartbeat packets to other nodes to inform them of its status, and it can also receive heartbeat packets sent by other nodes to obtain the status of other nodes.
[0171] The status here refers to the communication status on the one hand and the service status on the other.
[0172] In some embodiments, the master node (or current candidate node) determines the communication status of the slave node (or other node) by whether it can receive the heartbeat packet of the slave node (or other node). If so, it determines that the communication status indication of the slave node (or other node) is normal; if not, it determines that the communication status indication of the slave node (or other node) is abnormal.
[0173] In one implementation, the master node (or current candidate node) monitors whether a heartbeat packet from a slave node (or other node) is received within a preset period of time through an EBPF program deployed on the master node (or current candidate node).
[0174] For example, the master node monitors whether a heartbeat packet from a slave node is received within a preset period of time through an EBPF program deployed on the master node: the master node's EBPF program is mounted on the master node's data receiving path and captures the heartbeat packet sent by the slave node. If the heartbeat packet from the slave node is captured, it is determined that the heartbeat packet from the slave node has been received. If the heartbeat packet from the slave node is not captured, it is determined that the heartbeat packet from the slave node has not been received. If the heartbeat packet from the slave node is captured, the frequency of the heartbeat packet sent by the slave node and the heartbeat packet are saved in the eBPF map. Each time the heartbeat packet is captured, the corresponding information is saved in the eBPF map. In other words, the EBPF program on the master node's data receiving path monitors the heartbeat packets and the sending frequency sent by the slave node in real time and saves them in the eBPF map in a timely manner. The saving here neither overwrites the original information in the eBPF map (i.e., the heartbeat packets and the sending frequency sent by the slave node stored previously) nor replaces the original information. Instead, it adds the heartbeat packets and the sending frequency sent by the slave node. In this way, by continuously recording the heartbeat packets sent by the slave node and the sending frequency, it is possible to determine in real time whether the master node has received the heartbeat packets of the slave node, and then obtain the communication status of the slave node.
[0175] In other embodiments, the method further includes: determining whether the slave node can normally provide the first service.
[0176] In one implementation, if the communication status of the slave node indicates normal, it is determined whether the slave node can normally provide the first service according to the heartbeat packet of the slave node. If the communication status of the slave node indicates abnormal, it is determined that the slave node cannot normally provide the first service.
[0177] The heartbeat packet is used to indicate whether the node can normally provide the first service. It can be understood that the heartbeat packet indicates that the node sending the heartbeat packet can normally provide the first service.
[0178] Exemplarily, determining whether the slave node can provide the first service normally based on the heartbeat packet of the slave node includes: if the heartbeat packet indicates that the slave node can provide the first service normally, then the slave node can provide the first service normally; if the heartbeat packet indicates that the slave node cannot provide the first service normally, then the slave node cannot provide the first service normally.
[0179] Specifically, whether the slave node can normally provide the first service is determined through the heartbeat packet of the slave node stored in the ebpf map corresponding to the master node.
[0180] In this way, if the communication status indication of the slave node is normal, it is further determined based on the heartbeat packet whether the slave node can normally provide the first service, thereby more accurately determining the service status of the slave node.
[0181] In another implementation, the method further includes: determining whether the master node can normally provide the first service.
[0182] Exemplarily, the EBPF program of the master node is mounted on the data transmission path of the master node to capture the heartbeat packets sent by the master node. For example, when capturing the heartbeat packets sent by the master node, the frequency of the heartbeat packets sent by the master node and the heartbeat packets are saved to the eBPF map. It is understandable that the EBPF program on the data transmission path of the master node monitors the heartbeat packets and the frequency of the heartbeat packets sent by the master node in real time and saves them in the eBPF map in a timely manner. In this way, by recording the heartbeat packets sent by the master node (including the heartbeat packet information), it is determined whether the master node can normally provide the first service.
[0183] In this way, by determining the service status of the slave node and the service status of the master node, the service status of each node in the first edge cluster can be obtained, so that a clearer and more accurate judgment can be made on whether each node in the first edge cluster can provide the first service, thereby accurately obtaining service information.
[0184] Furthermore, if the heartbeat packet sent by the master node indicates that the master node can provide the first service normally, the master node can provide the first service normally; if the heartbeat packet sent by the master node indicates that the master node cannot provide the first service normally, the master node cannot provide the first service normally.
[0185] In some embodiments, the method further includes: synchronizing the service information to each node in the first edge cluster so that the service information stored by each node is consistent.
[0186] The service information is used to indicate whether each node in the first edge cluster can normally provide the first service.
[0187] In an embodiment of the present application, the service status of the slave node and the service status of the master node itself can be obtained based on the above-mentioned master node to form service information, and the service information can be synchronized to the slave node to synchronize the service information to each node in the first edge cluster, thereby making the service information stored by each node consistent.
[0188] In one implementation, the master node synchronizes service information to each node based on the Raft protocol, so that the service information stored by each node is consistent. In this way, the consistency of service information stored by each node is achieved through the Raft protocol.
[0189] Exemplarily, synchronizing the service information to each node in the first edge cluster so that the service information stored by each node is consistent includes: storing the heartbeat packet in the ebpf map into the memory through the raft protocol, so that each node obtains the service information by reading the heartbeat packet sent and received by the master node from the memory.
[0190] The heartbeat packets received and sent by the master node are stored in the eBPF map in the master node's Berkeley Packet Filter (EBPF) program. The heartbeat packets received and sent by the master node represent service information. In this way, based on the heartbeat packets received by the master node from the slave nodes and the heartbeat packets sent by the master node, the heartbeat packets received and sent by the master node cover the heartbeat packets of the slave nodes and the master node, thereby representing the service status of the slave nodes and the master node (i.e., whether the first service can be provided normally), and thus forming service information. By storing the heartbeat packets in the eBPF map in memory through the Raft protocol, each node can obtain service information by reading the heartbeat packets received and sent by the master node from memory, thereby ensuring that each node can obtain consistent service information.
[0191] Exemplarily, synchronizing service information to each node in the first edge cluster also includes: if the communication between the master node and the slave node is disconnected, then when the communication between the master node and the slave node is restored, the status information of each change after the communication is disconnected is synchronized to the slave node.
[0192] For example, at t0 on XX / XX / XX, the communication status of slave node 1 indicates an abnormality, meaning that communication with the master node is disconnected. At t1 on XX / XX / XX, the communication status of slave node 1 indicates a normal status, meaning that communication with the master node is restored. Between t0 and t1, the service status of slave node 1 was updated twice, corresponding to t0 on XX / XX / XX and t1 on XX / XX / XX, respectively. Therefore, the master node not only obtains the service status of slave node 1 at t1 on XX / XX / XX and synchronizes it to each node, but also obtains the service status of slave node 1 at t0 on XX / XX / XX and synchronizes it to each node. This further ensures that the status information stored by each node is consistent.
[0193] In this way, by synchronizing the service information to each node in the first edge cluster, a first edge cluster with data consistency is constructed, so that the service information stored by each node is consistent, and consensus is reached through the master node within the first edge cluster to improve the fault tolerance, availability and stability of the first edge cluster.
[0194] For example, when the service information is updated, the updated service information is synchronized to each node in the first edge cluster, thereby further achieving consistency in the service information stored by each node.
[0195] For example, at time t1, the service information includes: nodes 1-3 can normally provide the first service; nodes 4-7 cannot normally provide the first service. At this time, the master node synchronizes the service information to each node, so that the service information stored by each node includes: nodes 1-3 can normally provide the first service; nodes 4-7 cannot normally provide the first service.
[0196] For example, at time t2, the service information changes. The updated service information includes: Node 1-2 can provide the first service normally. Node 3-5 cannot provide the first service normally. Node 6-7 can provide the first service normally. At this time, the master node synchronizes the updated service information to each node, so that the status information stored by each node includes: Node 1-2 can provide the first service normally. Node 3-5 cannot provide the first service normally. Node 6-7 can provide the first service normally.
[0197] Exemplarily, the service information also includes a timestamp. A timestamp can be a time identifier with a specific format and standard. For example, a timestamp may include information such as year, month, day, hour, minute, second, millisecond, and microsecond. The timestamp is used to record the time when the node's service status changes. For example, at time t1 on XX / XX / XX, node 1 can normally provide the first service. At time t2 on XX / XX / XX, node 1 cannot normally provide the first service. Here, time t1 on XX / XX / XX and time t2 on XX / XX / XX are timestamps. In this way, the service information is more comprehensive and clear.
[0198] In some embodiments, the method further includes: sending service information to the cloud when the communication status indication of the cloud is normal.
[0199] In an embodiment of the present application, when the communication status indication of the cloud is normal, the master node can act as a representative of the first edge cluster to interact with the cloud, thereby sending service information to the cloud, so that the cloud can obtain the service status of each node in the first edge cluster that can normally provide the first service. There is no need to rely on the communication between the cloud and each node. The cloud only needs to communicate with the master node, reducing the misjudgment of the service status of each node due to unstable communication between the cloud and each node, and improving the accuracy of the cloud in obtaining the service status of each node.
[0200] In some other embodiments, the method further includes: obtaining a status acquisition request from the cloud, and sending service information to the cloud in response to the status acquisition request from the cloud.
[0201] The status acquisition request is used to request service information.
[0202] In an embodiment of the present application, (when the communication status indication of the cloud is normal) when a status acquisition request from the cloud is obtained, the master node sends service information to the cloud to report status information to the cloud at a specific time or in a specific scenario, thereby avoiding meaningless service information reporting by the master node and meeting the actual acquisition needs of the cloud, thereby making the service information obtained by the cloud more targeted and timely.
[0203] In other embodiments, the method further includes: when the service information is updated, sending the updated service information to the cloud.
[0204] In an embodiment of the present application, the service information reported by the master node to the cloud is the latest. For example, each time the service information is updated, the master node sends the updated service information to the cloud (when the communication status of the cloud indicates normal).
[0205] like Figure 5 As shown, an embodiment of the present application provides a flow chart of another edge cluster management method. The method is executed by a master node in a first edge cluster, and the method includes:
[0206] S501: Determine a communication status of a first edge cluster.
[0207] S502: Determine the communication status of the cloud.
[0208] S503: When the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates normal, the first edge cluster does not migrate the first service.
[0209] S504: When the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, the first edge cluster does not migrate the first service and / or triggers re-election of the master node.
[0210] S505: When the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud indicates a normality, the first edge cluster does not migrate the first service, or migrates the first service in the node where the communication status in the first edge cluster indicates an abnormality to the node where the communication status in the second edge cluster indicates a normality.
[0211] S506 , when the communication status of the first edge cluster indicates abnormality and the communication status of the cloud indicates abnormality, migrate the first service in the node in the first edge cluster where the communication status indicates abnormality to the node in the second edge cluster where the communication status indicates normality.
[0212] S507, triggering the re-election of the master node.
[0213] S508 , when the cloud communication status determined each time the master node is re-elected indicates an abnormality, migrate the first service in the node whose communication status in the first edge cluster indicates an abnormality to the node whose communication status in the second edge cluster indicates a normality.
[0214] S509: When the communication status of the cloud determined each time the master node is re-elected indicates that it is normal, the first edge cluster does not migrate the first service.
[0215] S510, sending service information to the cloud.
[0216] In the embodiment of the present application, the communication status of the first edge cluster and the cloud are used to flexibly and reasonably manage the first edge cluster. The master node interacts with the cloud and reports service information to the cloud, so that the cloud can obtain the service status of the first edge cluster and each node within it.
[0217] like Figure 6 As shown, an embodiment of the present application provides an edge cluster management device 200, which is applied to a master node in a first edge cluster. The first edge cluster includes a master node and a slave node, and both the master node and the slave node include a first service. The edge cluster management device 200 includes: a first determination module 21, a second determination module 22, and a management module 23. The first determination module 21 is used to determine the communication status of the first edge cluster, and the communication status of the first edge cluster is used to indicate whether the master node and the slave node can communicate normally. The second determination module 21 is used to determine the communication status of the cloud, and the communication status of the cloud is used to indicate whether the master node and the cloud can communicate normally. The management module 23 is used to manage the first service based on the communication status of the first edge cluster and the communication status of the cloud.
[0218] like Figure 7 As shown, an embodiment of the present application provides another computing device 500. The computing device 500 includes a processor 510 and a memory 520 for storing processor-executable instructions. When the processor 510 is configured to execute the instructions, the computing device 500 implements the method executed by the master node as described above.
[0219] In an embodiment of the present application, the computing device 500 is configured as the above-mentioned master node.
[0220] Figure 7 The computing device 500 shown is merely an example and should not limit the functionality and scope of use of the embodiments of the present application.
[0221] Computing device 500 is implemented as a general-purpose computing device. Components of computing device 500 may include, but are not limited to, one or more processors 510, memory 520, a communication bus 540 connecting various system components (including memory 520 and processor 510), and a communication interface 530.
[0222] Communication bus 540 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0223] The computing device 500 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device, including volatile and non-volatile media, removable and non-removable media.
[0224] The memory 520 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The computing device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Figure 7Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a floppy disk) and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read On ly Memory (hereinafter referred to as: CD-ROM), a Digital Versatile Disc Read On ly Memory (hereinafter referred to as: DVD-ROM) or other optical media) may be provided. In these cases, each drive can be connected to the communication bus 540 via one or more data medium interfaces. The memory 520 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.
[0225] A program / utility having a set (at least one) of program modules may be stored in memory 520. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. The program modules generally perform the functions and / or methods of the embodiments described herein.
[0226] The computing device 500 may also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), one or more devices that enable a user to interact with the computing device, and / or any device that enables the computing device to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through the communication interface 530. Furthermore, the computing device 500 may also communicate with the network adapter ( Figure 7 The network adapter can communicate with other modules of the computing device through the communication bus 540. It should be understood that although Figure 7 Not shown, other hardware and / or software modules may be used in conjunction with the computing device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, disk arrays (Redundant Arrays of Independent Drives; hereinafter referred to as: RAID) systems, tape drives, and data backup storage systems.
[0227] The processor 510 executes various functional applications and data processing by running the programs stored in the memory 520, such as implementing the method provided in the embodiment of the present application.
[0228] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely illustrative and does not constitute a structural limitation on the computing device 500. In other embodiments of the present application, the computing device 500 may also adopt a different interface connection method from the above embodiments, or a combination of multiple interface connection methods.
[0229] It is understandable that, in order to realize the above functions, the above-mentioned computing devices and the like include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily appreciate that, in combination with the various exemplary units and algorithm steps described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0230] The embodiment of the present application can divide the functional modules of the above-mentioned computing device etc. according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0231] The present application also provides a storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a computing device, the computing device implements the above-mentioned method.
[0232] An embodiment of the present application also provides a program product, which includes a computer program. When at least one processor executes the computer program, the at least one processor executes the method provided by the embodiment of the present application.
[0233] The computing device, storage medium or computer program product provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above and will not be repeated here.
[0234] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0235] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0236] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0237] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for managing an edge cluster, characterized in that: The method is applied to a master node in a first edge cluster, the first edge cluster including the master node and a slave node, the master node and the slave node both including a first service; and Determining a communication status of the first edge cluster, where the communication status of the first edge cluster is used to indicate whether the master node and the slave node can communicate normally; Determine the communication status of the cloud, where the communication status of the cloud is used to indicate whether the master node can communicate normally with the cloud; The first service is managed according to the communication status of the first edge cluster and the communication status of the cloud.
2. The management method according to claim 1, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: When the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, the first edge cluster does not migrate the first service, where the migration indicates migrating the first service from one or more nodes to other nodes.
3. The management method according to claim 1 or 2, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: When the communication status of the first edge cluster indicates normal and the communication status of the cloud indicates abnormal, re-election of the master node is triggered.
4. The management method according to any one of claims 1 to 3, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: When the communication status of the first edge cluster indicates an abnormality and the communication status of the cloud indicates a normality, the first edge cluster does not migrate the first service, where the migration indicates migrating the first service from one or more nodes to other nodes.
5. The management method according to any one of claims 1 to 4, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: When the communication status indication of the first edge cluster and the communication status indication of the cloud are abnormal, the first service in the node where the communication status indication is abnormal in the first edge cluster is migrated to the node where the communication status indication is normal in the second edge cluster, and the second edge node and the first edge node are located in the same local area network.
6. The management method according to claim 5, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud further includes: Triggering the re-election of the master node; When the communication status of the cloud determined each time the master node is re-elected indicates an abnormality, the first service in the node whose communication status in the first edge cluster indicates an abnormality is migrated to the node whose communication status in the second edge cluster indicates a normality.
7. The management method according to claim 6, characterized in that: Managing the first service according to the communication status of the first edge cluster and the communication status of the cloud includes: When the communication status of the cloud determined each time the master node is re-elected indicates that it is normal, the first edge cluster does not migrate the first service, and the migration is used to indicate migrating the first service from one or more nodes to other nodes.
8. The management method according to any one of claims 1 to 7, characterized in that: The determining the communication status of the first edge cluster includes: In the first edge cluster, when the communication status indication of the slave nodes greater than or equal to a first preset number indicates an abnormality, determining that the communication status indication of the first edge cluster is abnormal, the communication status of the slave node is used to indicate whether the slave node can communicate normally with the master node; In the first edge cluster, when the communication status indication of the slave nodes greater than or equal to the first preset number is normal, it is determined that the communication status indication of the first edge cluster is normal.
9. The management method according to claim 8, characterized in that: The determining the communication status of the first edge cluster further includes: Determining whether a heartbeat packet from the slave node is received within a preset time period; If a heartbeat packet of the slave node is received, determining that the communication status indication of the slave node is normal; If the heartbeat packet of the slave node is not received, it is determined that the communication status indication of the slave node is abnormal.
10. The management method according to any one of claims 1 to 9, characterized in that: Determining the communication status of the cloud includes: In the case of communication connection with the cloud, determining that the communication status indication of the cloud is normal; In a case where communication with the cloud is disconnected, it is determined that the communication status of the cloud indicates an abnormality.
11. The management method according to any one of claims 1 to 10, characterized in that: The method further comprises: Synchronize service information to each node in the first edge cluster so that the service information stored by each node is consistent, wherein the service information is used to indicate whether each node in the first edge cluster can normally provide the first service.
12. The management method according to claim 11, characterized in that: The heartbeat packets received and sent by the master node are stored in a data structure ebpf map in the Berkeley Packet Filter EBPF program of the master node, and the heartbeat packets received and sent by the master node represent the service information; The synchronizing the service information to each node in the first edge cluster so that the service information stored by each node is consistent includes: The heartbeat packets in the eBPF map are stored in the memory through the distributed consensus algorithm Raft protocol, so that each node obtains the service information by reading the heartbeat packets sent and received by the master node from the memory.
13. The management method according to claim 11 or 12, characterized in that: The method further comprises: When the communication status of the slave node indicates that the slave node is normal, determining whether the slave node can normally provide the first service according to a heartbeat packet of the slave node, wherein the heartbeat packet indicates whether the slave node can normally provide the first service; In a case where the communication status of the slave node indicates an abnormality, it is determined that the slave node cannot normally provide the first service.
14. The management method according to any one of claims 11 to 13, characterized in that: The method further comprises: When the communication status of the cloud indicates normal, the service information is sent to the cloud.
15. A management device for an edge cluster, characterized in that: The device is applied to a master node in the first edge cluster, the first edge cluster including the master node and a slave node, the master node and the slave node both including a first service; and the device includes: A first determining module is configured to determine a communication status of the first edge cluster, where the communication status of the first edge cluster is used to indicate whether the master node and the slave node can communicate normally; A second determining module is used to determine the communication status of the cloud, where the communication status of the cloud is used to indicate whether the master node and the cloud can communicate normally; A management module is configured to manage the first service according to the communication status of the first edge cluster and the communication status of the cloud.
16. A computing device, characterized in that include: a processor, a memory for storing instructions executable by the processor; When the processor is configured to execute the instructions, the computing device implements the method executed by the master node according to any one of claims 1 to 14.
17. An edge cluster, characterized in that: The edge cluster includes a master node and a slave node, and the master node and the slave node both include a first service, wherein: The master node is used to determine the communication status of the first edge cluster, where the communication status of the first edge cluster is used to indicate whether the master node and the slave node can communicate normally; to determine the communication status of the cloud, where the communication status of the cloud is used to indicate whether the master node and the cloud can communicate normally; and to manage the first service based on the communication status of the first edge cluster and the communication status of the cloud.