A health check method for a large-scale cluster and related devices
Patent Information
- Application Number
- CN202510390965.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]目前,由于每个负载均衡器都需要探测所有的业务节点,且每个业务节点的探测结果不能被其他负载均衡器复用,当业务集群大规模增长时,负载均衡器的计算资源被过多的分配给探活业务,导致降低了负载均衡器的路由业务的计算资源,从而降低负载均衡器的路由性能
[0023]该实现方式中,通过滑动窗口的方式在某时间段(如第二时长)内对第一业务节点进行多次探测,减少探测结果受瞬时网络波动、设备负载等偶然因素的影响,汇聚多个探测结果,减少异常值的干扰,从而提高第一业务节点的健康数据的准确性和可靠性。
Smart Images

Figure CN122845486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and related equipment for large-scale cluster health checks. Background Technology
[0002] Node health checks, also known as node probing, are the periodic checks performed by load balancers on the health of service nodes (such as servers, containers, and virtual machines) to ensure that service requests are only distributed to available nodes. In the realm of traffic routing, load balancers are responsible for distributing service requests to different member nodes of a service cluster and for centrally probing each member node. In other words, load balancers handle both routing and health checks for the service cluster. As service traffic grows rapidly, the size of the service cluster increases significantly, and the size of the load balancer cluster also grows accordingly.
[0003] Currently, because each load balancer needs to probe all business nodes, and the probe results of each business node cannot be reused by other load balancers, when the business cluster grows on a large scale, the computing resources of the load balancer are excessively allocated to the probe business, which reduces the computing resources of the load balancer's routing business, thereby reducing the routing performance of the load balancer. Summary of the Invention
[0004] This application provides a health check method and related equipment for large-scale clusters. By decoupling the detection services of the load balancing cluster, the detection cluster performs group detection on the service cluster, thereby improving the routing performance of the load balancing cluster.
[0005] Firstly, this application provides a health check method for a large-scale cluster, applied to a liveness detection center. The liveness detection center runs on a liveness detection cluster, which communicates with a load balancing cluster and a service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. The method includes: obtaining a task view, which includes a first association relationship between multiple first liveness detection nodes in the liveness detection cluster and a first service node in the service cluster; configuring multiple first liveness detection nodes to probe the first service node based on the first association relationship, obtaining multiple first probe results for the first service node, wherein one of the multiple first liveness detection nodes is used to probe multiple first service nodes; aggregating the multiple first probe results to obtain health data of the first service node; and sending the health data of the first service node to the load balancing cluster.
[0006] The task view contains the configuration information for the detection tasks performed by the detection nodes, including the first detection node and the multiple first service nodes that each first detection node is responsible for detecting, the detection strategy of the first detection node for detecting the first service nodes, and the first association relationship between the multiple first detection nodes and the first service nodes.
[0007] Based on the above scheme, the detection center configures multiple first detection nodes to probe the first business node according to the task view, obtains multiple first detection results, and aggregates the multiple first detection results to obtain the health data of the first business node. By having the detection center configure the detection nodes in the business cluster to detect the business node, instead of having the load balancing cluster perform the detection, the detection business of the load balancing cluster is decoupled. The computing resources of the load balancing cluster do not need to be used for the detection business, but are only responsible for routing and forwarding the business requests of the business node, thereby improving the routing performance of the load balancing cluster.
[0008] Furthermore, configuring multiple first-level detection nodes to probe the first-level service nodes can reduce the risk of single-point failures at the detection nodes, shortening the failure detection latency of service nodes from minutes to seconds, and improving the reliability of the detection cluster. One first-level detection node is used to probe multiple first-level service nodes, enabling grouped probing of the service cluster. This avoids the first detection node probing all service nodes, reducing the number of repeated probes and thus reducing the consumption of computing and bandwidth resources for the detection service.
[0009] In one possible implementation, the task view includes a second association between multiple second probe nodes in the probe cluster and a second business node in the business cluster; based on the second association, multiple second probe nodes are configured to probe the second business node, obtaining multiple second probe results for the second business node; the multiple second probe results are aggregated to obtain health data for the second business node; and the health data of the first business node and the health data of the second business node are aggregated to obtain health data for the business cluster. The first and second business nodes can be of the same type, such as both being database service nodes, or they can be of different types, such as the first business node being a database service node and the second business node being a web service node; the specific type is not limited here.
[0010] This implementation method is not limited to the health data of a single business node, but can provide the health data of the entire business cluster, which is helpful for troubleshooting or optimizing cluster configuration.
[0011] In one possible implementation, before obtaining the task view, a first health query request sent by the load balancing cluster is received. The first health query request includes the node identifier of the first business node and / or the node identifier of the second business node. The task view is determined based on the number of live nodes in the live node detection cluster, the node identifier of the first business node, and / or the node identifier of the second business node.
[0012] For example, based on the number of liveness detection nodes, the number of first business nodes, and / or the number of second business nodes, the liveness detection nodes can be evenly distributed among the first and / or second business nodes to ensure that the number of business nodes each liveness detection node is responsible for is within a reasonable range, avoiding an excessive or insufficient number of business nodes detected by a particular liveness detection node, and optimizing resource utilization.
[0013] In this implementation, the task view is determined based on the number of active nodes and node identifiers, which enables efficient task scheduling and load balancing, and improves resource utilization.
[0014] In one possible implementation, a second health query request sent by the load balancing cluster is received, the second health query request including the node identifier of the second service node; the health data of the second service node is determined based on the node identifier of the second service node and the health data of the service cluster; and the health data of the second service node is sent to the load balancing cluster.
[0015] In this implementation, by receiving the second health query request and returning the health data of the second business node, the efficiency of obtaining the health data of the second business node can be improved.
[0016] In one possible implementation, based on the first association relationship, multiple first probe nodes are configured to send multiple first probe requests to the first service node; in response to the multiple first probe requests, multiple first probe results sent by the first service node are received.
[0017] In this implementation, by configuring multiple first probe nodes to send multiple first probe requests to the first service node, the abnormal impact that a single first probe node or first probe request may cause can be effectively reduced, and probe failure caused by the failure of a single first probe node can be avoided, thereby improving the reliability of the probe cluster.
[0018] In one possible implementation, the task view further includes a probing strategy for probing the first service node and / or the second service node. The probing strategy includes a first duration and / or a first threshold and / or protocols supported by the probing request when probing the first service node and / or the second service node. The protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Protocol Secure (HTTPS), and service-based protocols. The first duration is the duration of probing the first service node and / or the second service node, and the first threshold is the upper limit of the number of failed probes and / or the upper limit of the number of successful probes for the first service node and / or the second service node.
[0019] In this implementation, the initial duration can be flexibly adjusted according to actual network conditions and business needs, avoiding misjudgments due to excessively short probe durations or unnecessary delays due to excessively long probe durations; the probe request supports multiple protocols, which can adapt to different types of business needs and network environments, enhancing adaptability to different types of business nodes; by setting an upper limit on the number of successful or failed continuous probes, false alarms caused by short-term network fluctuations or occasional failures can be avoided.
[0020] In one possible implementation, a decision is made based on multiple first probe results through a weighted and / or voting mechanism to obtain the health data of the first business node.
[0021] In this implementation, a weighted mechanism assigns different weights to different liveness detection nodes based on their reliability and accuracy, thereby increasing the detection results of highly reliable nodes and thus improving the credibility of the decision. The voting mechanism, based on the majority principle, reduces the potential errors of a single liveness detection node, thereby enhancing the credibility of the decision.
[0022] In one possible implementation, a sliding window is used to configure the first detection node to perform multiple detections on the first business node during the second time period, thereby obtaining multiple third detection results corresponding to the first business node during the second time period. The window size is the first time period, and the second time period is greater than or equal to the first time period. The multiple third detection results are aggregated to obtain the health data of the first business node corresponding to the second time period.
[0023] In this implementation, the first service node is probed multiple times within a certain time period (such as the second duration) by using a sliding window. This reduces the impact of instantaneous network fluctuations, equipment load and other accidental factors on the probe results, and aggregates multiple probe results to reduce interference from outliers, thereby improving the accuracy and reliability of the health data of the first service node.
[0024] Secondly, this application provides a liveness detection center that runs on a liveness detection cluster. The liveness detection cluster communicates with a load balancing cluster and a service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. The liveness detection center includes a scheduler, detectors, decision-makers, and transceivers.
[0025] The scheduler is used to obtain the task view, which includes the first association between multiple first liveness detection nodes in the liveness detection cluster and the first business node in the business cluster.
[0026] The detector is used to configure multiple first detection nodes to detect the first service node based on the first association relationship, and obtain multiple first detection results of the first service node. Among them, one of the multiple first detection nodes is used to detect the first service node.
[0027] The decision-maker is used to aggregate multiple first detection results to obtain the health data of the first business node;
[0028] The transceiver is used to send health data of the first service node to the load balancing cluster.
[0029] In one possible implementation, the task view includes a second association between multiple second detection nodes in the detection cluster and a second business node in the business cluster. The detector is also used to configure multiple second detection nodes to detect the second business node based on the second association, and obtain multiple second detection results of the second business node. The decision-maker is also used to aggregate the multiple second detection results to obtain the health data of the second business node. The detection center also includes an aggregator, which is used to aggregate the health data of the first business node and the health data of the second business node to obtain the health data of the business cluster.
[0030] In one possible implementation, before obtaining the task view, the transceiver is further configured to receive a first health query request sent by the load balancing cluster, the first health query request including the node identifier of the first service node and / or the node identifier of the second service node; the scheduler is further configured to determine the task view based on the number of live nodes in the live node probe cluster, the node identifier of the first service node and / or the node identifier of the second service node.
[0031] In one possible implementation, the transceiver is also used to receive a second health query request sent by the load balancing cluster, the second health query request including the node identifier of the second service node; the liveness detection center also includes an executor, which is used to determine the health data of the second service node based on the node identifier of the second service node and the health data of the service cluster; the transceiver is also used to send the health data of the second service node to the load balancing cluster.
[0032] In one possible implementation, the detector is further configured to send multiple first probe requests to the first service node based on the first association relationship; the transceiver is further configured to receive multiple first probe results sent by the first service node in response to the multiple first probe requests.
[0033] In one possible implementation, the task view further includes a probing strategy for probing the first service node and / or the second service node. The probing strategy includes a first duration and / or a first threshold and / or protocols supported by the probing request when probing the first service node and / or the second service node. The protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Security Protocol (HTTPS), and service-based protocols. The first duration is the duration of probing the first service node and / or the second service node, and the first threshold is the upper limit of the number of failed probes and / or the upper limit of the number of successful probes for the first service node and / or the second service node.
[0034] In one possible implementation, the decision-maker is also used to make decisions on multiple first probe results through a weighted and / or voting mechanism to obtain the health data of the first business node.
[0035] In one possible implementation, the detector is also configured to use a sliding window to configure the first detection node to perform multiple detections on the first service node during the second duration, thereby obtaining multiple third detection results corresponding to the first service node during the second duration. The window size is the first duration, and the second duration is greater than or equal to the first duration. The decision-maker is also configured to aggregate the multiple third detection results to obtain the health data of the first service node during the second duration.
[0036] Thirdly, this application provides a health check system, including a load balancing cluster, a life detection cluster, and a service cluster. The life detection cluster communicates with the load balancing cluster and the service cluster through a network. The life detection cluster and the load balancing cluster are different clusters.
[0037] The load balancing cluster is used to send health query requests to the health monitoring cluster. The health query request is used to obtain the health data of the first business node. In response to the health query request, the health data of the first business node sent by the health monitoring cluster is received.
[0038] The detection cluster is used to obtain a task view, which includes a first association between multiple first detection nodes in the detection cluster and a first business node in the business cluster. Based on the first association, multiple first detection nodes are configured to probe the first business node, obtaining multiple first probe results for the first business node. Among these, one of the multiple first detection nodes is used to probe the first business node. The multiple first probe results are aggregated to obtain the health data of the first business node. The health data of the first business node is sent to the load balancing cluster. The business cluster is used to receive multiple probe requests sent by the detection cluster, and in response to the multiple probe requests, sends multiple first probe results for the multiple probe requests. The multiple first probe results are used to determine the health data of the first business node.
[0039] The health check system implements the methods executed in the various possible implementations of the first aspect mentioned above.
[0040] Fourthly, this application provides a communication device including at least one processor for executing programs or instructions in a memory to enable the device to implement the methods executed in the first aspect and various possible implementations thereof.
[0041] Fifthly, this application provides a communication device including at least one logic circuit and an input / output interface; the logic circuit is used to perform the methods performed as described in the first aspect and various possible implementations of the first aspect.
[0042] In a sixth aspect, this application provides a computing device including a processor and a memory, the processor being coupled to the memory, the memory being used to store programs or instructions, and when the program or instructions are executed by the processor, causing the computing device to perform the methods performed as described in the first aspect and various possible implementations of the first aspect.
[0043] In a seventh aspect, this application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the methods performed as described in the first aspect and various possible implementations of the first aspect above.
[0044] Eighthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, execute the methods described in the first aspect and various possible implementations thereof.
[0045] Ninthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the methods as described in the first aspect and various possible implementations thereof. Attached Figure Description
[0046] Figure 1 This application provides a schematic diagram of the architecture of a health check system.
[0047] Figure 2 A flowchart illustrating a large-scale cluster health check method provided in this application;
[0048] Figure 3 A schematic diagram of the structure of a life detection center provided in this application;
[0049] Figure 4 A schematic diagram illustrating how to obtain health data of a business cluster at different time periods using a sliding window method, as provided in this application;
[0050] Figure 5 A schematic diagram of the structure of a computing device provided in this application;
[0051] Figure 6 This application provides a schematic diagram of the structure of a computing device cluster;
[0052] Figure 7 This is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. With the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0054] The terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of modules is not necessarily limited to those modules, but may include other modules not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0055] The terms "system" and "network" in this application are used interchangeably. "At least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B, and C" includes A, B, C, AB, AC, BC, or ABC.
[0056] In this application, the sequence number does not imply the order of execution. The execution order should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. It is understood that in the embodiments of this application, descriptions such as "in the case of," "if," "when," "if," etc., can be used interchangeably. Furthermore, these descriptions all refer to the corresponding processing that will be carried out under certain objective circumstances, and are not time-limited, nor do they require any judgment action during implementation, nor do they imply any other limitations.
[0057] The following describes the relevant terms and concepts that may be involved in the embodiments of this application.
[0058] (1) Load balancing
[0059] Load balancing (LB), also known as load scheduling, is a technology that distributes network traffic, requests, or workloads across multiple servers (such as the business nodes in this application), enabling these servers to collaboratively complete tasks. These servers can be File Transfer Protocol (FTP) servers, web servers, core enterprise application servers, and other primary task servers. By distributing traffic across multiple servers, load balancing ensures that each server bears the load evenly, reducing the load on any single server and thus improving the overall system's processing capacity and response speed.
[0060] Load balancing can be implemented in hardware or software. Hardware implementations use dedicated hardware devices (such as load balancers) to distribute requests. The load balancer monitors server load and distributes requests to different servers based on a pre-defined load balancing algorithm (such as round-robin or least connections). Software implementations, such as virtualized load balancers and Nginx, listen for client requests and distribute them to different servers according to a pre-defined load balancing algorithm.
[0061] According to the different layers of the Open Systems Interconnection (OSI) seven-layer model, load balancing can also include network layer load balancing, application layer load balancing, etc. Network layer load balancing distributes traffic to different servers through network layer devices (such as routers and switches), including IP-based load balancing and Domain Name System (DNS)-based load balancing. Application layer load balancing distributes traffic to different servers based on the content, protocol, or other characteristics of the requests through application layer devices (such as load balancers and reverse proxy servers).
[0062] (2) Node Activation
[0063] Node liveness detection, also known as node health check (HC), is the process by which a load balancer periodically checks the health status of service nodes to ensure that requests are only distributed to available service nodes. These service nodes can be servers, containers, virtual machines, etc.
[0064] Common health check methods include heartbeat detection, response timeout detection, and load detection. In heartbeat detection, the load balancer periodically sends heartbeat requests to service nodes to check if they are responding correctly. In response timeout detection, the load balancer sets a reasonable response timeout period; if a service node does not respond within the specified time, it is considered unavailable. In load detection, the load balancer monitors the load on service nodes, such as central processing unit (CPU) utilization and memory usage, to determine the availability of the service nodes.
[0065] (3) Service cluster / business cluster
[0066] A computer cluster, or simply cluster, is a loosely integrated group of computers connected together in a highly collaborative manner to perform computational tasks. The individual computers in a computer cluster are typically called nodes.
[0067] A service cluster refers to deploying multiple service instances (usually multiple replicas of the same service) into a cluster to provide features such as high availability, load balancing, and fault recovery. Service clusters are commonly found in distributed architectures, microservice architectures, and cloud computing environments. For example, a liveness detection service for business nodes can be deployed as multiple instances, with these instances belonging to the same service cluster (also known as a liveness detection cluster).
[0068] A business cluster is a collection of business function modules or business processes, rather than simply a deployment of technical services. It typically refers to integrating multiple business-related components or functions into a cluster, designed to work collaboratively to provide a specific business function or achieve a particular business goal. A business cluster includes multiple services; for example, an order processing cluster includes services such as inventory checking, order confirmation, and payment processing. These services together constitute the business cluster. Service clusters and business clusters can be interchanged as long as there is no logical conflict.
[0069] (4) Process call
[0070] Inter-process communication (IPC) includes local procedure calls (LPC) and remote procedure calls (RPC). LPC is a method of communication between different processes on the same machine. RPC is a method of communication between different machines on a network. IPC allows one process to call a function or method in another process. LPC provides efficient local communication, while RPC enables communication in distributed systems. In this embodiment, communication between the liveness detection client and the liveness detection center, as well as communication between liveness detection nodes, can be achieved through inter-process communication.
[0071] Currently, load balancers function to perform both liveness detection and routing. Because each load balancer needs to probe all service nodes, and the probe results for each service node cannot be reused by other load balancers, excessive computing resources are allocated to liveness detection, reducing the computing resources available for other services and impacting the load balancer's routing performance.
[0072] In view of this, this application provides a health check method and related equipment for large-scale clusters. By decoupling the liveness detection service of the load balancing cluster, the liveness detection cluster performs grouped detection on the service cluster, thereby improving the performance of the load balancing cluster and reducing the consumption of computing and bandwidth resources in the liveness detection service. Furthermore, embodiments of this application also provide a liveness detection center, a health check system, a computing device cluster, a computer program product, and a computer-readable storage medium.
[0073] Please see Figure 1 As shown, Figure 1 This is a schematic diagram of the architecture of a health check system provided in an embodiment of this application. The health check system includes a liveness detection cluster, a load balancing cluster, and a service cluster. This system is a distributed health detection and load scheduling platform for the service cluster, used to monitor the health status of the service cluster, ensure the reliability of the service cluster, and perform load scheduling by the load balancing cluster based on the health status of the service cluster. The specific functions of each part of the system are described below.
[0074] The business cluster comprises multiple business nodes, each capable of deploying different types of services, such as web services, database services, caching services, queue services, and file storage services. Web services handle business requests. Database services provide data storage, querying, and transaction processing services, such as MySQL, PostgreSQL, MongoDB, and Redis. Caching services accelerate data access and reduce database pressure, such as Redis and Memcached. Queue services can be used to handle message queues for asynchronous tasks, such as Kafka and RabbitMQ. File storage services handle large volumes of file storage and access requests, such as cloud storage services and traditional distributed file systems. For a more detailed introduction to the business cluster, please refer to the preceding sections; further details will not be provided here.
[0075] Business nodes can provide services through hardware or software, such as virtual machines, physical machines, service instances, and microservice containers, without further limitation here. A virtual machine is multiple independent runtime environments created on physical hardware using virtualization technology; each virtual machine can run an operating system and applications. A physical machine (physical server) refers to computing resources running directly on physical hardware. A service instance can be an instantiated object of a specific running service. In a microservice architecture, a microservice may have multiple instances running on different nodes to handle concurrent requests. In a containerized architecture, microservices are typically packaged into containers, and multiple containers may run on different physical machines or virtual machines.
[0076] In practical applications, business clusters can adopt a hybrid architecture, combining various computing resources such as physical machines, virtual machines, service instances, and microservice containers to form a flexible and scalable business cluster. On some large-scale cloud platforms, virtual machines and containerized microservices can be deployed simultaneously. For example, database services run on virtual machines, while web services and API gateways can be deployed as containers, and the entire system is managed and coordinated through a container orchestration platform (such as Kubernetes). In edge computing architectures, business clusters can combine physical machines, edge servers, and cloud platform resources to form a distributed computing and data processing network.
[0077] A load balancing cluster comprises multiple load balancers. Load balancers dynamically adjust traffic and computing resource allocation for service nodes based on their health status. For example, when a service node is determined to be unhealthy, the load balancer will redirect traffic from that node to other healthy service nodes to ensure service availability. Alternatively, the load balancer can allocate more computing resources to healthy service nodes, increasing their processing capacity, while reducing the computing resources allocated to unhealthy nodes. Furthermore, the load balancing cluster can implement fault tolerance mechanisms, automatically switching traffic from a failed load balancer to a backup load balancer to ensure high availability of the load balancing service. A load balancing cluster may include multiple grid routes (GRs), which are responsible for accurately routing external traffic to backend service nodes. Each GR instance handles a portion of the service traffic routing tasks and can make intelligent traffic distribution decisions based on dynamic factors such as request content, user location, and service load status. The various GR instances in the load balancing cluster work collaboratively through some communication mechanism (such as shared configuration or message queues) to ensure efficient traffic scheduling.
[0078] The liveness detection cluster comprises multiple liveness detection nodes. A liveness detection center runs on the cluster, configuring the nodes accordingly to perform liveness detection on the business cluster. When a failure is detected in a business node, the fault information is promptly reported to the load balancing cluster to ensure the reliability and stability of the business cluster. The liveness detection nodes or center can provide liveness detection services through hardware or software, such as virtual machines, physical machines, service instances, microservice containers, etc., without specific limitations here. Multiple liveness detection nodes can be deployed distributed across different computing devices or integrated within the same computing device, without specific limitations here. The specific functions of the liveness detection center are described in subsequent embodiments and will not be repeated here.
[0079] In addition, the liveness detection cluster also includes a liveness detection client. In practical applications, administrators can interact with the liveness detection center through the liveness detection client. For example, administrators can assign detection tasks to detection nodes and configure the detection nodes that execute the detection tasks, the corresponding business nodes of the detection nodes, detection policies, etc., through the application interface of the liveness detection client. Another example is that the liveness detection client can encapsulate the interfaces provided by the liveness detection center. The liveness detection client calls the internal interfaces of the liveness detection center through process calls, providing only a transparent business interface to the outside world for querying the health data of the business cluster, while the liveness detection center internally implements interfaces such as adding business nodes and querying the health data of the business cluster.
[0080] exist Figure 1In the system architecture shown, the load balancer sends a query request to the health detection center, which can be used to obtain health data of the business cluster. The health detection center receives the query request, and multiple health detection nodes in the health detection cluster send probe requests to the business cluster, performing grouped probes to obtain health data for the business cluster and then sending this health data to the load balancer. Based on the received health data of the business cluster, the load balancer schedules business requests to healthy business nodes, thus achieving health checks and load scheduling for the business cluster. This application decouples the health detection service of the load balancer, allowing the health detection cluster to perform health detection on the business nodes instead of the load balancer itself. The load balancer's computing resources are not used for health detection but are only responsible for routing and forwarding business requests from the business nodes, improving the routing performance of the load balancer. The health detection cluster performs health detection on the business cluster in groups, reducing the number of repeated probes on business nodes, thereby reducing the consumption of computing and bandwidth resources for health detection.
[0081] Combination Figure 1 The health check system shown below illustrates a large-scale cluster health check method provided in this application embodiment. Please refer to [link to relevant documentation]. Figure 2 As shown, Figure 2 This is a flowchart illustrating a health check method for a large-scale cluster provided in this application embodiment. The method is applied to a liveness detection center, which runs on a liveness detection cluster. The liveness detection cluster communicates with a load balancing cluster and a service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. The method specifically includes the following steps:
[0082] 201. Receive a first health query request sent by the load balancing cluster. The first health query request includes the node identifier of the first service node and / or the node identifier of the second service node.
[0083] A health query request is used to retrieve health data for a business cluster or a specific business node within a business cluster. A health query request can include query parameters and query conditions. Query parameters specify the business node being queried, such as the node identifier. Query conditions filter the query results; for example, a query condition could be "query business nodes in the business cluster that are in a healthy state." Health data includes the health status and metrics of the business nodes. Health status includes "healthy" or "unhealthy," and metrics include CPU utilization, memory usage, network connectivity, and request processing latency.
[0084] The first health query request includes the node identifier of the first business node and / or the node identifier of the second business node. The first health query request is used to obtain the health data of the first business node and / or the second business node. The first business node and the second business node can be business nodes of the same type, such as both being database service nodes, or they can be business nodes of different types, such as the first business node being a database service node and the second business node being a web service node. The specifics are not limited here.
[0085] 202. Obtain the task view, which includes the first association between multiple first detection nodes in the detection cluster and the first business node in the business cluster.
[0086] The task view contains the configuration information for the probe nodes to perform probe tasks. This includes the probe nodes and the multiple service nodes each probe node is responsible for probing, the probe strategies used by the probe nodes to probe the service nodes, and the relationships between the multiple probe nodes and the service nodes. After receiving a health query request from the load balancing cluster, the probe center can determine the task view based on the health query request. Specifically, the task view is determined based on the number of probe nodes in the probe cluster, the node identifier of the first service node, and / or the node identifier of the second service node.
[0087] For example, based on the number of probe nodes, the number of first service nodes, and / or the number of second service nodes, probe nodes can be evenly distributed among the first and / or second service nodes. This ensures that each probe node is responsible for a reasonable number of service nodes, preventing any single probe node from probing too many or too few service nodes, thus optimizing resource utilization. Determining the task view based on the number of probe nodes and node identifiers enables efficient task scheduling and load balancing, improving resource utilization. The task view includes the relationships between multiple probe nodes and service nodes; a service node is probed by multiple probe nodes, reducing the risk of single points of failure and improving the reliability of the probe cluster.
[0088] Optionally, the relationship between business nodes and activity detection nodes can be adjusted based on factors such as geographical location and business priority. For example, a business node can be assigned to the nearest activity detection node based on the geographical location information of the business node and the activity detection node. Another example is assigning more activity detection nodes to business nodes with higher business priority.
[0089] Optionally, task views can also be obtained from the database or configuration center that stores task views via a specific interface. When a new liveness detection node is added, the association between the liveness detection node and the business node in the database or configuration center can be adjusted. When a business node fails, the relevant information of that business node can be removed from the database or configuration center.
[0090] 203. Based on the first association relationship, configure multiple first detection nodes to probe the first service node, and obtain multiple first detection results for the first service node. Among them, one of the multiple first detection nodes is used to probe multiple first service nodes.
[0091] The task view also includes a probing strategy for probing the first and / or second business nodes. This strategy includes a first duration and / or a first threshold for probing the first and / or second business nodes, and / or the protocols supported by the probing requests. These protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Security Protocol (HTTPS), and business-based protocols. The business-based protocol is determined based on the specific business or service of the business node. For example, when the business node is a database service node, the protocol could be MySQL, PostgreSQL, etc.; when the business node is a cache service node, the protocol could be Redis, etc. The probe node sends corresponding requests to the business node according to the protocols supported by the business node to probe it. Specifically, when the probe request supports TCP or UDP, i.e., the business node is a network layer node, the probe node can probe the network connectivity of the network layer node using ping. When the probe request supports HTTP or HTTPS, i.e., the business node is an application layer node, such as a service instance, the probe node can probe it by sending HTTP or HTTPS requests to a specific application programming interface (API). The first duration is the duration for probing the first business node and / or the second business node, and the first threshold is the upper limit of the number of failed probes and / or successful probes of the first business node and / or the second business node. A probe node can perform multiple probe tasks on the same business node, and the probe result for that business node is determined based on the results of the multiple probes. For example, if the number of successful probes in multiple consecutive probe tasks exceeds the first threshold, then the probe result of the probe node on the business node is determined to be a successful probe.
[0092] The activity detection center configures multiple activity detection nodes to perform detection tasks based on the task view. It parses the task view to determine the association between the multiple activity detection nodes and the business nodes, the multiple business nodes that the activity detection nodes are responsible for detecting, and the detection strategies. Specifically, based on the first association, the activity detection center configures multiple first activity detection nodes to send multiple first detection requests to the first business nodes. The first activity detection nodes can send the first detection requests to the first business nodes according to the first duration and / or first threshold and / or the protocol supported by the detection request in the detection strategy. In response to the multiple first detection requests, the activity detection center receives multiple first detection results sent by the multiple first business nodes.
[0093] Upon receiving the first probe request, the first business node first parses it. This parsing process includes extracting key information such as the request type, the node identifier of the probe node, and the request's timestamp. Simultaneously, it verifies the legitimacy of the probe request, ensuring that only verified requests are processed further. Next, for each legitimate probe request, the first business node performs a self-state check and collects relevant data. This includes checking the system's hardware resource usage, such as CPU and memory usage; checking the operational status of the business node, such as whether the service has started normally and can process business requests correctly; and collecting business data, such as throughput and response time. This business data can be used to determine the operational status of the first business node. Based on the results of the state check and data collection, the first business node determines the probe result. For example, if CPU usage is too high or the business service experiences abnormal errors, the probe result indicates a potential risk or fault. Conversely, if all indicators are within the normal range, the probe result indicates normal operation. The probe result can be presented in the form of a clear status code or descriptive information.
[0094] The first probe node receives the probe results from the first business node. Based on the probe results, the health status of the business node can be determined. For example, if the probe result shows an HTTP 200 status code, the probe result can be determined as successful. Conversely, if the business node does not respond for an extended period or the response includes error description information, the probe result can be determined as failed.
[0095] In addition, one of the multiple first-probe nodes is used to probe multiple business nodes, realizing grouped probing of the business cluster. This avoids each probe node probing all business nodes, reducing the number of probe requests and lowering the consumption of computing and bandwidth resources.
[0096] 204. By aggregating multiple first detection results, the health data of the first business node is obtained.
[0097] The liveness detection center makes decisions based on the aggregated first probe results to obtain the health data of the first business node. Specifically, decisions are made on multiple first probe results through weighted and / or voting mechanisms to obtain the health data of the first business node. For example, different liveness detection nodes can be assigned different weights based on the importance of each node's probe task. The probe results of multiple liveness detection nodes can be weighted and averaged to determine the health data of the business node. For example, liveness detection node A has a weight of 0.1 and a corresponding CPU utilization rate of 90%, while liveness detection node B has a weight of 0.9 and a corresponding CPU utilization rate of 30%. The weighted average value yields a CPU utilization rate of 36% for the business node's health data. Alternatively, a "veto power" can be set for a liveness detection node, and the health data of the business node can be determined based on the probe results of that node. Furthermore, a majority vote can be used to decide the health data of the business node based on the probe results of multiple liveness detection nodes. Using statistical algorithms to make decisions on probe results to obtain the health data of the business node can improve the accuracy of node liveness detection. By using multiple liveness detection nodes to probe the same business node, the latency for detecting faults in the business node can be reduced from minutes to seconds, thus improving the reliability of the liveness detection cluster.
[0098] Steps 203 and 204 describe the process of probing the first service node and determining its health data. Optionally, a sliding window approach is used to perform multiple probing tasks on the first service node within a certain time period to obtain its health data for that time period. The duration of this time period is called the second duration, which is greater than or equal to the first duration. First, the size and sliding step of the sliding window are determined. The window size can be the first duration, meaning the window size is the duration of one probing of the first service node, with one window corresponding to one probing task. The sliding step is used to control the time interval between two consecutive probing attempts, thereby controlling the probing frequency. Next, using the sliding window approach, the first probing node is configured to perform multiple probing attempts on the first service node within the second duration, obtaining multiple third probing results corresponding to the first service node within the second duration. Specifically, during each probing attempt, the first probing node can send a third probing request to the first service node according to the probing strategy, and in response to the third probing request, receive the third probing result sent by the first service node. Based on a pre-defined sliding step size, the window continuously slides, and new third-probe results are constantly added to the new window. By aggregating multiple third-probe results, the health data of the first business node corresponding to the second time period is obtained. For example, the average of the indicator data obtained from multiple probes within this time period is used as the indicator data in the health data of the first business node for that time period. In this way, the health data of the first business node for that time period is obtained.
[0099] 205. Send the health data of the first service node to the load balancing cluster.
[0100] Steps 201 to 205 describe the process of determining the health data of the first business node, and the subsequent steps describe the process of determining the health data of the business cluster and the health data of a certain business node in the business cluster.
[0101] 206. Based on the second association relationship, configure multiple second detection nodes to detect the second business node and obtain multiple second detection results of the second business node.
[0102] The task view also includes the second association between multiple second probe nodes in the probe cluster and the second business node in the business cluster, the probe strategy for the second probe node to probe the second business node, etc. The probe strategy includes the first duration and / or the first threshold and / or the protocol supported by the probe request when probing the second business node. For a description of the probe strategy, please refer to step 203.
[0103] The activity detection center configures multiple second activity detection nodes to perform detection tasks based on the task view. It parses the task view to determine the association between the multiple second activity detection nodes and the second service nodes, the multiple second service nodes that the second activity detection nodes are responsible for detecting, and the detection strategy. Specifically, based on the second association, the center configures multiple second activity detection nodes to send multiple second detection requests to the second service nodes. The second activity detection nodes can send second detection requests to the second service nodes according to the first duration and / or first threshold and / or the protocol supported by the detection request in the detection strategy. In response to the multiple second detection requests, the activity detection center receives multiple second detection results sent by the multiple second service nodes. The process of determining the second detection results can be found in step 203, and will not be repeated here.
[0104] 207. By aggregating multiple second detection results, the health data of the second business node is obtained.
[0105] For details, please refer to step 204, which will not be repeated here.
[0106] 208. Aggregate the health data of the first business node and the health data of the second business node to obtain the health data of the business cluster.
[0107] In this step, health data from multiple business nodes is aggregated to obtain health data for the business cluster. Then, step 212 or step 209 can be executed. It should be understood that this application does not specifically limit the number of the first and second business nodes.
[0108] 209. Receive a second health query request sent by the load balancing cluster. The second health query request includes the node identifier of the second business node.
[0109] The second health query request is used to obtain the health data of the second business node. For a description of the second health query request, please refer to step 201, which will not be repeated here.
[0110] 210. Determine the health data of the second business node based on the node identifier of the second business node and the health data of the business cluster.
[0111] Based on the node identifier of the second business node, health data related to that node identifier is filtered out from the business cluster to obtain the health data of the second business node. Specifically, this can be done by querying a data table based on the node identifier, or by matching records related to that node identifier in a log file to determine the health data of the second business node.
[0112] 211. Send the health data of the second service node to the load balancing cluster.
[0113] 212. Send the health data of the service cluster to the load balancing cluster.
[0114] Accordingly, the load balancer receives health data from the service cluster. The load balancer can then perform traffic scheduling based on this health data. For example, when a service node is marked as "unhealthy," the load balancer can remove that node from the scheduling list, preventing traffic from being directed to it.
[0115] It should be noted that, Figure 2 In the method embodiment shown, steps 201, 206 to 212 are optional steps.
[0116] In summary, the liveness detection center obtains a task view and, based on this view, configures multiple first liveness detection nodes to probe the first business nodes, obtaining multiple first probe results. These results are then aggregated to obtain the health data of the first business nodes. By configuring liveness detection nodes within the business cluster to probe the business nodes through the liveness detection center, instead of relying solely on the load balancing cluster, the liveness detection service of the load balancing cluster is decoupled. The load balancing cluster's computing resources are no longer needed for liveness detection; instead, they are only responsible for routing and forwarding business requests from the business nodes, thereby improving the routing performance of the load balancing cluster. Furthermore, configuring multiple first liveness detection nodes to probe the first business nodes reduces the risk of single-point failures at these nodes, shortening the failure detection latency of business nodes from minutes to seconds, thus improving the reliability of the liveness detection cluster. One first liveness detection node is used to probe multiple first business nodes, enabling grouped probing of the business cluster. This avoids the first liveness detection node probing all business nodes, reducing the number of repeated probes and thus reducing the consumption of computing and bandwidth resources for the liveness detection service.
[0117] This application also provides a life detection center, a health check system, a computing device cluster, a computer program product, and a computer-readable storage medium.
[0118] Please see Figure 3 As shown, Figure 3 This is a schematic diagram of a liveness detection center 300 provided in an embodiment of this application. The liveness detection center runs on a liveness detection cluster, which communicates with a load balancing cluster and a service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. The liveness detection center 300 can be used to implement... Figure 2 The function of the method embodiment shown can also achieve the beneficial effects of the above method embodiment. The activity detection center 300 includes a scheduler 301, a detector 302, a decision maker 303, an aggregator 304, an actuator 305, and a transceiver 306.
[0119] The scheduler is used to obtain the task view, which includes the first association between multiple first liveness detection nodes in the liveness detection cluster and the first business node in the business cluster.
[0120] The detector is used to configure multiple first detection nodes to detect the first service node based on the first association relationship, and obtain multiple first detection results of the first service node. Among them, one of the multiple first detection nodes is used to detect multiple first service nodes.
[0121] The decision-maker is used to aggregate multiple first detection results to obtain the health data of the first business node;
[0122] The transceiver is used to send health data of the first service node to the load balancing cluster.
[0123] In one possible implementation, the task view includes a second association between multiple second detection nodes in the detection cluster and a second business node in the business cluster. The detector is also used to configure multiple second detection nodes to detect the second business node based on the second association, and obtain multiple second detection results of the second business node. The decision-maker is also used to aggregate the multiple second detection results to obtain the health data of the second business node. The aggregator is used to aggregate the health data of the first business node and the health data of the second business node to obtain the health data of the business cluster.
[0124] In one possible implementation, before obtaining the task view, the transceiver is further configured to receive a first health query request sent by the load balancing cluster, the first health query request including the node identifier of a first service node and / or the node identifier of a second service node. The scheduler is further configured to determine the task view based on the number of probe nodes in the probe cluster, the node identifier of the first service node, and / or the node identifier of the second service node.
[0125] In one possible implementation, the transceiver is further configured to receive a second health query request sent by the load balancing cluster, the second health query request including the node identifier of the second service node; the executor is configured to determine the health data of the second service node based on the node identifier of the second service node and the health data of the service cluster; the transceiver is further configured to send the health data of the second service node to the load balancing cluster.
[0126] In one possible implementation, the detector is further configured to send multiple first probe requests to the first service node based on the first association relationship; the transceiver is further configured to receive multiple first probe results sent by the first service node in response to the multiple first probe requests.
[0127] In one possible implementation, the task view further includes a probing strategy for probing the first service node and / or the second service node. The probing strategy includes a first duration and / or a first threshold and / or protocols supported by the probing request when probing the first service node and / or the second service node. The protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Security Protocol (HTTPS), and service-based protocols. The first duration is the duration of probing the first service node and / or the second service node, and the first threshold is the upper limit of the number of failed probes and / or the upper limit of the number of successful probes for the first service node and / or the second service node.
[0128] In one possible implementation, the decision-maker is also used to make decisions on multiple first probe results through a weighted and / or voting mechanism to obtain the health data of the first business node.
[0129] In one possible implementation, the detector is also configured to use a sliding window to configure the first detection node to perform multiple detections on the first service node during the second duration, thereby obtaining multiple third detection results corresponding to the first service node during the second duration. The window size is the first duration, and the second duration is greater than or equal to the first duration. The decision-maker is also configured to aggregate the multiple third detection results to obtain the health data of the first service node during the second duration.
[0130] In one possible implementation, the scheduler is also configured to receive heartbeat data sent by the detector and / or decision-maker and / or aggregator and / or actuator and / or transceiver; the scheduler is also configured to determine the state of the detector and / or decision-maker and / or aggregator and / or actuator and / or transceiver based on the heartbeat data.
[0131] The scheduler monitors the activity of probe nodes through inter-process communication. These probe nodes can be detectors, decision-makers, aggregators, executors, or transceivers. Specifically, probe nodes can periodically report heartbeat data to the scheduler. This heartbeat data includes the probe node's process heartbeat and / or metrics data of the probe node's execution of probe tasks.
[0132] The process heartbeat is a signal periodically sent by the processes of the probe nodes to determine the liveness status of the processes. The process heartbeat can include the identifier of the probe node. Specifically, when the scheduler detects a process heartbeat timeout or loss, it can determine that the process of that probe node has crashed or stopped running, and can assign the probe task to a normally functioning probe node for execution. Furthermore, the scheduler can also infer the load of the probe nodes based on the time interval of the process heartbeats. For example, if the scheduler receives a process heartbeat for an excessively long time, it can determine that the probe node is under high load.
[0133] Metrics data are used to determine the status of probe nodes performing probe tasks. This data can include execution time, task success rate, resource consumption, and data integrity. Execution time is the time taken to execute the probe task; whether the execution time exceeds a certain range determines whether the probe node is functioning correctly. Task success rate is used to statistically analyze the probability of a probe node successfully executing a task multiple times consecutively. Resource consumption includes the usage of CPU, memory, disk I / O, and other resources by the probe node during task execution. Data integrity is used to determine whether data transmitted during task execution is lost or inconsistent. The scheduler can compare the current probe node's metrics data with its historical metrics data. If the current metrics data deviates significantly from the historical metrics data, the probe node is considered to be in an abnormal state.
[0134] The scheduler determines the status of a probe node based on the heartbeat data sent by the probe node. By having the probe node periodically report heartbeat data to the scheduler, the scheduler monitors the status of the probe node and the execution of probe tasks in real time, promptly detecting failures of the probe node or abnormal probe tasks, thereby improving the stability of the probe cluster.
[0135] For a more detailed description of the scheduler 301, detector 302, decision maker 303, aggregator 304, actuator 305 and transceiver 306 mentioned above, please refer to the relevant descriptions in the foregoing method embodiments.
[0136] The following is combined with Figure 4 The diagram illustrates the process by which the health monitoring center obtains health data of the business cluster across different time periods using a sliding window approach. Detector 1 and Detector 2 each slide M windows to monitor multiple business nodes (e.g., business node 1 and business node 2) across M time periods. One detector monitors multiple business nodes, and one business node is monitored by multiple detectors. Within each time period, a detector can perform multiple monitoring tasks on a business node. The monitoring result for that business node within that time period is determined based on the results of these multiple monitoring tasks. For example, if the number of successful monitoring tasks exceeds a first threshold, the monitoring result for that business node within that time period is considered successful. This process yields the monitoring results for each business node across different time periods. The decision-maker uses a sliding window approach to make decisions based on the monitoring results of multiple detectors across different time periods, obtaining the health data of the business node within each time period. Next, the aggregator uses a sliding M window to aggregate the health data of all business nodes from N decision-makers across different time periods, obtaining the health data of the business cluster across M time periods, where N and M are positive integers. It should be noted that... Figure 4 For illustrative purposes only, the service nodes that detector 1 and detector 2 are responsible for detecting may be the same or different.
[0137] This application also provides a communication device, including at least one processor, which executes programs or instructions stored in a memory to enable the device to perform the above-described functions. Figure 2 The steps performed by the probing center in the health check method shown.
[0138] This application also provides another communication device, including at least one logic circuit and an input / output interface; the logic circuit is used to perform the above-described... Figure 2 The steps performed by the probing center in the health check method shown.
[0139] It should be understood that the division of units (such as schedulers, detectors, decision-makers, aggregators, actuators, and transceivers) in the above-mentioned device (such as a liveness detection center) is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely through software calls via processing elements; all units can be implemented entirely in hardware; or some units can be implemented through software calls via processing elements, while others can be implemented in hardware. For example, each unit can be a separately established processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, called and executed by a processing element within the device. Moreover, these units can be fully or partially integrated together, or they can be implemented independently. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through integrated logic circuits in the processor element or through software calls via processing elements.
[0140] It is worth noting that, for the sake of simplicity, the above method embodiments are described as a series of actions. However, those skilled in the art should know that this application is not limited to the order of the described actions. Furthermore, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0141] Other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the scope of protection of this application. Furthermore, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.
[0142] This application embodiment also provides a health check system, which can be as follows: Figure 1 The system shown includes a load balancing cluster, a probe cluster, and a service cluster. The probe cluster communicates with the load balancing cluster and the service cluster via the network. The probe cluster and the load balancing cluster are different clusters.
[0143] The load balancing cluster is used to send health query requests to the health monitoring cluster. The health query request is used to obtain the health data of the first business node. In response to the health query request, the health data of the first business node sent by the health monitoring cluster is received.
[0144] The detection cluster is used to obtain a task view, which includes a first association between multiple first detection nodes in the detection cluster and a first business node in the business cluster. Based on the first association, multiple first detection nodes are configured to probe the first business node, obtaining multiple first probe results for the first business node. Among these, one of the multiple first detection nodes is used to probe multiple first business nodes. The multiple first probe results are aggregated to obtain the health data of the first business node. The health data of the first business node is sent to the load balancing cluster. The business cluster is used to receive multiple probe requests sent by the detection cluster, and in response to the multiple probe requests, sends multiple first probe results for the multiple probe requests. The multiple first probe results are used to determine the health data of the first business node.
[0145] The health check system is used to perform the various steps of the aforementioned large-scale cluster health check method and achieve the same technical effect.
[0146] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 5 As shown, the computing device 500 includes a processor 501, a memory 502, a communication interface 503, and a bus 504. The processor 501, memory 502, and communication interface 503 are coupled via the bus. The memory 502 stores instructions. When the instructions in the memory 502 are executed, the computing device 500 performs the method described in the foregoing embodiments, that is, the computing device 500 is used to implement the functions of a load balancer, a health monitoring center, or a business node in the health check system.
[0147] The computing device 500 may be one or more integrated circuits configured to implement the methods described above, such as: one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these forms of integrated circuits. Furthermore, when the units in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units may be integrated together to implement a system-on-a-chip (SOC).
[0148] Processor 501 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0149] Memory 502 can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0150] The memory 502 stores executable program code, and the processor 501 executes the executable program code to implement the functions of the aforementioned units or modules, thereby realizing the aforementioned large-scale cluster health check method. That is, the memory 502 stores instructions for executing the aforementioned large-scale cluster health check method.
[0151] The communication interface 503 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 500 and other devices or communication networks.
[0152] In addition to the data bus, the 504 bus can also include a power bus, a control bus, and a status signal bus. The bus can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. The bus can be divided into address bus, data bus, and control bus.
[0153] Please see Figure 6 , Figure 6 This is a schematic diagram of a computing device cluster provided in an embodiment of this application. Figure 6 As shown, the computing device cluster 600 includes at least one computing device 500. The memory 502 of one or more computing devices 500 in the computing device cluster 600 may store the same instructions for performing the aforementioned health check method for large-scale clusters.
[0154] In some possible implementations, the memory 502 of one or more computing devices 500 in the computing device cluster 600 may also store partial instructions for executing the aforementioned large-scale cluster health check method. In other words, a combination of one or more computing devices 500 can jointly execute the instructions for executing the aforementioned large-scale cluster health check method.
[0155] It should be noted that the memory 502 in different computing devices 500 within the computing device cluster 600 can store different instructions, each used to execute a portion of the functions of the aforementioned detection center. That is, the instructions stored in the memory 502 of different computing devices 500 can implement the functions of one or more modules among the scheduler, detector, decision-maker, aggregator, executor, and transceiver.
[0156] In some possible implementations, one or more computing devices 500 in the computing device cluster 600 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.
[0157] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating the network connection of computer devices in a computer cluster, as provided in an embodiment of this application. Figure 7 As shown, the two computing devices 500A and 500B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0158] In one possible implementation, the memory in computing device 500A stores instructions for executing scheduler and detector functions. Meanwhile, the memory in computing device 500B stores instructions for executing decision-maker, aggregator, executor, and transceiver functions.
[0159] It should be understood that Figure 7 The functions of computing device 500A shown can also be performed by multiple computing devices. Similarly, the functions of computing device 500B can also be performed by multiple computing devices.
[0160] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the method performed by the detection center in the foregoing embodiments.
[0161] In another embodiment of this application, a computer program product is also provided, which includes computer-executable instructions stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device performs the method executed by the exploration center in the foregoing embodiments.
[0162] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0163] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.
[0164] In the embodiments of this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, and in the various implementation methods / methods / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various implementation methods / methods / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various implementation methods / methods / implementations within each embodiment can be combined according to their inherent logical relationships to form new embodiments, implementation methods, methods, or implementation approaches. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0165] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0166] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0167] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for health checks on large-scale clusters, characterized in that, The method is applied to a liveness detection center, which runs on a liveness detection cluster. The liveness detection cluster communicates with a load balancing cluster and a service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. Obtain a task view, which includes a first association between multiple first detection nodes in the detection cluster and a first business node in the business cluster; Based on the first association relationship, the plurality of first detection nodes are configured to detect the first service node, and a plurality of first detection results of the first service node are obtained, wherein one of the plurality of first detection nodes is used to detect a plurality of the first service nodes. By aggregating the multiple first detection results, the health data of the first service node is obtained; Send the health data of the first service node to the load balancing cluster.
2. The method according to claim 1, characterized in that, The task view includes a second association relationship between multiple second detection nodes in the detection cluster and second service nodes in the service cluster; the method further includes: Based on the second association relationship, the plurality of second detection nodes are configured to detect the second service node, and a plurality of second detection results of the second service node are obtained; By aggregating the multiple second detection results, the health data of the second service node is obtained; The health data of the business cluster is obtained by aggregating the health data of the first business node and the health data of the second business node.
3. The method according to claim 2, characterized in that, Prior to obtaining the task view, the method further includes: Receive a first health query request sent by the load balancing cluster, the first health query request including the node identifier of the first service node and / or the node identifier of the second service node; The process of obtaining the task view includes: The task view is determined based on the number of active nodes in the active cluster, the node identifier of the first service node, and / or the node identifier of the second service node.
4. The method according to any one of claims 2 to 3, characterized in that, The method further includes: Receive a second health query request sent by the load balancing cluster, the second health query request including the node identifier of the second service node; The health data of the second service node is determined based on the node identifier of the second service node and the health data of the service cluster. Send the health data of the second service node to the load balancing cluster.
5. The method according to any one of claims 1 to 4, characterized in that, Based on the first association relationship, the plurality of first detection nodes are configured to probe the first service node, and multiple first detection results of the first service node are obtained, including: Based on the first association relationship, configure the multiple first detection nodes to send multiple first detection requests to the first service node; In response to the plurality of first probe requests, receive the plurality of first probe results sent by the first service node.
6. The method according to any one of claims 1 to 5, characterized in that, The task view also includes a detection strategy for probing the first service node and / or the second service node. The detection strategy includes a first duration and / or a first threshold and / or protocols supported by the detection request when probing the first service node and / or the second service node. The protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Security Protocol (HTTPS), and service-based protocols. The first duration is the duration of probing the first service node and / or the second service node. The first threshold is the upper limit of the number of failed probes and / or the upper limit of the number of successful probes for the first service node and / or the second service node.
7. The method according to any one of claims 1 to 6, characterized in that, The aggregation of the multiple first detection results to obtain the health data of the first service node includes: The health data of the first service node is obtained by making decisions based on the multiple first detection results through a weighted and / or voting mechanism.
8. The method according to any one of claims 6 to 7, characterized in that, The method further includes: By using a sliding window, the first detection node is configured to perform multiple detections on the first service node during the second duration, thereby obtaining multiple third detection results corresponding to the first service node during the second duration. The size of the window is the first duration, and the second duration is greater than or equal to the first duration. By aggregating the results of the multiple third detections, the health data of the first service node corresponding to the second duration is obtained.
9. A detection center, characterized in that, The activity detection center runs on an activity detection cluster, which communicates with a load balancing cluster and a service cluster via a network. The activity detection cluster and the load balancing cluster are different clusters. The activity detection center includes: A scheduler is used to obtain a task view, which includes a first association relationship between multiple first detection nodes in the detection cluster and a first service node in the service cluster. The detector is configured to detect the first service node based on the first association relationship, and obtain multiple first detection results of the first service node, wherein one of the multiple first detection nodes is used to detect multiple first service nodes. A decision-maker, which aggregates the multiple first detection results to obtain the health data of the first service node; A transceiver, used to send health data of the first service node to the load balancing cluster.
10. The detection center according to claim 9, characterized in that, The task view includes a second association between multiple second detection nodes in the detection cluster and second business nodes in the business cluster. The detection center also includes an aggregator. The detector is also used to configure the plurality of second detection nodes to detect the second service node based on the second association relationship, and to obtain a plurality of second detection results of the second service node; The decision-maker is also used to aggregate the multiple second detection results to obtain the health data of the second service node; The aggregator is used to aggregate the health data of the first business node and the health data of the second business node to obtain the health data of the business cluster.
11. The detection center according to claim 10, characterized in that, Before obtaining the task view, the transceiver is also used to receive a first health query request sent by the load balancing cluster, the first health query request including the node identifier of the first service node and / or the node identifier of the second service node; The scheduler is also used to determine the task view based on the number of liveness detection nodes in the liveness detection cluster, the node identifier of the first service node, and / or the node identifier of the second service node.
12. The liveness detection center according to any one of claims 10 to 11, characterized in that, The liveness detection center also includes actuators. The transceiver is also used to receive a second health query request sent by the load balancing cluster, the second health query request including the node identifier of the second service node; The actuator is used to determine the health data of the second service node based on the node identifier of the second service node and the health data of the service cluster; The transceiver is also used to send health data of the second service node to the load balancing cluster.
13. The liveness detection center according to any one of claims 9 to 12, characterized in that, The detector is also configured to send multiple first detection requests to the first service node based on the first association relationship; The transceiver is also configured to receive multiple first probe results sent by the first service node in response to the multiple first probe requests.
14. The liveness detection center according to any one of claims 9 to 13, characterized in that, The task view also includes a detection strategy for probing the first service node and / or the second service node. The detection strategy includes a first duration and / or a first threshold and / or protocols supported by the detection request when probing the first service node and / or the second service node. The protocols include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Security Protocol (HTTPS), and service-based protocols. The first duration is the duration of probing the first service node and / or the second service node. The first threshold is the upper limit of the number of failed probes and / or the upper limit of the number of successful probes for the first service node and / or the second service node.
15. The liveness detection center according to any one of claims 9 to 14, characterized in that, The decision-maker is also used to make decisions on the multiple first detection results through a weighted and / or voting mechanism to obtain the health data of the first service node.
16. The liveness detection center according to any one of claims 14 to 15, characterized in that, The detector is also configured to use a sliding window to configure the first detection node to perform multiple detections on the first service node during the second duration, thereby obtaining multiple third detection results corresponding to the first service node during the second duration. The size of the window is the first duration, and the second duration is greater than or equal to the first duration. The decision-maker is also used to aggregate the multiple third detection results to obtain the health data of the first service node corresponding to the second duration.
17. A health check system, characterized in that, It includes a load balancing cluster, a liveness detection cluster, and a service cluster. The liveness detection cluster communicates with the load balancing cluster and the service cluster via a network. The liveness detection cluster and the load balancing cluster are different clusters. The load balancing cluster is used to send a health query request to the detection cluster. The health query request is used to obtain the health data of the first service node. In response to the health query request, the health data of the first service node sent by the detection cluster is received. The activity detection cluster is used to obtain a task view, which includes a first association relationship between multiple first activity detection nodes in the activity detection cluster and a first business node in the business cluster. Based on the first association relationship, the plurality of first detection nodes are configured to detect the first service node, and a plurality of first detection results of the first service node are obtained. Among them, one of the plurality of first detection nodes is used to detect the first service node. The plurality of first detection results are aggregated to obtain the health data of the first service node. The health data of the first service node is sent to the load balancing cluster. The service cluster is used to receive multiple probe requests sent by the probe cluster, and in response to the multiple probe requests, send multiple first probe results for the multiple probe requests, the multiple first probe results being used to determine the health data of the first service node.
18. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 8.
19. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 8.
20. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 8.