Method and system for realizing high availability of main node of Kubernetes cluster
By configuring the reverse proxy and custom health check module in the Nginx server, real-time health monitoring and failure isolation of kube-apiserver instances in the Kubernetes cluster master nodes has been solved, and the stability and availability of the master nodes in the existing technology has been significantly improved.
Patent Information
- Application Number
- CN202510184319.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
The high availability scheme of existing Kubernetes clusters is difficult to detect health status in real time when there are abnormalities such as IO failures in the master node, resulting in long-term inability to operate the client, affecting service interruptions and workload migration, and affecting user experience and business continuity.
By configuring reverse proxy, custom health check module, health status detection, client request proxy and performance monitoring in the Nginx server, real-time health monitoring and fault isolation of back-end kube-apiserver instances in the Kubernetes cluster master node.
It significantly improves the stability and availability of the cluster, enhances the real-time nature of fault discovery and processing, supports the expansion and complexity of cluster size, and ensures the continuity of service and user experience.
Smart Images

Figure CN120075026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud native technology, and specifically to a method and system for achieving high availability of a master node in a Kubernetes cluster. Background Art
[0002] With the continuous development and popularization of cloud computing technology, containerization technology has become an important means of building, deploying and managing applications. In this context, Kubernetes has rapidly emerged as a leader in the field of container orchestration with its powerful container orchestration capabilities. However, in the deployment and operation and maintenance of Kubernetes clusters, how to ensure the high availability of the cluster has always been an important issue that needs attention. Kubernetes clusters usually consist of multiple nodes, among which the master node is responsible for the management and control of the entire cluster. Specifically, the master node contains core components such as kube-apiserver, kube-controller-manager and kube-scheduler, which are responsible for providing API services, managing resources in the cluster and scheduling Pods. Since these components play a vital role in the stability and operating efficiency of the cluster, once the master node fails, it will have a serious impact on the entire cluster.
[0003] Traditional Kubernetes cluster high availability solutions are usually implemented by redundant deployment of multiple master nodes and load balancing. In this solution, multiple master nodes are connected through a load balancer, and client requests are randomly distributed to one of the master nodes for processing. However, this solution has some limitations and problems in actual applications. When a long connection (such as HTTP / 2) is used between nginx and kube-apiserver, it is difficult for nginx to detect changes in the health status of the server in real time in some cases, especially when anomalies such as IO failures occur in the master node, causing the client (such as kubelet) to be unable to operate for a long time, and even triggering the node NotReady state, causing service interruption and workload migration, affecting user experience and business continuity, and affecting the stability of the entire cluster.
[0004] How to improve the high availability, stability, and usability of the Kubernetes cluster is a technical problem that needs to be solved. Summary of the invention
[0005] The technical task of the present invention is to address the above shortcomings and provide a method and system for achieving high availability of the master node of the Kubernetes cluster to solve the technical problem of how to improve the high availability, stability and availability of the Kubernetes cluster.
[0006] In a first aspect, a method for implementing high availability of the master node of a Kubernetes cluster includes the following steps:
[0007] Configure Nginx reverse proxy: Deploy an Nginx server as a reverse proxy in the Kubernetes cluster, configure a load balancing policy, client request forwarding rules, and routing parameters of the backend kube-apiserver instances in the Nginx server, and build a kube-apiserver server list based on all backend kube-apiserver instances;
[0008] Customize the health check mechanism: Compile and integrate a custom health detection module in the Nginx server. The health detection module is used to perform health detection on the backend kube-apiserver instances based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instances. The health status of the backend kube-apiserver instances includes two types: normal and unavailable;
[0009] Health status detection: Read the kube-apiserver server list configured in the Nginx server, regularly detect the health status of each backend kube-apiserver instance through the health detection module, and update the health status information of the backend kube-apiserver instances;
[0010] Client request proxy: After the Nginx server listens to the client request, proxy the client request to a backend kube-apiserber instance with a normal health status based on the health status of all current backend kube-apiserver instances, the load balancing policy, and the client request forwarding rules;
[0011] Performance monitoring: For the client requests listened to, collect the metrics related to the client requests through the Nginx server, perform performance analysis on the backend kube-apiserver instances based on the metrics and the health status detected by the health check module, and perform warning processing on the backend kube-apiserver instances with performance analysis results lower than the threshold and unavailable.
[0012] Preferably, the health detection module is used to allocate a detection subprocess for all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters;
[0013] Among them, the instance detection parameters include detection method, detection time interval, detection timeout, number of failure retries, number of successful detections, and detection path. The detection method includes two methods: tcp and http. When configuring the detection method as http, configure the detection path, and the detection path is the path in the client request of the http method. The detection timeout refers to the timeout for waiting for the kube-apiserver instance to respond to the request when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance performs health detection. If there is no response within the specified detection timeout, the detection result of this time is determined to be a non-alive state; the number of failure retries refers to the number of times when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance reaches the number of failure retries when the detection results are all non-alive states for several consecutive times, and marks the health status of the current kube-apiserver instance as unavailable. The setting of the number of failure retries can avoid the jitter of the service health status caused by the failure of the kube-apiserver instance to respond to the health detection subprocess of the health detection module in a timely manner when there are many requests; the number of successful detections refers to the number of consecutive successful times when the subprocess of the health detection module responsible for detecting the health status of the backend kube-apiserver instance returns healthy as expected by the kube-apiserver.
[0014] Preferably, during health status detection, initialize the health check module, read the kube-apiserver server list through the health check module, poll the health status of the backend kube-apiserver instance according to the configured detection time interval and detection method, continuously monitor the health status of the backend kube-apiserver instance. After reaching the number of failure retries for detection, immediately mark the server status as unavailable and update the health status information of the backend kube-apiserver instance. Among them, the session_drop instruction can be configured in the health detection module. The session_drop instruction can accept the values on and off. When the value is on, it means to detect the health status of the established downstream server in real time in the is_alive method of the health detection module. When the value is off, it means not to detect the health status of the established downstream server in the is_alive method of the health detection module, and both return 1.
[0015] Preferably, the client request proxy includes the following steps:
[0016] After the Nginx server listens to the client request through its server module, it parses the client request to obtain the kube-apiserver instance corresponding to the client request, and forwards the client request to the relevant upstream module;
[0017] After the upstream module receives a client request, it calls the health detection module to detect the health status of the kube-apiserver instance corresponding to the client request at the backend. If the health status of the kube-apiserver instance corresponding to the client request is normal, the upstream module forwards the client request to the corresponding kube-apiserver instance at the backend. If the health status of the kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding kube-apiserver instance at the backend and redirects the client request. When redirecting the client request, based on the configured load balancing policy, a normal kube-apiserver instance at the backend is selected as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
[0018] Preferably, during performance monitoring, the metrics of the health detection module are set in Prometheus, and an alarm is set when the number of healthy kube-apiserver instances at the available backend is lower than the threshold. The Nginx server collects the metrics related to the client request. Based on the metrics and the health status detected by the health check module, performance analysis is performed on the kube-apiserver instances at the backend, and an alarm is issued when the number of healthy kube-apiserver instances at the available backend is lower than the threshold.
[0019] Among them, the metrics include the number of established connections, the number of processed requests, the CPU usage rate of the server, and the memory usage of each kube-apiserver instance at the backend.
[0020] In a second aspect, a system for implementing high availability of the master node of a Kubernetes cluster according to the present invention includes a client, an Nginx server, and a backend server;
[0021] The backend server is used to provide kube-apiserver instances at the backend;
[0022] The Nginx server is deployed as a reverse proxy in the Kubernetes cluster. In the Nginx server, the load balancing policy, client request forwarding rules, and routing parameters of the backend kube-apiserver instances are configured, and a kube-apiserver server list composed of all backend kube-apiserver instances is built. A custom health detection module is compiled and integrated in the Nginx server. The health detection module is used to perform health detection on the backend kube-apiserver instances based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instances. The health status of the backend kube-apiserver instances includes two types: normal and unavailable. The health detection module is used to read the kube-apiserver server list configured in the Nginx server, regularly detect the health status of each backend kube-apiserver instance through the health detection module, and update the health status information of the backend kube-apiserver instances.
[0023] The client is used to access each component of the kube-apiserver instance of the Kubernetes cluster master node through the Nginx server. Each component includes kubelet, kube-proxy, kube-scheduler, and kube-controller-manager in the Kubernetes cluster. Various components establish long connections with the backend kube-apiserver instances in the HTTP2 manner.
[0024] After the Nginx server listens to the client request, it is used to proxy the client request to a backend kube-apiserber instance with a normal health status based on the health status of all current backend kube-apiserver instances, the load balancing policy, and the client request forwarding rules.
[0025] For the listened client requests, the Nginx server is used to collect the metrics related to the client requests, perform performance analysis on the backend kube-apiserver instances based on the metrics and the health status detected by the health check module, and perform alarm processing on the backend kube-apiserver instances with performance analysis results lower than the threshold and unavailable.
[0026] Preferably, the health detection module is used to allocate a detection subprocess for all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters.
[0027] Among them, the instance detection parameters include detection method, detection time interval, detection timeout, number of failure retries, number of successful detections, and detection path. The detection methods include two types: tcp and http. When configuring the detection method as http, the detection path needs to be configured. The detection path is the path in the client request of the http method. The detection timeout refers to the timeout for waiting for the kube-apiserver instance to respond to the request when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance is performing health detection. If there is no response within the specified detection timeout, the detection result of this time is determined to be a non-alive state; the number of failure retries refers to the number of times when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance reaches the number of failure retries when the detection results are all non-alive states for several consecutive times, and the health status of the current kube-apiserver instance is marked as unavailable. The setting of the number of failure retries can avoid the jitter of the service health status caused by the failure of the kube-apiserver instance to respond to the health detection subprocess of the health detection module in a timely manner when there are many requests; the number of successful detections refers to the number of consecutive successful times when the subprocess of the health detection module responsible for detecting the health status of the backend kube-apiserver instance returns healthy as expected by the kube-apiserver.
[0028] Preferably, during health status detection, the Nginx server is used to initialize the health check module. By reading the kube-apiserver server list within the health check module, it polls the health status of the backend kube-apiserver instance according to the configured detection time interval and detection method, continuously monitors the health status of the backend kube-apiserver instance. After reaching the number of failure retries, it immediately marks the server status as unavailable and updates the health status information of the backend kube-apiserver instance. Among them, the session_drop directive can be configured in the health detection module. The session_drop directive can accept the values of on and off. When the value is on, it means to detect the health status of the established downstream server in real time in the is_alive method of the health detection module. When the value is off, it means not to detect the health status of the established downstream server in the is_alive method of the health detection module, and both return 1.
[0029] Preferably, after the Nginx server listens to the client request through its server module, it parses the client request to obtain the kube-apiserver instance corresponding to the client request and forwards the client request to the relevant upstream module;
[0030] After the upstream module receives a client request, it calls the health detection module to detect the health status of the kube-apiserver instance corresponding to the client request at the backend. If the health status of the kube-apiserver instance corresponding to the client request is normal, the client request is forwarded to the corresponding kube-apiserver instance at the backend. If the health status of the kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding kube-apiserver instance at the backend and redirects the client request. When redirecting the client request, based on the configured load balancing policy, a normal kube-apiserver instance at the backend is selected as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
[0031] Preferably, Prometheus is set to collect the metrics of the health detection module and has the following rules: issue an alarm when the number of healthy kube-apiserver instances at the available backend is lower than the threshold; the Nginx server is used to collect the metrics related to client requests, perform performance analysis on the kube-apiserver instances at the backend based on the metrics and the health status detected by the health check module, and issue an alarm when the number of healthy kube-apiserver instances at the available backend is lower than the threshold;
[0032] Among them, the metrics include the number of established connections, the number of processed requests, the CPU usage rate of the server, and the memory usage of each kube-apiserver instance at the backend.
[0033] The method and system for implementing the high availability of the master node of the Kubernetes cluster of the present invention have the following advantages:
[0034] 1. Significantly improve the stability and availability of the cluster: Through the real-time monitoring and detection of the kube-apiserver instances at the backend in the Kubernetes master node by the health monitoring module, potential faults can be quickly discovered and isolation measures can be taken immediately to ensure that the cluster is always in a healthy and stable state. This real-time health check and fault isolation mechanism greatly improves the overall availability and stability of the cluster;
[0035] 2. Enhance the real-time nature of fault discovery and handling: The health detection module can detect the health status of the backend kube-apiserver instances in real time. Once a fault is detected, it will immediately disconnect from that instance and forward requests to other healthy backend kube-apiserver instances. This real-time fault discovery and handling capability greatly reduces the impact of faults on the stability and availability of the cluster, ensuring service continuity;
[0036] 3. Support the expansion of cluster scale and increased complexity: As the scale of the Kubernetes cluster expands and its complexity increases, the requirements for high availability also become higher and higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] The present invention will be further described below with reference to the drawings.
[0039] Figure 1 It is a flowchart of a method for implementing high availability of the master node of a Kubernetes cluster in Embodiment 1;
[0040] Figure 2 It is a block diagram of the working principle of a method for implementing high availability of the master node of a Kubernetes cluster in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the embodiments cited are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0042] The embodiments of the present invention provide a method and system for implementing high availability of the master node of a Kubernetes cluster, which are used to solve the technical problem of how to improve the high availability, stability and availability of the Kubernetes cluster.
[0043] Embodiment 1:
[0044] A method for implementing high availability of the master node of a Kubernetes cluster according to the present invention includes five steps: configuring Nginx reverse proxy, customizing a health check mechanism, health status detection, client request proxy, and performance monitoring.
[0045] Step S100 configures Nginx reverse proxy: Deploy an Nginx server as a reverse proxy in the Kubernetes cluster. Configure the load balancing policy, client request forwarding rules, and routing parameters of the backend kube-apiserver instances in the Nginx server, and build a kube-apiserver server list based on all backend kube-apiserver instances.
[0046] In this embodiment, an nginx server is configured as a reverse proxy. When using the nginx reverse proxy to proxy the kube-apiserver instances of the master node in the Kubernetes cluster, the four-layer forwarding ability of the nginx server is used. Configure the upstream parameters of the backend kube-apiserver instances in the stream module of the nginx server. The configurable items include the load balancing policy, the IP address and port of the backend server.
[0047] Step S200 customizes the health check mechanism: Compile and integrate a custom health detection module in the Nginx server. The health detection module is used to perform health detection on the backend kube-apiserver instances based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instances. The health status of the backend kube-apiserver instances includes two types: normal and unavailable.
[0048] In this embodiment, the health detection module is used to allocate a detection subprocess for all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters.
[0049] Among them, the instance detection parameters include detection method, detection time interval, detection timeout, number of failure retries, number of successful detections, and detection path. The detection method includes two methods: tcp and http. When configuring the detection method as http, the detection path needs to be configured. The detection path is the path in the client request of the http method. The detection timeout refers to the timeout for waiting for the kube-apiserver instance to respond to the request when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance performs health detection. If there is no response within the specified detection timeout, the detection result of this time is determined to be a non-surviving state. The number of failure retries refers to the number of times when the subprocess of the health detection module for detecting the health status of the backend kube-apiserver instance marks the current health status of the kube-apiserver instance as unavailable when the number of consecutive detection results of non-surviving states reaches the number of failure retries. Setting the number of failure retries can avoid the jitter of the service health status caused by the kube-apiserver instance not responding to the health detection subprocess of the health detection module in time when there are many requests. The number of successful detections refers to the consecutive successful times when the subprocess of the health detection module responsible for detecting the health status of the backend kube-apiserver instance returns health as expected by the kube-apiserver.
[0050] In this embodiment, a custom health detection module nginx_dynamic_healthcheck is compiled and integrated into the nginx server. Instance detection parameters of the backend kube-apiserver instance are configured in the upstream module of the nginx server. The instance detection parameters include detection method, detection time interval, detection timeout, number of failure retries, number of successful detections, and detection path. The detection method includes two methods: tcp and http. When the detection method is configured as the http method, the detection path needs to be configured. The detection path is the path in the http request. For example, the detection path for configuring the health status of the healthy kube-apiserver instance in this method is / healthz. The detection timeout refers to the timeout for waiting for the kube-apiserver instance to respond to the request when the subprocess of nginx_dynamic_healthcheck for detecting the health status of the kube-apiserver instance performs a health check. If there is no response within the specified detection timeout, the detection result of this time is a non-surviving state. The number of failure retries refers to the number of times when the subprocess of nginx_dynamic_healthcheck for detecting the health status of the kube-apiserver instance marks the current health status of the kube-apiserver instance as unavailable when the number of consecutive non-surviving detection results reaches the number of failure retries. The setting of the number of failure retries can avoid the service health status jitter caused by the kube-apiserver instance not responding to the health check subprocess of the nginx_dynamic_healthcheck module in time when there are many requests. The number of successful detections refers to the number of consecutive successful detections when the subprocess of the nginx_dynamic_healthcheck module responsible for detecting the health status of the kube-apiserver instance returns healthy as expected.
[0051] Step S300 Health status detection: Read the list of kube-apiserver servers configured in the Nginx server, and regularly detect the health status of each backend kube-apiserver instance through the health detection module and update the health status information of the backend kube-apiserver instance.
[0052] In this embodiment, during the health status detection, the health check module is initialized. The list of kube-apiserver server instances is read from the health check module, and the health status of the backend kube-apiserver instances is polled according to the configured detection time interval and detection method, continuously monitoring the health status of the backend kube-apiserver instances. After reaching the detection failure retry count, the server status is immediately marked as unavailable, and the health status information of the backend kube-apiserver instances is updated. Among them, the session_drop instruction can be configured in the health detection module. The session_drop instruction can accept the values on and off. When the value is on, it means that the health status of the established downstream server is detected in real time in the is_alive method of the health detection module. When the value is off, it means that the health status of the established downstream server is not detected in the is_alive method of the health detection module, and 1 is returned in both cases.
[0053] Step S400: Client request proxying: After the Nginx server listens to the client request, the client request is proxied to a backend kube-apiserber instance with a normal health status according to the health status of all current backend kube-apiserver instances, based on the load balancing policy, and the client request forwarding rule.
[0054] In this embodiment, the client request proxying includes the following steps:
[0055] (1) After the Nginx server listens to the client request through its server module, the client request is parsed to obtain the kube-apiserver instance corresponding to the client request, and the client request is forwarded to the relevant upstream module;
[0056] (2) After the upstream module receives a client request, it calls the health check module to detect the health status of the kube-apiserver instance corresponding to the client request. If the health status of the kube-apiserver instance corresponding to the client request is normal, the upstream module forwards the client request to the corresponding kube-apiserver instance. If the health status of the kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding kube-apiserver instance and redirects the client request. When redirecting the client request, based on the configured load balancing policy, a normal kube-apiserver instance is selected as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
[0057] In this embodiment, the upstream module of the Nginx server processes client requests to the kube-apiserver instance. When the upstream module processes a long connection request from the client, the client sends request data to nginx. Before the upstream module of nginx forwards the request to the backend service kube-apiserver instance, it calls the is_alive method of the nginx_dynamic_healthcheck module. The call parameters include the socket information of the kube-apiserver instance. Whether to forward the request is determined according to the return result of the is_alive method. When the return value of the is_alive method is 1, the upstream forwards the request to the corresponding kube-apiserver instance. When the return value of the is_alive method is 0, the upstream releases this connection resource and redirects the request to other normal kube-apiserver instances. When redirecting the request, the configured load balancing policy is executed, and a healthy kube-apiserver instance is selected from the surviving kube-apiserver instances according to the policy, a connection is established with it, and the client request is forwarded to this connection.
[0058] Step S500 Performance Monitoring: For the monitored client requests, relevant metrics of the client requests are collected through the Nginx server, and performance analysis is performed on the backend kube-apiserver instances based on the metrics and the health status detected by the health check module. Alarm processing is performed on the backend kube-apiserver instances with performance analysis results lower than the threshold and unavailable.
[0059] In this embodiment, during performance monitoring, set the metrics of the health detection module in Prometheus, and set an alarm when the number of healthy backend kube-apiserver instances available is lower than the threshold. Collect the metrics related to client requests through the Nginx server, and perform performance analysis on the backend kube-apiserver instances based on the metrics and the health status detected by the health check module. Alarm when the number of healthy backend kube-apiserver instances available is lower than the threshold based on the performance analysis results; where the metrics include the number of established connections, the number of processed requests, the CPU usage rate of the server, and the memory usage of each backend kube-apiserver instance.
[0060] The method of this embodiment combines the reverse proxy function of Nginx and the real-time detection ability of nginx_dynamic_healthcheck to achieve real-time health check and fault isolation of the kube-apiserver instances in the Kubernetes cluster master node. This real-time health check and fault isolation mechanism can ensure that each client request is always forwarded to an available kube-apiserver instance, thus improving the stability and availability of the Kubernetes cluster.
[0061] Embodiment 2:
[0062] A system for implementing high availability of the master node of a Kubernetes cluster includes a client, an Nginx server, and a backend server.
[0063] The backend server is used to provide backend kube-apiserver instances.
[0064] The Nginx server is deployed as a reverse proxy in the Kubernetes cluster. The Nginx server is configured with a load balancing policy, client request forwarding rules, and routing parameters for the backend kube-apiserver instances. A kube-apiserver server list composed of all backend kube-apiserver instances is built. A custom health detection module is compiled and integrated in the Nginx server. The health detection module is used to perform health detection on the backend kube-apiserver instances based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instances. The health status of the backend kube-apiserver instances includes two types: normal and unavailable. The health detection module is used to read the kube-apiserver server list configured in the Nginx server, regularly detect the health status of each backend kube-apiserver instance through the health detection module, and update the health status information of the backend kube-apiserver instances.
[0065] In this embodiment, the nginx server is configured as a reverse proxy. When using the nginx reverse proxy to access the kube-apiserver instances of the master node in the Kubernetes cluster, the four-layer forwarding capability of the nginx server is used, and the upstream parameters of the backend kube-apiserver instances are configured in the stream module of the nginx server. Among them, the configurable items include the load balancing policy, the IP address and port of the backend server.
[0066] The health detection module is used to allocate a detection subprocess for all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters.
[0067] The client is used to access each component of the kube-apiserver instance of the Kubernetes cluster master node through the Nginx server. Each component includes kubelet, kube-proxy, kube-scheduler, and kube-controller-manager in the Kubernetes cluster. Various components establish long connections with the backend kube-apiserver instances in the HTTP2 manner.
[0068] After the Nginx server listens to the client request, it is used to proxy the client request to a kube-apiserver instance with a normal health status according to the health status of all current backend kube-apiserver instances, the load balancing policy, and the client request forwarding rule.
[0069] For the listened client request, the Nginx server is used to collect the metrics related to the client request, perform performance analysis on the backend kube-apiserver instance based on the metrics and the health status detected by the health check module, and perform alarm processing on the backend kube-apiserver instance with a performance analysis result lower than the threshold and unavailable.
[0070] Among them, the instance detection parameters include the detection method, detection time interval, detection timeout, failure retry times, detection success times, and detection path. The detection method includes two methods: tcp and http. When the detection method is configured as the http method, the detection path is configured. The detection path is the path in the http method client request. The detection timeout refers to the timeout time for waiting for the kube-apiserver instance to respond to the request when the subprocess of the health check module for detecting the health status of the backend kube-apiserver instance performs the health check. If there is no response within the specified detection timeout, the detection result of this time is determined to be a non-surviving state; the failure retry times refer to the number of times when the subprocess of the health check module for detecting the health status of the backend kube-apiserver instance marks the health status of the current kube-apiserver instance as unavailable when the number of consecutive detection results of non-surviving states reaches the failure retry times. The setting of the failure retry times can avoid the service health status jitter caused by the kube-apiserver instance not responding to the health check subprocess of the health check module in time when there are many requests; the detection success times refer to the consecutive success times when the subprocess of the health check module responsible for detecting the health status of the backend kube-apiserver instance returns healthy as expected by the kube-apiserver.
[0071] In this embodiment, a custom health detection module nginx_dynamic_healthcheck is compiled and integrated into the nginx server. Instance detection parameters of the backend kube-apiserver instance are configured in the upstream module of the nginx server. The instance detection parameters include detection method, detection time interval, detection timeout, number of failure retries, number of successful detections, and detection path. The detection method includes two methods: tcp and http. When the detection method is configured as http, the detection path needs to be configured. The detection path is the path in the http request. For example, the detection path for configuring the health status of the healthy kube-apiserver instance in this method is / healthz. The detection timeout refers to the timeout for waiting for the kube-apiserver instance to respond to the request when the subprocess of nginx_dynamic_healthcheck for detecting the health status of the kube-apiserver instance performs a health check. If there is no response within the specified detection timeout, the detection result of this time is a non-surviving state. The number of failure retries means that when the number of consecutive non-surviving detection results of the subprocess of nginx_dynamic_healthcheck for detecting the health status of the kube-apiserver instance reaches the number of failure retries, the current health status of the kube-apiserver instance is marked as unavailable. The setting of the number of failure retries can avoid the service health status jitter caused by the kube-apiserver instance not responding to the health check subprocess of the nginx_dynamic_healthcheck module in time when there are many requests. The number of successful detections refers to the number of consecutive successful times when the subprocess of the nginx_dynamic_healthcheck module responsible for detecting the health status of the kube-apiserver instance returns healthy as expected.
[0072] In this embodiment, when performing health status detection, the Nginx server is used to initialize the health check module. By reading the kube-apiserver server list within the health check module, it polls the health status of the backend kube-apiserver instances at the configured detection time interval and detection method, continuously monitors the health status of the backend kube-apiserver instances. After reaching the detection failure retry count, it immediately marks the server status as unavailable and updates the health status information of the backend kube-apiserver instances. Among them, the session_drop instruction can be configured in the health detection module. The session_drop instruction can accept on and off values. When the value is on, it means to detect the health status of the established downstream server in real time in the is_alive method of the health detection module. When the value is off, it means not to detect the health status of the established downstream server in the is_alive method of the health detection module, and both return 1.
[0073] In this embodiment, when the Nginx server executes the client request proxy, it performs the following operations:
[0074] (1) After the Nginx server listens to the client request through its server module, it parses the client request to obtain the kube-apiserver instance corresponding to the client request, and forwards the client request to the relevant upstream module;
[0075] (2) After the upstream module receives the client request, it calls the health detection module to detect the health status of the backend kube-apiserver instance corresponding to the client request. If the health status of the backend kube-apiserver instance corresponding to the client request is normal, it forwards the client request to the corresponding backend kube-apiserver instance. If the health status of the backend kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding backend kube-apiserver instance, and redirects the client request. When redirecting the client request, based on the configured load balancing policy, it selects a normal backend kube-apiserver instance as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
[0076] In this embodiment, the upstream module of the Nginx server processes the requests from the client to the kube-apiserver instance. When the upstream module processes the long connection requests from the client, the client sends the request data to nginx. Before forwarding the request to the backend service kube-apiserver instance, the upstream module of nginx calls the is_alive method of the nginx_dynamic_healthcheck module. The call parameters include the socket information of the kube-apiserver instance. Whether to forward the request is determined according to the return result of the is_alive method. When the return value of the is_alive method is 1, the upstream forwards the request to the corresponding kube-apiserver instance; when the return value of the is_alive method is 0, the upstream releases the connection resource and redirects the request to other normal kube-apiserver instances. When redirecting the request, the configured load balancing policy is executed, and a healthy kube-apiserver instance is selected according to the policy from the surviving kube-apiserver instances, a connection is established with it, and the client request is forwarded to this connection.
[0077] In this embodiment, during performance monitoring, the metrics of the health detection module are set in prometheus, and an alarm is set when the number of healthy backend kube-apiserver instances available is lower than the threshold. The metrics related to the client requests are collected by the Nginx server. The performance of the backend kube-apiserver instances is analyzed based on the metrics and the health status detected by the health check module. An alarm is given when the number of healthy backend kube-apiserver instances available is lower than the threshold based on the performance analysis results. Among them, the metrics include the number of established connections, the number of processed requests, the CPU usage rate of the server, and the memory usage of each backend kube-apiserver instance.
[0078] The system of this embodiment can implement the high availability of the master node of the Kubernetes cluster by executing the method disclosed in Embodiment 1.
[0079] The above has introduced in detail the method and system for implementing high availability of the master node in a Kubernetes cluster. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for achieving high availability of a Kubernetes cluster master node, characterized in that: The steps include: Configure Nginx reverse proxy: Deploy Nginx server as reverse proxy in the Kubernetes cluster, configure load balancing policy, client request forwarding rules, and routing parameters of the backend kube-apiserver instance in the Nginx server, and build a kube-apiserver server list based on all backend kube-apiserver instances; Custom health check mechanism: Compile and integrate a custom health check module in the Nginx server. The health check module is used to perform health checks on the backend kube-apiserver instance based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instance. The health status of the backend kube-apiserver instance includes normal and unavailable. Health status detection: read the kube-apiserver server list configured in the Nginx server, regularly detect the health status of each backend kube-apiserver instance through the health detection module, and update the health status information of the backend kube-apiserver instance; Client request proxy: After listening to the client request through the Nginx server, the client request is proxied to a backend kube-apiserber instance with a normal health status based on the health status of all current backend kube-apiserver instances, load balancing policies, and client request forwarding rules; Performance monitoring: For monitored client requests, the Nginx server collects metrics related to client requests, performs performance analysis on the backend kube-apiserver instance based on the metrics and the health status detected by the health check module, and issues alarms for backend kube-apiserver instances whose performance analysis results are below the threshold or are unavailable.
2. The method for achieving high availability of a Kubernetes cluster master node according to claim 1, characterized in that: The health detection module is used to assign a detection subprocess to all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters; Among them, the instance detection parameters include detection mode, detection interval, detection timeout, number of failed retries, number of successful detections, and detection path. The detection mode includes TCP and HTTP. When the detection mode is configured as HTTP, the detection path is configured. The detection path is the path in the HTTP client request. The detection timeout refers to the timeout for the kube-apiserver instance to respond to the request when the subprocess of the health detection module detects the health status of the backend kube-apiserver instance. If there is no response within the specified detection timeout, the detection result is determined to be non-survival state; the number of failed retries is It means that when the subprocess of the health detection module that detects the health status of the backend kube-apiserver instance is in a non-survival state for several consecutive times, the number of failed retries reaches the number of failed retries. The health status of the current kube-apiserver instance is marked as unavailable. The setting of the number of failed retries can avoid the jitter of the service health status caused by the failure to respond to the health detection subprocess of the health detection module in time when there are many requests to the kube-apiserver instance; the number of successful detections refers to the number of consecutive successes of the subprocess of the health detection module responsible for detecting the health status of the backend kube-apiserver instance when the kube-apiserver returns health as expected.
3. The method for achieving high availability of a Kubernetes cluster master node according to claim 2, characterized in that: During health status detection, the health check module is initialized, the kube-apiserver server list is read from the health check module, the health status of the backend kube-apiserver instance is polled according to the configured detection time interval and detection method, and the health status of the backend kube-apiserver instance is continuously monitored. After the number of detection failure retries is reached, the server status is immediately marked as unavailable, and the health status information of the backend kube-apiserver instance is updated. The session_drop instruction can be configured in the health detection module. The session_drop instruction can accept on and off values. When the value is on, it means that the health status of the downstream server that has been connected is detected in real time in the is_alive method of the health detection module. When the value is off, it means that the health status of the downstream server that has been connected is not detected in the is_alive method of the health detection module, and both return 1.
4. The method for achieving high availability of a Kubernetes cluster master node according to claim 1, characterized in that: The client request agent includes the following steps: After the Nginx server listens to the client request through its server module, it parses the client request, obtains the kube-apiserver instance corresponding to the client request, and forwards the client request to the relevant upstream module; After receiving the client request, the upstream module calls the health detection module to detect the health status of the backend kube-apiserver instance corresponding to the client request. If the health status of the backend kube-apiserver instance corresponding to the client request is normal, the client request is forwarded to the corresponding backend kube-apiserver instance. If the health status of the backend kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding backend kube-apiserver instance and redirects the client request. When redirecting the client request, based on the configured load balancing policy, a normal backend kube-apiserver instance is selected as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
5. The method for achieving high availability of a Kubernetes cluster master node according to claim 1, characterized in that: When monitoring performance, set the metrics indicators of the health detection module in Prometheus, and set an alarm when the number of available backend kube-apiserver instances is lower than the threshold. Collect metrics indicators related to client requests through the Nginx server, perform performance analysis on the backend kube-apiserver instances based on the metrics indicators and the health status detected by the health check module, and issue an alarm when the number of available backend kube-apiserver instances is lower than the threshold based on the performance analysis results; Among them, the metrics include the number of connections established by each backend kube-apiserver instance, the number of requests processed, the server CPU usage, and the memory usage.
6. A system for achieving high availability of a Kubernetes cluster master node, characterized in that: Including client, Nginx server and backend server; The backend server is used to provide the backend kube-apiserver instance; The Nginx server is deployed in the Kubernetes cluster as a reverse proxy. The Nginx server is configured with load balancing strategies, client request forwarding rules, and routing parameters of the backend kube-apiserver instance, and a kube-apiserver server list consisting of all backend kube-apiserver instances is constructed. A custom health detection module is compiled and integrated in the Nginx server. The health detection module is used to perform health detection on the backend kube-apiserver instance based on predefined instance detection parameters to obtain the health status of the backend kube-apiserver instance. The health status of the backend kube-apiserver instance includes normal and unavailable. The health detection module is used to read the kube-apiserver server list configured in the Nginx server, and regularly detect the health status of each backend kube-apiserver instance through the health detection module, and update the health status information of the backend kube-apiserver instance. The client is used to access the components of the kube-apiserver instance of the master node of the Kubernetes cluster through the Nginx server. The components include kubelet, kube-proxy, kube-scheduler, and kube-controller-manager in the Kubernetes cluster. Various components establish a persistent connection with the backend kube-apiserver instance in HTTP2 mode. After the Nginx server monitors the client request, it is used to proxy the client request to a backend kube-apiserber instance with a normal health status based on the health status of all current backend kube-apiserver instances, load balancing policies, and client request forwarding rules; For monitored client requests, the Nginx server is used to collect metrics related to client requests, perform performance analysis on the backend kube-apiserver instance based on the metrics and the health status detected by the health check module, and issue alarms for backend kube-apiserver instances whose performance analysis results are below the threshold or are unavailable.
7. The system for realizing high availability of master nodes of Kubernetes cluster according to claim 6, characterized in that: The health detection module is used to assign a detection subprocess to all backend kube-apiserver instances and perform health detection on all backend kube-apiserver instances based on predefined instance detection parameters; Among them, the instance detection parameters include detection mode, detection interval, detection timeout, number of failed retries, number of successful detections, and detection path. The detection mode includes TCP and HTTP. When the detection mode is configured as HTTP, the detection path is configured. The detection path is the path in the HTTP client request. The detection timeout refers to the timeout for the kube-apiserver instance to respond to the request when the subprocess of the health detection module detects the health status of the backend kube-apiserver instance. If there is no response within the specified detection timeout, the detection result is determined to be non-survival state; the number of failed retries is It means that when the subprocess of the health detection module that detects the health status of the backend kube-apiserver instance is in a non-survival state for several consecutive times, the number of failed retries reaches the number of failed retries. The health status of the current kube-apiserver instance is marked as unavailable. The setting of the number of failed retries can avoid the jitter of the service health status caused by the failure to respond to the health detection subprocess of the health detection module in time when there are many requests to the kube-apiserver instance; the number of successful detections refers to the number of consecutive successes of the subprocess of the health detection module responsible for detecting the health status of the backend kube-apiserver instance when the kube-apiserver returns health as expected.
8. The system for realizing high availability of master nodes of Kubernetes cluster according to claim 7, characterized in that: During health status detection, the Nginx server is used to initialize the health check module, read the kube-apiserver server list in the health check module, poll the health status of the backend kube-apiserver instance according to the configured detection time interval and detection method, and continuously monitor the health status of the backend kube-apiserver instance. After reaching the number of detection failure retries, the server status is immediately marked as unavailable, and the health status information of the backend kube-apiserver instance is updated. The session_drop instruction can be configured in the health detection module. The session_drop instruction can accept on and off values. When the value is on, it means that the health status of the downstream server that has been connected is detected in real time in the is_alive method of the health detection module. When the value is off, it means that the health status of the downstream server that has been connected is not detected in the is_alive method of the health detection module, and both return 1.
9. The system for realizing high availability of master nodes of Kubernetes cluster according to claim 6, characterized in that: After the Nginx server listens to the client request through its server module, it parses the client request, obtains the kube-apiserver instance corresponding to the client request, and forwards the client request to the relevant upstream module; After receiving the client request, the upstream module calls the health detection module to detect the health status of the backend kube-apiserver instance corresponding to the client request. If the health status of the backend kube-apiserver instance corresponding to the client request is normal, the client request is forwarded to the corresponding backend kube-apiserver instance. If the health status of the backend kube-apiserver instance corresponding to the client request is unavailable, the upstream module releases the connection resources with the corresponding backend kube-apiserver instance and redirects the client request. When redirecting the client request, based on the configured load balancing policy, a normal backend kube-apiserver instance is selected as the target kube-apiserver instance. The upstream module establishes a connection with the target kube-apiserver instance and forwards the client request to the target kube-apiserver instance.
10. The system for realizing high availability of master nodes of Kubernetes cluster according to claim 6, characterized in that: Prometheus is set up with metrics indicators for collecting health detection modules, and the following rules are set: an alarm is issued when the number of available backend kube-apiserver instances is lower than the threshold; the Nginx server is used to collect metrics indicators related to client requests, and perform performance analysis on the backend kube-apiserver instances based on metrics indicators and the health status detected by the health check module. Based on the performance analysis results, an alarm is issued when the number of available backend kube-apiserver instances is lower than the threshold; Among them, the metrics include the number of connections established by each backend kube-apiserver instance, the number of requests processed, the server CPU usage, and the memory usage.