Method and device for realizing Promtheus high availability based on cloud native environment

By dividing multiple Kubernetes clusters in a cloud-native environment and deploying child Prometheus and central Prometheus, the problem of difficult to detect Prometheus failure is solved, and the high availability of monitoring systems and data reliability is achieved.

CN119960910APending Publication Date: 2025-05-09SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510039376.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, Prometheus is difficult to quickly detect and recover after a failure, resulting in the high availability of the monitoring system being limited.

Method used

By dividing multiple Kubernetes clusters in a cloud native environment, select one cluster as the primary cluster and the other clusters as normal clusters. A sub-Prometheus without mount storage is deployed in each normal cluster for self-monitoring and data collection; a central Prometheus with mount storage is deployed in the main cluster for mount storage is used to pull monitoring data of the normal cluster and configure global alarm notifications.

Benefits of technology

It realizes the rapid discovery and recovery of Prometheus failures, ensures high availability of the monitoring system and data reliability, and can respond in a timely manner by receiving alarm notifications when a sub-Prometheus or central Prometheus failure occurs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960910A_ABST
    Figure CN119960910A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for realizing Promtheus high availability based on a cloud native environment, which can solve the problem that Promtheus faults are difficult to find. The invention discloses a method for realizing Promtheus high availability based on a cloud native environment, which comprises the following steps of: dividing a plurality of Kubenes clusters based on a Prometheus federated cluster scheme, and selecting one cluster as a main cluster and other clusters as common clusters; in each common cluster, a single pod is used for deploying a sub Prometheus; wherein the sub Prometheus is not mounted and stored, and the sub Prometheus is used for starting self-monitoring of the corresponding common cluster and collecting monitoring data of the corresponding common cluster; the method comprises the following steps of: deploying a center Prometheus in a main cluster; wherein the central Prometheus is mounted and stored, and the central Prometheus is used for pulling the monitoring data of the sub Prometheus in the common cluster, and the central Prometheus is used for monitoring the data of the sub Prometheus in the common cluster. And a global alarm notification is configured in the main cluster, and when the monitoring data of the central Prometheus or the monitoring data of the sub Prometheus meet an alarm condition, the main cluster generates the alarm notification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and in particular to a method and device for implementing Promtheus high availability based on a cloud native environment. Background Art

[0002] Cloud native is a modern software development and operation methodology that encourages taking advantage of cloud computing to build and run scalable, reliable, and easy-to-manage applications. Kubernetes has become the de facto standard for cloud native application management, providing a powerful and highly automated platform for deploying, managing, and scaling applications in cloud native environments. As an open source monitoring framework, Prometheus provides out-of-the-box monitoring capabilities for the Kubernetes container orchestration platform. For multiple Kubernetes cluster environments across data centers, Prometheus provides a federated cluster architecture implementation. However, for actual production environments, federated clusters do not implement high availability of monitoring systems that allow single Prometheus failures, such as difficulty in detecting Prometheus failures. Summary of the invention

[0003] In order to solve the above technical problems, the present invention is proposed. The embodiments of the present invention provide a method and device for implementing Promtheus high availability based on a cloud native environment, which can solve the problem that Promtheus faults are difficult to find.

[0004] According to one aspect of the present invention, a method for implementing Promtheus high availability based on a cloud native environment is provided, including: dividing multiple Kubenetes clusters based on a Prometheus federated cluster solution, and selecting one cluster as the main cluster, and the other clusters as ordinary clusters; in each ordinary cluster, using a single pod to deploy a sub-Prometheus; wherein the sub-Prometheus does not mount storage, and the sub-Prometheus is used to enable self-monitoring of the corresponding ordinary cluster and collect monitoring data of the corresponding ordinary cluster; deploying a central Prometheus in the main cluster; wherein the central Prometheus mounts storage, and the central Prometheus is used to pull monitoring data of the sub-Prometheus in the ordinary cluster; configuring a global alarm notification in the main cluster, and when the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm condition, the main cluster generates an alarm notification.

[0005] In one embodiment, in each common cluster, a single pod is used to deploy a sub-Prometheus, and the method further includes: each common cluster uses Deployment to manage a single Pod to deploy a sub-Prometheus, and configures the monitoring configuration of the corresponding common cluster; the startup mode of the sub-Prometheus is modified from directly starting the default service to first deleting the wal directory and then starting the service; the default configuration of the sub-Prometheus is modified, the time series database storage retention time is configured to a preset minimum value, and the time series database storage retention size is configured to a preset minimum value.

[0006] In one embodiment, the storage deployment center Prometheus is stored in the main cluster, including: using Statefulset in the main cluster to configure two Pods to deploy two center Prometheus; mounting storage on both center Prometheus; configuring federation for the two center Prometheus to pull monitoring data of the sub-Prometheus in the common cluster.

[0007] In one embodiment, the method for implementing Promtheus high availability based on a cloud native environment also includes: adding a sync container to the Pods of the two central Prometheus to record in real time the latest time of the monitoring data stored by itself; configuring the sync container to read the address information of the other central Prometheus; when any one of the two central Prometheus fails, reading the monitoring data of the other central Prometheus to restore its own monitoring data.

[0008] In one embodiment, the method for implementing Promtheus high availability in a cloud-native environment also includes: when the Pod of the central Prometheus is restarted, the sync container compares whether its own write time is consistent with the current time of the pod; if the own write time is inconsistent with the current time of the pod, it is considered that the corresponding central Prometheus has failed.

[0009] In one embodiment, when any one of the two central Prometheus fails, the monitoring data of the other central Prometheus is read to restore its own monitoring data, including: when the write time of any one of the central Prometheus is inconsistent with the current time of the pod, the central Prometheus with inconsistent time is determined to be the faulty central Prometheus; the central Prometheus whose write time is consistent with the current time of the pod is determined to be the normal central Prometheus; the faulty central Prometheus reads the missing monitoring data of the normal central Prometheus from the said own write time to the current time of the pod; the missing monitoring data is written into the faulty central Prometheus to restore the monitoring data of the faulty central Prometheus.

[0010] In one embodiment, a method for implementing Promtheus high availability based on a cloud native environment also includes: configuring a sub-Prometheus in the main cluster; wherein the sub-Prometheus in the main cluster is used to monitor the central Prometheus; the sub-Prometheus in the main cluster configures the fault alarm rules of the central Prometheus; wherein, when the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm condition, the main cluster generates an alarm notification, including: when the monitoring data of the central Prometheus meets the first alarm condition, the sub-Prometheus in the main cluster generates an alarm and sends a global alarm notification.

[0011] In one embodiment, a method for implementing Promtheus high availability based on a cloud native environment also includes: the sub-Prometheus in the main cluster is associated with the metrics interface of the central Prometheus; the sub-Prometheus in the main cluster pulls the default indicators provided by the metrics interface; wherein the default indicators are used to troubleshoot central Prometheus service problems and as a sign of whether the central Prometheus is available.

[0012] In one embodiment, a global alarm notification is configured in the main cluster. When the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm condition, the main cluster generates an alarm notification, which also includes: configuring an alarm rule for the sub-Prometheus fault alarm in the central Prometheus of the main cluster; when the monitoring data of any sub-Prometheus pulled by the central Prometheus meets the second alarm condition, the central Prometheus generates an alarm and sends a global alarm notification.

[0013] According to another aspect of the present invention, a device for implementing Promtheus high availability based on a cloud native environment is provided, including: a partitioning module, used to partition multiple Kubenetes clusters based on the Prometheus federated cluster solution, and select one cluster as the main cluster, and the other clusters are ordinary clusters; a first deployment module, used to deploy a sub-Prometheus in each ordinary cluster using a single pod; wherein the sub-Prometheus does not mount storage, and the sub-Prometheus is used to enable self-monitoring of the corresponding ordinary cluster and collect monitoring data of the corresponding ordinary cluster; a second deployment module, used to deploy a central Prometheus in the main cluster; wherein the central Prometheus mounts storage, and the central Prometheus is used to pull monitoring data of the sub-Prometheus in the ordinary cluster; an alarm module, used to configure a global alarm notification in the main cluster, and when the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm condition, the main cluster generates an alarm notification.

[0014] The present invention provides a method and device for realizing Promtheus high availability based on a cloud native environment, and a main cluster performs a unified alarm. In this way, no matter a sub-Prometheus failure of a cluster or a central Prometheus failure of the main cluster occurs, the operation and maintenance personnel can receive an alarm notification and thus promptly learn that there is a Prometheus failure in the environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 It is a flowchart of a method for implementing Promtheus high availability based on a cloud native environment provided by an exemplary embodiment of the present invention.

[0017] Figure 2 It is a flowchart of a method for implementing Promtheus high availability based on a cloud native environment provided by another exemplary embodiment of the present invention.

[0018] Figure 3 It is a structural diagram of a Promtheus high-availability device based on a cloud native environment provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0019] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.

[0020] Kubernetes is an open source system for automatically deploying, scaling, and managing containerized applications. The core functions of Kubernetes are to automatically deploy applications to multiple nodes in a cluster, automatically scale up or scale down the number of application instances based on demand, manage clusters of multiple containerized applications, and ensure that they run efficiently on different nodes. Prometheus is an open source service monitoring system and time series database that provides a common data model and a fast data collection, storage, and query interface. Its core component, the Prometheus server, regularly pulls monitoring indicator data from statically configured monitoring targets or targets automatically configured based on service discovery, and persists it to the time series database TSDB. Moreover, Prometheus has become the core of the monitoring system in the Kubernetes ecosystem. It can effectively monitor the Kubernetes cluster itself and the various applications running on it. Prometheus supports the Kubernetes-based service discovery mechanism and can dynamically discover and monitor resources in various dimensions such as Pod, Service, and Ingress in the Kubernetes cluster. In Kubernetes, Prometheus collects container indicator data through kubelet (cAdvisor), such as CPU usage, memory usage, and network message sending / receiving / discarding rates. At the same time, it can also use node_exporter to collect host machine metrics data, such as average load, CPU, memory, disk, network and other information. Prometheus supports the definition of alarm rules. When the monitoring data meets the alarm conditions, an alarm will be generated and sent to Alertmanager for aggregation and distribution. This helps to promptly discover and handle potential problems in the Kubernetes cluster. Alertmanager is a component of Prometheus that is responsible for processing alarms sent by clients, grouping, suppressing, silencing and routing.

[0021] For multiple Kubernetes cluster environments across data centers, Prometheus provides a federated cluster architecture implementation. Each Kubernetes cluster deploys a child Prometheus to collect fine-grained data (instance level) of the cluster. Multiple Kubernetes clusters select a master cluster to deploy one or more central Prometheus, which is responsible for collecting and aggregating data from each child Promtheus (task level) and storing the aggregated data. The advantage of the Prometheus federated cluster is that it combines multiple independent Prometheus instances into a unified monitoring system, thereby achieving comprehensive monitoring across data centers.

[0022] The present invention proposes a method for realizing high availability of Promtheus based on cloud native environment, which makes many optimizations on the basis of native federated cluster, including adding central Prometheus monitoring and alarm to the sub-Prometheus of the main cluster, optimizing the startup logic of the sub-Prometheus to shorten the restart time to solve the problem of data missing and avoid the problem of startup failure caused by data damage, and adding a mechanism for synchronizing the data of the failure time period with each other's data after restarting the two central Prometheus. In this way, whether it is a sub-Prometheus failure of a cluster or a central Prometheus failure of the main cluster, the operation and maintenance personnel can receive alarm notifications to timely learn that there is a Prometheus failure in the environment, and the sub-Prometheus failure can realize automatic repair of the sub-Prometheus without losing monitoring data with the help of the Kubernetes environment pod drift mechanism, and the central Prometheus failure can realize that the central Prometheus does not lose monitoring data after repair with the help of manual repair by the operation and maintenance personnel and the self-synchronization mechanism. Therefore, whether the sub-Prometheus or the central Prometheus fails, the upper-level users who actually use the Prometheus monitoring system will not be aware of it, ensuring the continuity and reliability of the Prometheus service, the reliability and recoverability of the Prometheus data, and thus achieving high availability of the cloud-native environment monitoring system.

[0023] Figure 1 This is a flow chart of a method for implementing Promtheus high availability based on a cloud native environment provided by an exemplary embodiment of the present invention. Figure 1 The method of implementing Promtheus high availability based on cloud native environment is explained. The method of implementing Promtheus high availability in cloud native environment includes: dividing multiple Kubenetes clusters based on Prometheus federated cluster solution, and selecting one cluster as the main cluster, and the other clusters are ordinary clusters (see Figure 1In each normal cluster, a single pod is used to deploy the child Prometheus (see Figure 1 The child Prometheus does not mount storage. The child Prometheus is used to enable self-monitoring of the corresponding common cluster and collect monitoring data of the corresponding common cluster. Deploy the central Prometheus in the main cluster (see Figure 1 S300). The central Prometheus mounts storage and is used to pull monitoring data from the sub-Prometheus in the common cluster. Configure global alarm notifications in the main cluster. When the monitoring data of the central Prometheus or the sub-Prometheus meets the alarm conditions, the main cluster generates an alarm notification (see Figure 1 S400).

[0024] Pod is a core concept in Kubernetes, which represents a collection of closely related containers. These containers share the same network namespace, storage volume and other resources, and are usually scheduled to run on the same node (computer). Pod is the smallest deployable computing unit in Kubernetes and the basic unit for running applications in a cluster.

[0025] In one embodiment, S200 may include: each common cluster uses Deployment to manage a single Pod to deploy a sub-Prometheus, and configures the monitoring configuration of the corresponding common cluster; modifies the startup method of the sub-Prometheus, from directly starting the default service to deleting the wal directory before starting the service; modifies the default configuration of the sub-Prometheus, configures the time series database storage retention time to a preset minimum value, and configures the time series database storage retention size to a preset minimum value.

[0026] Based on the Prometheus federated cluster solution, Kubernetes cluster roles are divided. One Kubernetes cluster is selected as the main cluster, and the other clusters are common clusters. Each common cluster uses Deployment to manage a single Pod. The sub-Prometheus is deployed without mounting storage, and the monitoring configuration required by this cluster is configured, specifically the monitoring objects associated with the monitoring content, such as monitoring Kubernetes nodes, Pods, etc. Not mounting memory means that the Pod is not connected to any external storage space (storage volume). A Pod is created to run Prometheus, but no storage space is configured for it to save data. This can avoid the problem that Prometheus's own data is damaged due to unknown failures and cause its own service to be unavailable. Therefore, it is ensured that the sub-Prometheus can be quickly restored when a failure occurs, providing the availability of the Prometheus service itself. Modify the startup mode to not load wal data to achieve fast startup, and modify the startup configuration to basically not store data to avoid data damage and cause the problem of self-service unavailability.

[0027] One possible way to modify the startup mode is to modify the startup mode of the child Prometheus from the default service direct startup to deleting the wal directory before starting the service. The default startup mode requires loading the cached data in the wal directory, so the startup is slow. Deleting the wal directory and then starting the service greatly shortens the startup time of Prometheus. When any failure occurs in the subsequent Kubernetes cluster, the Prometheus service can be quickly started after the Prometheus Pod drifts, ensuring that the central Prometheus will not fail to pull the child Prometheus data, resulting in interruption of monitoring data.

[0028] One possible way to modify the startup configuration is to modify the default configuration of the child Prometheus, set storage.tsdb.retention.time to the minimum allowed value of 1ms, and set storage.tsdb.retention.size to the minimum allowed value of 1KB. The purpose is to ensure that the child Prometheus generates almost no data storage. On the one hand, it reduces the size of cached data in the wal directory and speeds up the execution speed of deleting the wal directory in the startup method first. On the other hand, it avoids the problem that Prometheus's own data is damaged due to unknown failures, resulting in the unavailability of its own service. Therefore, it ensures that the child Prometheus can recover quickly when a failure occurs, and provides the availability of the Prometheus service itself.

[0029] In one embodiment, S300 may include: using Statefulset to configure two Pods in the main cluster to deploy two central Prometheus; mounting storage on both central Prometheus; configuring federation for the two central Prometheus to pull monitoring data from the sub-Prometheus in the common cluster.

[0030] The main cluster uses Statefulset to configure two Pods to mount storage and deploy the central Prometheus. One Pod deploys one central Prometheus. The storage supports local storage or external storage. Federation is configured to pull the sub-Prometheus data of each cluster to provide global monitoring. The alarm rules and alarm push alertmanager required for this environment are configured to implement global alarms and alarm notifications. The sub-Prometheus does not mount storage, so the monitoring data of the sub-Prometheus is pulled by the central Prometheus. The main Prometheus pulls the monitoring data of the sub-Prometheus and stores the monitoring data of the sub-Prometheus, providing the sub-Prometheus with a basis for optimizing the startup logic and shortening the restart time, solving the problem of data missing and avoiding the problem of startup failure caused by data damage. The sub-Prometheus failure can achieve automatic repair of the sub-Prometheus without losing monitoring data with the help of the pod drift mechanism of the Kubernetes environment. The two central Prometheus add a mechanism to synchronize their own data during the failure period with the help of each other's data after restart. In this way, no matter whether a sub-Prometheus failure of a cluster or a central Prometheus failure of the main cluster occurs, the operation and maintenance personnel can receive alarm notifications to promptly learn that there is a Prometheus failure in the environment. The central Prometheus failure can be manually repaired by the operation and maintenance personnel and its own synchronization mechanism to ensure that the monitoring data is not lost after the central Prometheus is repaired.

[0031] Figure 2 is a flow chart of a method for implementing Promtheus high availability based on a cloud native environment provided by another exemplary embodiment of the present invention, such as Figure 2 As shown, the method for implementing Promtheus high availability based on a cloud native environment can also include: adding a sync container to the Pods of the two central Prometheus to record the latest time of the stored monitoring data in real time (see Figure 2 S500). Configure the sync container to read the address information of the peer center Prometheus (see Figure 2When any of the two central Prometheus fails, it reads the monitoring data of the other central Prometheus to restore its own monitoring data (see Figure 2 S700).

[0032] The two Pods of the central Prometheus add sync containers and configure them to read each other's Pod address information. When the central Prometheus is running, the sync container records the latest time of its own storage monitoring data in real time. When the Pod of the central Prometheus is restarted, the sync container compares its own write time with the current time of the pod. If the write time is inconsistent with the current time of the pod, it is considered that the corresponding central Prometheus has failed.

[0033] In one embodiment, S700 may include: when the write time of any central Prometheus is inconsistent with the current time of the pod, determining the central Prometheus with inconsistent time as the faulty central Prometheus; determining the central Prometheus whose write time is consistent with the current time of the pod as the normal central Prometheus; the faulty central Prometheus reads the missing monitoring data of the normal central Prometheus from its own write time to the current time of the pod; and writing the missing monitoring data into the faulty central Prometheus to restore the monitoring data of the faulty central Prometheus.

[0034] Read the data of the Prometheus in another center during this time range and write it to the local Prometheus, so that when any single point failure occurs in the two center Prometheus, the node data can be restored with the help of other nodes, and finally the data consistency of the two center Prometheus is guaranteed. When the upper-level personnel query the monitoring data, they will not be aware of the failure of the center Prometheus; if it is consistent, it is considered that no failure has occurred and no processing is required. Therefore, after the failure of any center Prometheus, the Prometheus Pod data of another center can be read to restore the monitoring data of its own failure time period. Even if a single point failure occurs, the data during the failure period can be automatically restored to ensure the continuity of the monitoring data, maintain the stability of the main cluster, and improve the availability and stability of the Prometheus monitoring system in complex scenarios with multiple Kubenetes clusters across data centers, thereby meeting the high availability requirements of Prometheus in the production environment.

[0035] In one embodiment, the method for implementing Promtheus high availability based on a cloud native environment may also include: configuring a sub-Prometheus in the main cluster; wherein the sub-Prometheus in the main cluster is used to monitor the central Prometheus; the sub-Prometheus in the main cluster configures the fault alarm rules of the central Prometheus; wherein S400 may include: when the monitoring data of the central Prometheus meets the first alarm condition, the sub-Prometheus in the main cluster generates an alarm and sends a global alarm notification.

[0036] The sub-Prometheus of the main cluster is configured with job="prometheus-center". The sub-Prometheus in the main cluster is associated with the metrics interface of the central Prometheus. The sub-Prometheus in the main cluster pulls the default indicators provided by the metrics interface. The default indicators are used to troubleshoot the central Prometheus service problems and as a sign of whether the central Prometheus is available, and are provided for alarm use. For example, if an alarm indicator appears in the indicator, an alarm is generated based on the alarm indicator and a global alarm notification is sent. The sub-Prometheus of the main cluster enables the alarm function, configures the alarm rules for any failure of the central Prometheus, and uses up{job="prometheus-center"}==0 as the alarm expression. At this time, if any of the two central Prometheus fails, the sub-Prometheus can generate an alarm and send an Alertmanger. The operation and maintenance personnel receive the alarm notification sent by the Alertmanger and log in to the main cluster to troubleshoot the central Prometheus problem. After manual repair, the Pod restart triggers data synchronization. The sub-Prometheus of the main cluster implements the monitoring and alarm of the central Prometheus. Therefore, the dual Pods of the central Prometheus of the main cluster avoid single point failures. The sub-Prometheus monitoring can detect its own failures through alarms. Even if a single point failure occurs, it can automatically restore the data during the failure to ensure the continuity of monitoring data and improve the overall cluster stability.

[0037] In one embodiment, S400 may further include: configuring an alarm rule for a sub-Prometheus fault alarm in the central Prometheus of the main cluster; when the monitoring data of any sub-Prometheus pulled by the central Prometheus meets the second alarm condition, the central Prometheus generates an alarm and sends a global alarm notification.

[0038] The central Prometheus configures alarm rules to add sub-Prometheus failure alarms. The alarm expression uses up{job="prometheus"}==0. At this time, if a sub-Prometheus of any cluster fails, the central Prometheus can generate an alarm and send it to Alertmanager. After receiving the alarm notification sent by Alertmanger, the operation and maintenance personnel log in to the Kubernetes cluster where the faulty Prometheus is located to troubleshoot the problem.

[0039] That is to say, under normal circumstances, the central Prometheus checks and warns the monitoring data uploaded by all sub-Prometheus. When the monitoring data of the sub-Prometheus meets the alarm conditions, the central Prometheus generates an alarm and sends a global alarm notification. When the central Prometheus fails, the sub-Prometheus configured in the main cluster generates an alarm. When the monitoring data of the central Prometheus meets the alarm conditions, the sub-Prometheus in the main cluster generates an alarm and sends a global alarm notification. This linkage method not only improves the stability and controllability of the sub-Prometheus, but also ensures the stability and recoverability of the central Prometheus. When a failure occurs, it will not affect the upper-level users of the Prometheus monitoring system.

[0040] A possible implementation process for implementing Promtheus high availability based on cloud native environment: First, divide the Kubenetes cluster roles based on the Prometheus federated cluster solution, select one cluster as the main cluster, and the main cluster provides central Prometheus storage and Alertmanager notification to achieve global monitoring and alarm capabilities for cloud native environment. Secondly, in each Kubernetes cluster, use a single Pod without mounting storage to deploy sub-Prometheus, configure the monitoring configuration required by this cluster, and increase the monitoring of its own availability; modify the startup mode to not load wal data to achieve fast startup, and modify the startup configuration to basically not store data to avoid data damage and cause self-service unavailability. Then use two Pods to mount storage and deploy central Prometheus in the main cluster, configure federation to pull sub-Prometheus data of each cluster to provide global monitoring, configure the alarm rules and alarm push alertmanager required for this environment to achieve global alarms and alarm notifications, and at the same time, add sub-Prometheus failure alarms to the alarm rules to detect sub-Prometheus failures. Then, the Pod of the central Prometheus adds a sync container to record the latest time of its own stored monitoring data in real time, and reads the data of another central Prometheus Pod when its own Pod restarts to restore its own monitoring data during the failure period. Finally, modify the configuration of the sub-Prometheus of the main cluster to increase the monitoring of whether the central Prometheus is available; enable the alarm function and push alertamagner, and configure the alarm rule to alarm the central Prometheus failure to detect the central Prometheus failure.

[0041] Prometheus high availability is achieved in the cloud native environment. For complex scenarios with multiple kubernetes environments, based on the division of roles in the main cluster and the optimization of the Prometheus federated cluster solution, the sub-Prometheus does not store data and starts quickly when restarting. The sub-Prometheus in the main cluster implements the monitoring and alarm of the central Prometheus, and the central Prometheus implements automatic recovery of data in the failure time period after a single point failure. This method is both an extension of the original solution and makes up for the shortcomings of the original solution. It truly realizes the high availability of the Prometheus federated cluster method in complex scenarios with multiple Kubernetes clusters in the cloud native environment. It is a very important requirement for the application of the Prometheus monitoring system in the production environment. This method can improve the availability and stability of the Prometheus monitoring system in complex scenarios with multiple Kubenetes clusters across data centers, thereby meeting the requirements for high availability of Prometheus in the production environment.

[0042] Based on the above modifications, operation and maintenance personnel can detect sub-Prometheus or central Prometheus failures in real time, automatically realize the rapid recovery of uninterrupted monitoring capabilities after sub-Prometheus failures, and automatically synchronize monitoring data during the failure period after the central Prometheus failure is repaired. Therefore, the Prometheus monitoring system can ensure high availability when either sub-Prometheus or central Prometheus fails.

[0043] Figure 3 is a schematic diagram of a structure of a Promtheus high-availability device based on a cloud native environment provided by an exemplary embodiment of the present invention. Figure 3 As shown, a Promtheus high-availability device 3 for implementing a cloud-native environment includes: a division module 31, which is used to divide multiple Kubenetes clusters based on the Prometheus federated cluster solution, and select one cluster as the main cluster, and the other clusters are ordinary clusters; a first deployment module 32, which is used to deploy a sub-Prometheus in each ordinary cluster using a single pod; wherein the sub-Prometheus does not mount storage, and the sub-Prometheus is used to enable self-monitoring of the corresponding ordinary cluster and collect monitoring data of the corresponding ordinary cluster; a second deployment module 33, which is used to deploy a central Prometheus in the main cluster; wherein the central Prometheus mounts storage, and the central Prometheus is used to pull the monitoring data of the sub-Prometheus in the ordinary cluster; an alarm module 34, which is used to configure a global alarm notification in the main cluster, and when the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm condition, the main cluster generates an alarm notification.

[0044] The present invention provides a Promtheus high-availability device based on a cloud-native environment, and a main cluster performs a unified alarm. In this way, no matter a sub-Prometheus failure of a cluster or a central Prometheus failure of the main cluster occurs, the operation and maintenance personnel can receive an alarm notification to promptly learn that there is a Prometheus failure in the environment.

[0045] In one embodiment, the first deployment module 32 can be configured as follows: each ordinary cluster uses Deployment to manage a single Pod to deploy a sub-Prometheus, and configures the monitoring configuration of the corresponding ordinary cluster; modifies the startup method of the sub-Prometheus, from directly starting the default service to deleting the wal directory before starting the service; modifies the default configuration of the sub-Prometheus, configures the time series database storage retention time to the preset minimum value, and configures the time series database storage retention size to the preset minimum value.

[0046] In one embodiment, the second deployment module 33 can be configured as follows: use Statefulset to configure two Pods in the main cluster to deploy two central Prometheus; mount storage on both central Prometheus; configure federation for the two central Prometheus to pull monitoring data from the sub-Prometheus in the common cluster.

[0047] In one embodiment, the Promtheus high-availability device 3 implemented based on the cloud native environment can be configured as follows: the Pods of the two central Prometheus both add a sync container to record the latest time of their own stored monitoring data in real time; the sync container is configured to read the address information of the other central Prometheus; when any of the two central Prometheus fails, the monitoring data of the other central Prometheus is read to restore its own monitoring data.

[0048] In one embodiment, the Promtheus high availability device 3 based on the cloud native environment can also be configured as follows: when the Pod of the central Prometheus is restarted, the sync container compares its own write time with the current time of the pod to see if it is consistent; if the own write time is inconsistent with the current time of the pod, it is considered that the corresponding central Prometheus has failed.

[0049] In one embodiment, the Promtheus high-availability device 3 implemented based on the cloud-native environment can also be configured as follows: when the write time of any central Prometheus is inconsistent with the current time of the pod, the central Prometheus with inconsistent time is determined to be the faulty central Prometheus; the central Prometheus whose write time is consistent with the current time of the pod is determined to be the normal central Prometheus; the faulty central Prometheus reads the missing monitoring data of the normal central Prometheus from its own write time to the current time of the pod; and the missing monitoring data is written into the faulty central Prometheus to restore the monitoring data of the faulty central Prometheus.

[0050] In one embodiment, the Promtheus high-availability device 3 based on the cloud-native environment can also be configured as: configuring a sub-Prometheus in the main cluster; wherein the sub-Prometheus in the main cluster is used to monitor the central Prometheus; the sub-Prometheus in the main cluster configures the fault alarm rules of the central Prometheus; wherein the alarm module 34 can be correspondingly configured as: when the monitoring data of the central Prometheus meets the first alarm condition, the sub-Prometheus in the main cluster generates an alarm and sends a global alarm notification.

[0051] In one embodiment, the Promtheus high-availability device 3 based on the cloud-native environment can also be configured as: the sub-Prometheus in the main cluster is associated with the metrics interface of the central Prometheus; the sub-Prometheus in the main cluster pulls the default indicators provided by the metrics interface; wherein, the default indicators are used to troubleshoot the central Prometheus service problems and as a sign of whether the central Prometheus is available.

[0052] In one embodiment, the alarm module 34 can be configured as follows: configuring the alarm rules for sub-Prometheus fault alarms in the central Prometheus of the main cluster; when the monitoring data of any sub-Prometheus pulled by the central Prometheus meets the second alarm condition, the central Prometheus generates an alarm and sends a global alarm notification.

[0053] The embodiment provides a Promtheus high-availability device based on a cloud-native environment. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From a hardware perspective, in addition to the CPU, memory, network interface, and non-volatile memory, the device in the embodiment can generally include other hardware, such as a forwarding chip responsible for processing messages, etc. Taking software implementation as an example, as a device in a logical sense, it is formed by the CPU of the device in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them.

[0054] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute any of the above embodiments of the method for implementing Promtheus high availability based on a cloud native environment.

[0055] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method for implementing Promtheus high availability based on a cloud native environment according to various embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0056] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; and a processor for executing any of the above-mentioned embodiments of the method for implementing Promtheus high availability based on a cloud native environment.

[0057] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of a method for implementing Promtheus high availability based on a cloud native environment according to various embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0058] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for implementing Promtheus high availability based on a cloud native environment, characterized in that: include: Divide multiple Kubenetes clusters based on the Prometheus federated cluster solution, select one cluster as the main cluster, and the other clusters as common clusters; In each common cluster, a sub-Prometheus is deployed using a single pod; wherein the sub-Prometheus does not mount storage, and the sub-Prometheus is used to enable self-monitoring of the corresponding common cluster and collect monitoring data of the corresponding common cluster; Deploy a central Prometheus in the main cluster; wherein the central Prometheus mounts storage, and the central Prometheus is used to pull monitoring data of the sub-Prometheus in the common cluster; Configure global alarm notifications in the main cluster. When the monitoring data of the central Prometheus or the monitoring data of the child Prometheus meets the alarm conditions, the main cluster generates an alarm notification.

2. According to the method for realizing Promtheus high availability based on cloud native environment according to claim 1, it is characterized in that: In each normal cluster, a single pod is used to deploy a child Prometheus, which also includes: Each common cluster uses Deployment to manage a single Pod deployment sub-Prometheus and configures the monitoring configuration of the corresponding common cluster; Modify the startup method of child Prometheus, from directly starting the default service to deleting the wal directory before starting the service; Modify the default configuration of the child Prometheus, configure the time series database storage retention time to the preset minimum value, and configure the time series database storage retention size to the preset minimum value.

3. According to the method for realizing Promtheus high availability based on cloud native environment according to claim 1, it is characterized in that: Deploy the central Prometheus in the main cluster, including: Use Statefulset to configure two Pods in the main cluster to deploy two central Prometheus; both central Prometheus are mounted with storage; Configure federation for the two central Prometheus to pull monitoring data from the sub-Prometheus in the common cluster.

4. According to claim 3, the method for realizing Promtheus high availability based on a cloud native environment is characterized in that: Methods for implementing Promtheus high availability based on cloud native environments also include: The Pods of Prometheus in both centers have added sync containers to record the latest time of their own stored monitoring data in real time; Configure the sync container to read the address information of the peer center Prometheus; When any of the two central Prometheus fails, the monitoring data of the other central Prometheus is read to restore its own monitoring data.

5. According to claim 4, the method for implementing Promtheus high availability based on a cloud native environment is characterized in that: The cloud-native environment also implements the following methods to achieve high availability of Promtheus: When the central Prometheus Pod is restarted, the sync container compares its own write time with the pod's current time to see if they are consistent; If the write time itself is inconsistent with the current time of the pod, it is considered that the corresponding central Prometheus has failed.

6. According to the method for realizing high availability of Promtheus based on cloud native environment according to claim 5, it is characterized in that: When any of the two central Prometheus fails, the monitoring data of the other central Prometheus is read to restore its own monitoring data, including: When the write time of any center Prometheus is inconsistent with the current time of the pod, the center Prometheus with inconsistent time is determined as the fault center Prometheus; Determine that the central Prometheus whose write time is consistent with the current time of the pod is the normal central Prometheus; The fault center Prometheus reads the missing monitoring data from the normal center Prometheus between the time it writes itself and the current time of the pod; Write the missing monitoring data to the fault center Prometheus to restore the monitoring data of the fault center Prometheus.

7. According to claim 1, the method for implementing Promtheus high availability based on a cloud native environment is characterized in that: The method of implementing Promtheus high availability based on cloud native environment also includes: Configure a sub-Prometheus in the main cluster. The sub-Prometheus of the main cluster is used to monitor the central Prometheus. The sub-Prometheus in the main cluster configures the fault alarm rules of the central Prometheus; When the monitoring data of the central Prometheus or the monitoring data of the child Prometheus meets the alarm conditions, the main cluster generates an alarm notification, including: When the monitoring data of the central Prometheus meets the first alarm condition, the sub-Prometheus in the main cluster generates an alarm and sends a global alarm notification.

8. According to claim 7, the method for implementing Promtheus high availability based on a cloud native environment is characterized in that: The method of implementing Promtheus high availability based on cloud native environment also includes: The sub-Prometheus in the main cluster is associated with the metrics interface of the central Prometheus; The sub-Prometheus in the main cluster pulls the default indicators provided by the metrics interface; wherein the default indicators are used to troubleshoot the central Prometheus service problems and serve as a sign of whether the central Prometheus is available.

9. According to claim 1, the method for implementing Promtheus high availability based on a cloud native environment is characterized in that: Configure global alarm notifications in the main cluster. When the monitoring data of the central Prometheus or the monitoring data of the child Prometheus meets the alarm conditions, the main cluster generates an alarm notification, which also includes: Configure the alarm rules for sub-Prometheus failure alarms in the central Prometheus of the main cluster; When the monitoring data of any child Prometheus pulled by the central Prometheus meets the second alarm condition, the central Prometheus generates an alarm and sends a global alarm notification.

10. A high-availability device for Promtheus based on a cloud-native environment, characterized in that: include: The partitioning module is used to partition multiple Kubenetes clusters based on the Prometheus federated cluster solution, and select one cluster as the main cluster, and the other clusters are common clusters; The first deployment module is used to deploy a sub-Prometheus using a single pod in each common cluster; wherein the sub-Prometheus does not mount storage, and the sub-Prometheus is used to enable self-monitoring of the corresponding common cluster and collect monitoring data of the corresponding common cluster; The second deployment module is used to deploy a central Prometheus in the main cluster; wherein the central Prometheus mounts storage, and the central Prometheus is used to pull monitoring data of the sub-Prometheus in the common cluster; The alarm module is used to configure global alarm notifications in the main cluster. When the monitoring data of the central Prometheus or the monitoring data of the sub-Prometheus meets the alarm conditions, the main cluster generates an alarm notification.