High-availability monitoring method based on container cloud environment
By deploying Remote Storge data packets and application configurations in a container cloud environment, combined with components such as Prometheus-Operator and Grafana, efficient monitoring and alarming of Kubernetes cluster containers is achieved, and the problem of inefficient management in the existing technology is solved.
Patent Information
- Application Number
- CN202311819184.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
In a container cloud environment, it is difficult for existing monitoring systems to achieve refined management of Kubernetes cluster containers, resulting in inefficient problem investigation and processing.
A highly available monitoring method based on the container cloud environment is adopted, and real-time monitoring and alarming of the container's operating status is achieved by deploying cloud component Remote Storge data packets on the container cloud and deploying application configurations on the user side. The method includes components such as Prometheus-Operator, Grafana, kube-state-metrics, and other components, which are used to collect, store and display monitoring data, and implement alarm and log monitoring through Alertmanager and Loki components.
It realizes refined management of Kubernetes cluster containers, improves the efficiency of problem investigation and processing, and reduces the difficulty and cost of secondary development of cloud environments.
Smart Images

Figure CN120223559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and more specifically, to a highly available monitoring method based on a container cloud environment. Background Art
[0002] With the rapid development of cloud computing technology, containerization technology has gradually become mainstream. In a container cloud environment, the number of containers is huge and they are widely distributed. Therefore, a highly available monitoring system is needed to monitor the status, resource utilization, network traffic, and log information of containers in real time to ensure the stability and reliability of the container cloud environment. Through the API provided by the container management platform, the status of containers is queried regularly and displayed in the monitoring center. At the same time, by setting thresholds, abnormal states of containers are detected and corresponding measures are taken in a timely manner. To improve availability, a primary and standby monitoring center model can be adopted. When the primary monitoring center has problems, it can be automatically switched to the standby monitoring center. A monitoring agent program is installed on each host to collect resource utilization data of each container. The resource utilization of each container is displayed through a visualization interface, and the threshold alarm function is supported to detect containers with abnormal resource occupancy in a timely manner. At the same time, a cluster resource management tool can be considered to monitor and schedule the resource utilization of the entire cluster. Through the network plugin of the container cloud, network data such as the ingress traffic and egress traffic of containers are collected. These data are transmitted to the monitoring center, and the network traffic situation of containers is presented in the form of charts to detect abnormal traffic in a timely manner. To improve availability, multiple network plugins can be used simultaneously to avoid single point of failure. A log collector is integrated into the container to send the log information of the container to the log center. By analyzing the log data of the container, possible abnormal situations are detected, and query and analysis functions are provided to quickly locate problems. At the same time, to improve availability, a distributed log system can be adopted to avoid single point of failure. In the monitoring solution, strict access control policies are implemented to restrict access to the monitoring system to only authorized personnel. At the same time, fine-grained permission control can be performed according to user roles. The transmission of monitoring data should use encryption protocols such as to ensure the security of data transmission, prevent illegal tampering or theft, and perform log auditing on operations and accesses in the monitoring system to record relevant operation information for tracking and analysis. Summary of the Invention
[0003] In order to solve the above problems, the present invention proposes a highly available monitoring method based on a container cloud environment, which can achieve refined management of Kubernetes cluster containers, facilitate troubleshooting the source of problems and handling problems in a timely manner, improve the generality and efficiency of collecting and monitoring different manufacturers' cloud environments, and reduce the difficulty and cost of secondary development of cloud environments.
[0004] The technical solution adopted by the present invention is: a highly available monitoring method based on a container cloud environment, including:
[0005] Install and deploy the cloud component Remote Storge data packet in the container cloud to obtain and manage various resource objects related to the container environment and their corresponding monitoring data;
[0006] Deploy the application configuration at the user end to establish a connection with the container cloud, monitor the running status of the container in real time and display it to the user, and give an over-limit alarm.
[0007] The cloud component Remote Storge data packet includes the following key components:
[0008] The Prometheus-Operator component is used to regularly, real-timely and accurately manage the monitoring data collected by Prometheus; according to different types of data collection tasks, use the Prometheus component cluster to divide the tasks into each Prometheus sub-service; the Prometheus sub-services are divided into the following types of information collection and monitoring: physical layer information, container information, alarm information, and log information; judge whether the monitoring data reaches the alarm threshold, if so, alarm and notify the user to handle it;
[0009] The Prometheus-Service component is used to obtain, store and query monitoring data, capture and store time series data, capture metrics from the configured target location through the HTTP protocol, statically configure and manage monitoring targets, or dynamically manage monitoring targets in cooperation with the Service Discovery method, and obtain data from these monitoring targets; Prometheus-Service is a time series database that stores the collected monitoring data in the local disk in a time series manner;
[0010] The Grafana component is used to, according to the specified dashboard accessed by the user, send an http request to the corresponding Prometheus-Operator to obtain metric data and feedback and display it at the specified position on the dashboard;
[0011] The kube-state-metrics component is used to listen to the API Server to generate status metrics related to resource objects: Deployment, Node, and Pod status metrics; the metric data is used to characterize the metadata status related to the business; the metadata is the replica status, replicas scheduling quantity, and pod; kube-state-metrics only provides one metrics data and does not store these metric data, and uses Prometheus to capture the relevant data and then store it;
[0012] The Alertmanager component is used to receive client instructions and execute the alarm limits, notification content, and notification forms input by the user;
[0013] The Loki component collects detailed log information during the operation of Pods, which is used to monitor the running status of Pods, predict potential risks based on historical trends, and thus make timely adjustments and optimizations;
[0014] The Cadvisor component is a tool for monitoring container resources, which is used to collect, aggregate, process, and export information about running containers in real time, including CPU usage, memory usage, network throughput, and file system usage.
[0015] The Prometheus-Operator component is used to collect relevant metric data of applications and expose it through the metrics interface; after the ServiceMonitor registers with the Prometheus-Operator, the Prometheus-Operator will regularly collect monitoring data.
[0016] The registration of the ServiceMonitor with the Prometheus-Operator is a passive discovery process. The Prometheus-Operator scans all ServiceMonitors in the cluster. When a new created node is found, the address for obtaining metric data of the application is stored in the Prometheus-Operator, and then the Prometheus-Operator regularly pulls the metric data.
[0017] The Prometheus-Operator moves the collected metric data to the remote storage Remote-storage and then displays the data through Grafana.
[0018] The notification forms are email and short message forms.
[0019] Deploying the application configuration on the client means configuring the following settings in the Grafana configuration component:
[0020] a. Set the alarm channel and input the alarm limit, notification content, and notification form;
[0021] b. Configure the data visualization module dashboard, including configuring charts and dashboards, which are used to intuitively display monitoring data to users, including the running status and performance of the container environment;
[0022] c. Set the alarm threshold, which is used to automatically trigger the alarm mechanism when the monitoring data exceeds the preset threshold and send alarm information to relevant personnel in a timely manner.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention provides a highly available monitoring method based on a container cloud environment, which monitors the aggregated metric data of the same service distributed on different machine nodes, then sends the monitored aggregated monitoring data to users in real time in the form of alarms, and displays these aggregated monitoring data in different ways, so as to realize the refined management of Kubernetes cluster containers, facilitate troubleshooting the problem source and handling problems in a timely manner. Brief Description of the Drawings
[0025] Figure 1 is a flowchart of the highly available monitoring method based on a container cloud environment according to an embodiment of the present invention;
[0026] Figure 2 is a flowchart of the container cloud environment architecture according to an embodiment of the present invention; Detailed Embodiments
[0027] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation methods of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0029] As Figure 1 shown, an embodiment of the present invention provides a highly available monitoring method based on a container cloud environment, including:
[0030] Step 1: Deploy components on the container platform, including Prometheus-Operator, Grafana, and kube-state-metrics, where Prometheus-Operator is used to collect monitoring data, Grafana is used to display monitoring data, and kube-state-metrics is used to obtain resource objects and corresponding monitoring data of the K8S platform;
[0031] Step 2: Configure an alarm channel for Grafana and set a Prometheus data source; preferably, the alarm channels include WeChat, SMS, and email;
[0032] Step 3: Deploy the application, synchronize the deployment components, and then collect monitoring data regularly through Prometheus-Operator; among them, the deployed application includes a database, middleware, or an application system, and the components include the Exporter component and the ServiceMonitor. The Exporter component is responsible for exposing the corresponding metrics externally, and the ServiceMonitor is responsible for registering with Prometheus-Operator.
[0033] Step 4: Grafana configures the data visualization module dashboard and sets the alarm threshold;
[0034] Step 5: Determine whether the monitoring data reaches the alarm threshold. If so, send an alarm notification to the user for handling.
[0035] Preferably, after the application is deployed, Prometheus-Operator will collect relevant metric data of the application (such as information on CPU, memory, http requests, etc.), and these metric data are exposed externally through the / metrics interface. After the ServiceMonitor registers with Prometheus-Operator, Prometheus-Operator will collect monitoring data regularly.
[0036] Among them, the registration of the ServiceMonitor with Prometheus-Operator is a passive discovery process. Prometheus-Operator will scan all ServiceMonitors in the cluster. After discovering a newly created one, it will store the address for obtaining metric data of the corresponding application in Prometheus-Operator, and then Prometheus-Operator will pull the metric data regularly.
[0037] Preferably, the metric data collected by Prometheus-Operator will be stored at the location specified for saving metrics during the installation of Prometheus-Operator, and then the data will be displayed through Grafana.
[0038] Preferably, when the user accesses the specified dashboard, Grafana will initiate an http request to access Prometheus-Operator to obtain metric data and display it at the specified location on the dashboard. In the specific dashboard, set the alarm threshold. When the monitoring data reaches the alarm threshold, an alarm will be triggered. The user will handle the alarm in a timely manner through the set alarm channel.
[0039] Such as Figure 2As shown, it includes a data source layer, a data collection layer, a data display layer, a data storage layer, and an alarm notification layer, with the functions as follows:
[0040] The data source layer monitors the memory, CPU, network IO, disk IO, etc. of the Cadvisor container resource monitoring container and provides a WEB page to view the real-time running status of the container. It mainly displays the monitoring data at two levels, namely Host and container, as well as the historical change data. The Node Exporter is a metric data collection component responsible for collecting data from the target Jobs and converting the collected data into the time series data format supported by Prometheus. It is mainly used to collect the hardware and system metrics of the UNIX-like kernel. The Kube-state-metrics is used to listen to the API Server to generate status metrics related to resource objects, mainly some metadata related to the business, such as the status of Deployments, Pod replicas, how many replicas are scheduled, the status of how many pods are running / stopped / terminated, and how many times the pods have been restarted, etc. Metrics provides the data foundation for the monitoring of microservices and is a set of standard metric libraries used to provide all-round, multi-dimensional, real-time, and accurate metric services from the operating system, virtual machine, container to the application.
[0041] The data collection layer is mainly responsible for scraping and storing time series data from various data sources. These data sources can be statically configured or dynamically found through service discovery. The data source layer module scrapes metrics from the target location through the HTTP protocol and stores these data in the local disk in the form of time series. Another important function is to scrape and store the monitoring data. The monitoring targets can be managed statically or dynamically in combination with the Service Discovery method, and data can be obtained from these monitoring targets. And it provides the Exporters tool, which can export the metrics of common services into the format that can be scraped by Prometheus.
[0042] The data display layer is mainly used to monitor and record various systems and applications, and provides a powerful visualization interface that can help users analyze and display large amounts of data. It can be easily integrated with various data sources, including Prometheus, InfluxDB, etc. Users can create various dashboards and visualization reports to better understand the status and performance of their systems and applications. Loki presents a large amount of log data in a visual way so that users can better understand, monitor, and debug their systems and applications. By using these components, users can easily create various dashboards and reports to better understand the status and performance of their systems and applications. In addition, these components also provide powerful data integration and query functions, which can help users better manage and analyze their log data.
[0043] The data storage layer stores data in a stable storage medium so that data will not be lost in case of system crashes or failures. Data persistence storage usually uses hardware devices such as disks and tapes as storage media and adopts various data protection technologies such as RAID and redundant backups to ensure the reliability and security of data. RemoteStorage stores data in a remote storage system to achieve centralized management and sharing of data. By using the RemoteStorage data storage layer, the data of multiple applications or systems can be centrally stored on one or more remote storage devices, facilitating unified management and access to data. In addition, the RemoteStorage data storage layer can also provide functions such as data backup, recovery, and fault tolerance to ensure the reliability and availability of data.
[0044] The alert notification layer is responsible for receiving alerts sent by the Prometheus server and performing further processing and notification. Its functions include:
[0045] 1. Manage alerts: It can receive, store, and route alert information from the Prometheus server. It can manage alerts from multiple Prometheus servers and perform operations such as aggregation, suppression, or silencing according to preset rules.
[0046] 2. Send notifications: It has built-in support for multiple notification methods, including email, Slack, Webhook, etc. When specific rule conditions are met, it can send notifications to relevant personnel or systems so that problems can be addressed in a timely manner.
[0047] 3. Deduplication and aggregation: It can eliminate duplicate alert information and group similar alerts. This helps reduce the number of notifications and makes it easier for administrators to identify and solve potential problems.
[0048] 4. Routing Notification: Alerts can be routed to different recipients or groups according to preset rules, enabling administrators to send relevant alerts to specific personnel or teams for handling based on actual situations.
[0049] In summary, the alert notification layer plays a crucial role in the Prometheus monitoring system, ensuring that alerts are sent to the corresponding recipients in a reliable manner, helping to detect problems in a timely manner and take measures for handling.
[0050] It should be noted that for this embodiment, for the sake of simplicity of description, it is expressed as a series of action combinations. However, those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
Claims
1. A highly available monitoring method based on a container cloud environment, characterized in that, Including: Install and deploy the cloud component Remote Storge data packet in the container cloud to obtain and manage various resource objects related to the container environment and their corresponding monitoring data; Deploy the application configuration at the user end to establish a connection with the container cloud, monitor the running status of the container in real time and display it to the user, and give an out-of-limit alarm.
2. The high-availability monitoring method based on a container cloud environment according to claim 1, wherein The cloud component Remote Storge data packet includes the following key components: The Prometheus-Operator component is used to regularly, real-time and accurately manage Prometheus to collect monitoring data; according to different types of data collection tasks, use the Prometheus component cluster to divide the tasks into each Prometheus sub-service; the Prometheus sub-service is divided into the following types of information collection and monitoring: physical layer information, container information, alarm information, log information; judge whether the monitoring data reaches the alarm threshold, if so, alarm and notify the user to handle; The Prometheus-Service component is used to obtain, store and query monitoring data, capture and store time series data, grab metrics from the configured target location through the HTTP protocol, statically configure and manage monitoring targets, or dynamically manage monitoring targets in cooperation with the Service Discovery method, and obtain data from these monitoring targets; Prometheus-Service is a time series database, and stores the collected monitoring data in the local disk in the form of time series; The Grafana component is used to access the Prometheus-Operator to obtain metric data according to the specified dashboard accessed by the user, and feedback and display it at the specified position of the dashboard; The kube-state-metrics component is used to listen to the API Server to generate status metrics related to resource objects: Deployment, Node, Pod status metrics; the metric data is used to represent the metadata status related to the business; The metadata is the replica status, replicas scheduling quantity, pod; kube-state-metrics only provides one metrics data and does not store these metric data, and Prometheus is used to grab the relevant data and then store it; The Alertmanager component is used to receive the instruction from the user end and execute the input alarm limit value, notification content and notification form; The Loki component collects the detailed log information during the operation of the Pod, is used to monitor the running status of the Pod, predict potential dangers according to the historical situation, and make timely adjustments and optimizations; The Cadvisor component is a tool for monitoring container resources, and is used to collect, aggregate, process and export information about the running containers in real time, including CPU usage, memory usage, network throughput and file system usage.
3. The high-availability monitoring method based on the container cloud environment according to claim 2, wherein The Prometheus-Operator component is used to collect relevant metric data of the application and expose it externally through the metrics interface; after the ServiceMonitor registers with the Prometheus-Operator, the Prometheus-Operator will collect monitoring data regularly.
4. The high-availability monitoring method based on a container cloud environment according to claim 4, wherein The registration of the ServiceMonitor with the Prometheus-Operator is a passive discovery process. The Prometheus-Operator scans all ServiceMonitors in the cluster. When a new node is created, it stores the address for the application to obtain metric data in the Prometheus-Operator, and then the Prometheus-Operator pulls the metric data regularly.
5. The highly available monitoring method based on a container cloud environment according to claim 1, characterized in that The Prometheus-Operator moves the collected metric data to the remote storage Remote-storage and then displays the data through Grafana.
6. The high-availability monitoring method based on a container cloud environment according to claim 1, characterized in that The notification forms are email and short message.
7. The highly available monitoring method based on a container cloud environment according to claim 1, wherein Deploying the application configuration at the user end means configuring the following settings in the Grafana configuration component: a. Set the alarm channel, input the alarm limit value, notification content, and notification form; b. Configure the data visualization module dashboard, including configuring charts and dashboards, to visually display the monitoring data to the user, including the running status and performance of the container environment; c. Set the alarm threshold, which is used to automatically trigger the alarm mechanism when the monitoring data exceeds the preset threshold and send alarm information to relevant personnel in a timely manner.