Method and system for improving observability of cloud native platform

Through the integration of Prometheus and Exporter, the construction of eBPF technology and the construction of Grafana monitoring board, the data island, real-time and alarm mechanism problems in the monitoring solution of the cloud-native platform are solved, and the comprehensive monitoring and intelligent alarm of the cloud-native platform are realized, and observability is improved.

CN120295850APending Publication Date: 2025-07-11QIMING INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411608851.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the monitoring solutions of cloud-native platforms have problems such as data islands, insufficient real-time, insufficient monitoring granularity and unintelligent alarm mechanisms, which leads to the inability to achieve comprehensive monitoring and intelligent alarms, affecting the observability of cloud-native platforms.

Method used

The integration of Prometheus and Exporter is adopted, and data collection and real-time acquisition of network information is combined with eBPF technology. A Grafana monitoring board is built, and an intelligent alarm mechanism is configured to achieve comprehensive monitoring of cloud-native platforms.

Benefits of technology

Real-time monitoring and intelligent alarms for cloud-native platforms are realized, providing fine-grained data collection and visual analysis, reducing data collection delays, reducing false alarm rates, and improving operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295850A_ABST
    Figure CN120295850A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for improving observability of a cloud native platform. The method comprises the following steps of S1, integrating Prometheus and Export; s2, acquiring network information in real time through the eBPF; s3, collecting and storing data; s4, a Grafana monitoring billboard is constructed; and S5, configuring an intelligent alarm mechanism to realize real-time monitoring of the cloud native platform. According to the invention, real-time monitoring of key system resources such as a CPU, a memory, a disk, a network and the like, various middleware and a database can be realized, early warning is carried out on clusters and user applications in combination with a Grafana visual board and an intelligent warning mechanism, and data support is provided for troubleshooting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud native platform monitoring, and particularly to a method and system for improving the observability of cloud native platforms. Background Art

[0002] With the rapid development of cloud computing and microservices architecture, the complexity of cloud native platforms has increased significantly. Cloud native applications usually run in a multi-node and multi-container environment, and the monitoring of system resources and cluster resources has become particularly important. In this architecture, ensuring the effective utilization and health status of the cluster and the system is crucial for ensuring the performance and stability of the application. The current traditional monitoring solutions face the following problems: 1. Data silos: Data between different monitoring tools cannot be integrated, resulting in a one-sided monitoring perspective.

[0003] 2. Lack of real-time performance: Delays in data collection and processing prevent operations personnel from discovering problems in a timely manner.

[0004] 3. Insufficient monitoring granularity: Lack of monitoring of fine-grained metrics makes it difficult to deeply analyze the root cause of problems.

[0005] 4. Unintelligent alerting mechanism: Traditional alerting rules often lead to false alarms, increasing the burden on operations and maintenance.

[0006] Therefore, there is an urgent need for an integrated solution that can achieve comprehensive monitoring and intelligent alerting to improve the observability of cloud native platforms. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for improving the observability of cloud native platforms to solve the technical problem of how to achieve comprehensive monitoring of cloud native platforms.

[0008] The present invention is implemented by the following technical solutions: A method for improving the observability of cloud native platforms includes the following steps: S1: Integrate Prometheus with Exporter; S2: Obtain network information in real time through eBPF; S3: Collect and store data; S4: Build a Grafana monitoring dashboard; S5: Configure an intelligent alerting mechanism to achieve real-time monitoring of cloud native platforms.

[0009] Further, step S1 includes the following sub-steps: S11: In the cloud native platform, deploy Prometheus, configure it as a monitoring center, and deploy Kube-state-metrics in the cluster to monitor pod resources and anomalies; S12: Select Exporter to monitor resource usage at the node level and business metrics.

[0010] Further, step S12 is specifically as follows: Monitor resource usage at the node level through Node Exporter, including CPU, memory, and disk; collect specific business metrics through Application Exporter, including kafka-exporter and redis-exporter.

[0011] Further, step S2 is specifically as follows: Obtain real-time network traffic, socket connection information, network latency, and packet loss information through eBPF.

[0012] Further, step S3 is specifically as follows: Prometheus pulls data from various Exporters through the HTTP protocol, forms time-series data, and stores it in the TSDB library to support high-performance real-time query and analysis.

[0013] Further, step S3 also includes: Configure the scraping interval and data storage strategy of Prometheus to optimize performance and resource usage.

[0014] Further, step S4 is specifically as follows: Deploy Grafana and configure Prometheus as the data source, and design a personalized monitoring dashboard to display the changes of key metrics, where the key metrics include CPU usage rate, memory usage, disk IO performance, and network traffic and latency.

[0015] Further, the key metrics also include Pod exception information and middleware itself metrics. Among them, Pod exception information includes abnormal Pod restarts, evictions, and resource shortages; middleware itself metrics include kafka production and consumption situations and redis hit situations.

[0016] Further, step S5 is specifically as follows: Configure alerting rules in Prometheus and configure Alertmanager to define alert receivers, notification methods, and repeated alerts; the alerting rules include CPU usage rate thresholds, memory usage rate thresholds, disk available space thresholds, network latency thresholds, and Pod exception situations.

[0017] A system for enhancing the observability of a cloud-native platform, used to implement the method for enhancing the observability of a cloud-native platform described above, includes an integrated data collection module, a visualization monitoring module, and an intelligent alerting module, where, An integrated data collection module that uses Prometheus as the central monitoring system, combines kube-stat-metrics and various Exporters to collect key metrics of CPU, memory, disk, and network for clusters, nodes, and pods, and uses eBPF technology to obtain real-time network traffic, latency, and packet loss information; A visualization monitoring module that uses Grafana to build a real-time monitoring dashboard, displays the data collected by Prometheus, supports various chart forms, and provides a user interface for easy data analysis and problem troubleshooting by operation and maintenance personnel; An intelligent alerting module that uses Prometheus Alertmanager to implement alerts and supports configuring alert policies to ensure timely information delivery.

[0018] The beneficial effects of the present invention are as follows: The present invention can achieve real-time monitoring of key system resources such as CPU, memory, disk, and network, as well as various middleware and databases. Combining the visualization dashboard and intelligent alerting mechanism of Grafana, it can provide early warnings for clusters and user applications, and provide data support for problem troubleshooting.

[0019] The advantages of the present invention are specifically reflected in: Comprehensive monitoring: By integrating Prometheus, various Exporters, and eBPF, it can comprehensively monitor key resources such as CPU, memory, disk, and network at the Node and Pod levels, providing fine-grained data collection.

[0020] Real-time performance: eBPF technology supports real-time acquisition of network status, reduces data collection latency, and enables operation and maintenance personnel to quickly respond to problems.

[0021] Visualization analysis: Combining Grafana to build an intuitive monitoring dashboard and being able to view data for a past period of time, which is convenient for users to quickly view the system situation, analyze problems, and provide data support for future system prediction.

[0022] Intelligent alerting: Through flexible alert rule configuration and various notification methods, it ensures that operation and maintenance personnel can timely learn about potential problems and reduce the false alarm rate.

[0023] Dynamic scalability: The system architecture is flexible and can dynamically adjust monitoring metrics and alert rules according to needs to adapt to the constantly changing cloud-native environment. Description of the Drawings

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0025] Figure 1 This is the flowchart of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0027] It should be noted that: similar reference numerals and letters denote similar items in the following accompanying drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0028] The following will describe in detail some implementation manners of the present invention with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0029] See Figure 1 , a method for improving the observability of a cloud-native platform, including the following steps: S1: Integrate Prometheus with Exporter; S2: Obtain network information in real time through eBPF; S3: Collect and store data; S4: Build a Grafana monitoring dashboard; S5: Configure an intelligent alarm mechanism to achieve real-time monitoring of the cloud-native platform.

[0030] This method comprehensively uses Prometheus, Alertmanager, and various Exporters and eBPF technologies to monitor key resources such as CPU, memory, disk, and network in real time, and combines Grafana to build a visual monitoring dashboard and an intelligent alarm system. The purpose of this method is to improve the observability, operation and maintenance efficiency, and system warning of cloud-native applications, and provide effective data support for troubleshooting and performance optimization.

[0031] In this embodiment, step S1 is specifically as follows: In a cloud-native environment, select a suitable node to deploy Prometheus and configure it as the monitoring system center; deploy Kube-state-metrics in the cluster to monitor pod resources and anomalies. Select Node Exporter to monitor resource usage at the node level, including CPU, memory, disk, etc.; select Application Exporter, an exporter developed according to specific application requirements, to collect specific business metrics, such as kafka-exporter, redis-exporter, etc.

[0032] In this embodiment, step S2 is specifically as follows: Obtain real-time network traffic, socket connection information, network latency, packet loss, etc. through eBPF. Using eBPF programs, detailed data of each network connection can be monitored, including the latency and traffic conditions of each request. eBPF (Extended Berkeley Packet Filter) is a powerful kernel technology that allows users to run small programs in the kernel to monitor and analyze system behavior.

[0033] In this embodiment, step S3 is specifically as follows: Prometheus pulls data from various Exporters through the HTTP protocol to form time-series data, which is stored in the TSDB library to support high-performance real-time query and analysis. Configure the scraping interval and data storage strategy of Prometheus to optimize performance and resource usage.

[0034] Furthermore, Prometheus uses local disk to store metric data, and there may be some challenges for large-scale long-term storage and horizontal scaling. To solve this problem, the remote storage adapter of Prometheus can be used to push metric data to an external storage system, such as cloud storage or a distributed database. By pushing data to an external storage system, the problem of limited local storage capacity of Prometheus can be solved, and long-term storage and horizontal scaling can be achieved. To improve availability, the federation function and multi-instance deployment of Prometheus can be used to achieve high availability. Through the federation function, multiple Prometheus instances can be aggregated into a whole for global query and monitoring. Multi-instance deployment and load balancing: Deploy multiple Prometheus instances and use a load balancer (such as Nginx or HAProxy) to distribute requests to these instances to ensure high availability and fault recovery capabilities.

[0035] In addition, Prometheus mainly focuses on real-time monitoring and short-term storage, and has relatively limited support for long-term storage and historical data query. To address this issue, Prometheus can be used in combination with other long-term storage systems (such as time series databases like InfluxDB, OpenTSDB) or object storage (such as Amazon S3). By configuring the Remote Write Adapter of Prometheus, metric data can be written to an external long-term storage system, and specialized tools or query languages can be used for historical data query and analysis. Utilize the data export of Prometheus: Prometheus supports data export in various formats, such as CSV, JSON, Prometheus data format, etc. By exporting data regularly, the data can be retained in an external system for long-term storage and historical data query.

[0036] In this embodiment, step S4 is specifically as follows: Deploy Grafana and configure Prometheus as the data source, design a personalized monitoring dashboard to display the changes in key metrics, including: CPU usage rate: Real-time monitor the CPU load and analyze the resource consumption of each service.

[0037] Memory usage: Monitor the memory usage and promptly detect memory leak problems.

[0038] Disk I / O performance: Monitor the disk read and write performance and evaluate the health status of the storage system.

[0039] Network traffic and latency: Display network traffic, latency, and packet loss rate based on network data obtained through eBPF.

[0040] Pod exception information: Such as abnormal Pod restarts, evictions, resource shortages, etc.

[0041] Metrics of various middleware itself: Such as kafka production and consumption situations, redis hit situations.

[0042] In this embodiment, step S5 is specifically as follows: Configure alerting rules in Prometheus and configure Alertmanager to define the recipients, notification methods, and policies for repeated alerts. Alerting rules are as follows: Trigger an alert when the CPU usage rate exceeds 85%.

[0043] Trigger an alert when the memory usage rate exceeds 90%.

[0044] Trigger an alert when the available disk space is below 10%.

[0045] An alarm is triggered when the network latency exceeds the set threshold.

[0046] An abnormal situation occurs in the Pod.

[0047] The present invention also provides a system for improving the observability of a cloud native platform, which is used to implement the method for improving the observability of a cloud native platform described above, including an integrated data collection module, a visualization monitoring module, and an intelligent alarm module. Among them, The integrated data collection module is used to use Prometheus as the central monitoring system, combine kube-stat-metrics and various exporters to collect key metrics such as CPU, memory, disk, and network of the cluster, nodes, and pods, as well as specific metrics of middleware and databases, and use eBPF technology to obtain information such as network traffic, latency, and packet loss in real time.

[0048] The visualization monitoring module is used to build a real-time monitoring dashboard using Grafana to display the data collected by Prometheus, supporting various chart forms. And it provides a user-friendly interface to facilitate data analysis and problem troubleshooting for operation and maintenance personnel.

[0049] The intelligent alarm module is used to implement alarms through Prometheus Alertmanager, and users can configure alarm rules according to specific needs. It supports configuring alarm strategies. For example, for general alarms, a notification message will only be sent when the threshold is reached for 10 minutes to avoid alarm storms. It supports multiple alarm notification methods, including email, DingTalk, WeChat, Webhook, etc., to ensure timely information transmission.

[0050] The present invention can realize real-time monitoring of key system resources such as CPU, memory, disk, and network, as well as various middleware and databases. Combined with the visualization dashboard and intelligent alarm mechanism of Grafana, it can give early warnings to the cluster and user applications, providing data support for problem troubleshooting. It is mainly reflected in: comprehensive monitoring: by integrating Prometheus, various exporters, and eBPF, it can comprehensively monitor key resources such as CPU, memory, disk, and network at the Node and Pod levels, providing fine-grained data collection. Real-time performance: eBPF technology supports real-time acquisition of network status, reducing data collection latency, enabling operation and maintenance personnel to quickly respond to problems. Visualization analysis: combined with Grafana to build an intuitive monitoring dashboard, and the data of the past period can be viewed, which is convenient for users to quickly view the system situation, analyze problems, and provide data support for future system prediction. Intelligent alarm: through flexible alarm rule configuration and multiple notification methods, it ensures that operation and maintenance personnel can timely learn about potential problems and reduce the false alarm rate. Dynamic scalability: the system architecture is flexible and can dynamically adjust monitoring metrics and alarm rules according to needs to adapt to the ever-changing cloud native environment.

[0051] For the foregoing embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.

[0052] In the above embodiments, the basic principles, main features and advantages of the present invention are described. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications and changes made by those skilled in the art that do not depart from the spirit and scope of the present invention should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for enhancing the observability of a cloud-native platform, characterized in that, It includes the following steps: S1: Integrate Prometheus with Exporter; S2: Obtain network information in real time through eBPF; S3: Collect and store data; S4: Build a Grafana monitoring dashboard; S5: Configure an intelligent alert mechanism to achieve real-time monitoring of the cloud-native platform.

2. The method for improving the observability of a cloud native platform according to claim 1, wherein Step S1 includes the following sub-steps: S11: In the cloud-native platform, deploy Prometheus, configure it as a monitoring center, and deploy Kube-state-metrics in the cluster to monitor pod resources and anomalies; S12: Select Exporter to monitor resource usage at the node level and business metrics.

3. A method for enhancing the observability of a cloud native platform according to claim 2, characterized in that, Specifically, step S12 is: Monitor resource usage at the node level, including CPU, memory, and disk, through Node Exporter; collect specific business metrics, including kafka-exporter and redis-exporter, through Application Exporter.

4. A method for improving the observability of a cloud native platform according to claim 3, characterized in that, Specifically, step S2 is: Obtain real-time network traffic, socket connection information, network latency, and packet loss information through eBPF.

5. A method for improving the observability of a cloud-native platform as claimed in claim 4, characterized in that, Specifically, step S3 is: Prometheus pulls data from various Exporters through the HTTP protocol to form time-series data and stores it in the TSDB library, supporting high-performance real-time query and analysis.

6. The method for enhancing the observability of a cloud native platform according to claim 5, wherein Step S3 also includes: Configure the scraping interval and data storage strategy of Prometheus to optimize performance and resource usage.

7. A method for enhancing the observability of a cloud-native platform as described in claim 6, characterized in that, Specifically, step S4 is: Deploy Grafana and configure Prometheus as the data source, and design a personalized monitoring dashboard to display the changes in key metrics, where the key metrics include CPU usage rate, memory usage, disk I / O performance, and network traffic and latency.

8. A method for improving the observability of a cloud native platform according to claim 7, characterized in that, The key metrics also include Pod anomaly information and middleware itself metrics. Among them, Pod anomaly information includes Pod abnormal restart, being evicted, and resource shortage; middleware itself metrics include kafka production and consumption situation and redis hit situation.

9. A method for improving the observability of a cloud native platform according to claim 8, characterized in that, Specifically, step S5 is: Configure alert rules in Prometheus and configure Alertmanager to define alert receivers, notification methods, and repeated alerts; the alert rules include CPU usage rate threshold, memory usage rate threshold, disk available space threshold, network latency threshold, and Pod anomaly situation.

10. A system for enhancing the observability of a cloud-native platform, which is used to implement the method for enhancing the observability of a cloud-native platform described in any one of claims 1 to 9, characterized in that, It includes an integrated data collection module, a visualization monitoring module, and an intelligent alert module. Among them, The integrated data collection module uses Prometheus as the central monitoring system, combines kube-stat-metrics and various Exporters to collect key metrics of CPU, memory, disk, and network of the cluster, nodes, and pods, and obtains real-time network traffic, latency, and packet loss information through eBPF technology; Visualization monitoring module, which is used to build a real-time monitoring dashboard using Grafana, display the data collected by Prometheus, support various chart forms, and provide a user interface to facilitate data analysis and problem troubleshooting by operation and maintenance personnel; Intelligent alarm module, which is used to implement alarms through Prometheus Alertmanager and support the configuration of alarm strategies to ensure the timely transmission of information.

Citation Information

Cited By

  • Observable system for container environment

    CN121151253A