Method, device and equipment for realizing Prometheus cluster

By configuring tags for Prometheus instances and pre-calculated sharding and hierarchical configurations, combined with the main and standby instance mechanism, the problems of complex operation and maintenance and low reliability of native Prometheus clusters are solved, and cluster management with high reliability, easy scalability and low resource utilization are achieved, and query performance is improved.

CN120256367APending Publication Date: 2025-07-04CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290484.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The operation and maintenance of native Prometheus clusters are complex and have low reliability, making it difficult to meet the requirements of high reliability, ease of scalability and low resource utilization.

Method used

By configuring the first tag, the second tag and the third tag for the Prometheus instance, pre-calculated sharding and hierarchical configuration using the acquisition configuration template and environment variables, the query components are configured for data query and aggregation, and the main and backup instance mechanism is introduced, and the tag information is used for deduplication and backup.

Benefits of technology

It simplifies the management of Prometheus cluster, improves reliability and scalability, reduces memory loss, improves query performance, and realizes high-availability and high-performance cluster management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256367A_ABST
    Figure CN120256367A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device and equipment for realizing a Prometheus cluster. The method comprises the following steps: configuring a first label, a second label and a third label for each Prometheus instance according to an acquisition configuration template and environment variables during operation; the first label and the second label are used for determining a collection target; the third label is used for identifying the type of the Prometheus instance; according to the pre-calculation configuration template, pre-calculation fragmentation configuration and pre-calculation layering configuration are carried out on the Prometheus instance; and a query component is configured to complete query and aggregation of Prometheus data. Therefore, duplicate removal and backup can be carried out in advance by utilizing Prometheus label information, and the standby instance can be queried only when a certain main instance does not respond, so that half of memory loss can be saved, duplicate removal logic is avoided, and the query performance is better; and the requirements of high reliability, easy expansion, low resource utilization rate and the like are well met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for implementing a Prometheus cluster. Background Art

[0002] As an open-source monitoring system and alerting system, the Prometheus cluster is applicable to monitoring hardware metrics such as servers and also to monitoring highly dynamic service-oriented architectures. Therefore, it is widely used in various monitoring systems.

[0003] The cluster methods provided by native Prometheus include: collecting configurations and using the hash remainder method for sharding. Since Prometheus itself provides a federated query function, the sharded Prometheus can be aggregated.

[0004] However, the cluster method of native Prometheus relies entirely on manual adjustment for collecting configurations and federated deployment, which is operationally complex and has low reliability. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for implementing a Prometheus cluster that can meet requirements such as high reliability, easy expansion, and low resource utilization.

[0006] In a first aspect, this application provides a method for implementing a Prometheus cluster, and the method includes:

[0007] Configuring a first label, a second label, and a third label for each Prometheus instance according to a collection configuration template and runtime environment variables; wherein, the first label and the second label are used to determine a collection target; the third label is used to identify the type of the Prometheus instance; the type of the Prometheus instance includes a standby instance and a primary instance;

[0008] Performing precomputation sharding configuration and precomputation hierarchical configuration on the Prometheus instance according to a precomputation configuration template;

[0009] Configuring a query component to complete the query and aggregation of Prometheus data.

[0010] In one of the embodiments, the configuring a first label, a second label, and a third label for each Prometheus instance according to a collection configuration template and runtime environment variables includes:

[0011] Deploy Prometheus instances on each target server and set the IP address of each target server as the environment variable when the Prometheus instance runs;

[0012] Load the collection configuration template in a hot reload manner to generate the first label, the second label, and the third label corresponding to each Prometheus instance;

[0013] Form a hash ring according to the Prometheus instance specified by the second label;

[0014] Generate multiple collection targets according to the collection configuration template, and each collection target includes address information;

[0015] Calculate the hash value according to the address information included in the collection target, and find the first label corresponding to the closest hash value in the hash ring in a clockwise direction;

[0016] Only retain the collection target when the first label corresponding to the found hash value is the same as the first label corresponding to the collection target.

[0017] In one embodiment, the pre-computation sharding configuration of the Prometheus instance according to the pre-computation configuration template includes:

[0018] Identify the Prometheus instances that need to be sharded and the Prometheus instances that do not need to be sharded by configuring the fourth label, and identify the pre-computation tasks for querying local pre-computed tasks and querying remote data by adding new fields;

[0019] For the Prometheus instances that need to be sharded, determine the pre-computation tasks to be executed by using the consistent hashing allocation method;

[0020] For the Prometheus instances that do not need to be sharded, execute each pre-computation task.

[0021] In one embodiment, the pre-computation hierarchical configuration of the Prometheus instance according to the pre-computation configuration template includes:

[0022] Plan a hierarchical label for each Prometheus instance; among them, all pre-computation queries will only match the Prometheus in the upper layer.

[0023] In one embodiment, the configuration of the query component to complete the query and aggregation of Prometheus data includes:

[0024] Obtain the query filtering conditions and determine the Prometheus instances that meet the hierarchy;

[0025] Group according to the group name label of the Prometheus instance, and sort the Prometheus instances in the same group in ascending order according to the priority within the group;

[0026] Execute the query task in a concurrent manner for multiple groups, and query data in sequence within each group;

[0027] When a Prometheus instance in the same group successfully responds, stop querying data for the remaining Prometheus instances in the same group;

[0028] Aggregate the query data fed back by all groups to obtain the aggregated query data.

[0029] In one embodiment, the method further includes:

[0030] Specify the address of the standby instance by configuring the fifth label;

[0031] The primary instance sends a heartbeat signal to the address specified by the fifth label at a preset period;

[0032] If the standby instance does not receive a heartbeat signal within a preset time, it is determined that the primary instance has an abnormality;

[0033] In the case where the primary instance has an abnormality, the standby instance resumes the original precomputation frequency;

[0034] In the case where the primary instance does not have an abnormality, the standby instance executes a downsampling strategy; the downsampling strategy includes: feeding back data at least once within the retrospective time window.

[0035] In a second aspect, the present application further provides an implementation device for a Prometheus cluster, and the device includes:

[0036] An acquisition configuration module, configured to configure a first label, a second label, and a third label for each Prometheus instance according to an acquisition configuration template and runtime environment variables; wherein, the first label and the second label are used to determine an acquisition target; the third label is used to identify the type of the Prometheus instance; the type of the Prometheus instance includes a standby instance and a primary instance;

[0037] A precomputation configuration module, configured to perform precomputation sharding configuration and precomputation hierarchical configuration on the Prometheus instance according to a precomputation configuration template;

[0038] A query module, configured to configure a query component to complete the query and aggregation of Prometheus data.

[0039] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0040] Configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the environment variables at runtime. Among them, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances;

[0041] Perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template;

[0042] Configure a query component to complete the query and aggregation of Prometheus data.

[0043] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0044] Configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the environment variables at runtime. Among them, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances;

[0045] Perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template;

[0046] Configure a query component to complete the query and aggregation of Prometheus data.

[0047] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0048] Configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the environment variables at runtime. Among them, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances;

[0049] Perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template;

[0050] Configure a query component to complete the query and aggregation of Prometheus data.

[0051] The above-mentioned method, device, computer device, computer-readable storage medium, and computer program product for implementing a Prometheus cluster configure the first label, second label, and third label for each Prometheus instance according to the collection configuration template and the environmental variables at runtime; wherein, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances; thus, each Prometheus instance can be managed through labels, making the management of the Prometheus cluster simpler and easier to expand. Perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template; thus, abandon the collection sharding ability of Prometheus itself, and add data sharding management and hierarchical management for pre-computation, which is convenient for implementing a highly available and high-performance cluster. Configure a query component to complete the query and aggregation of Prometheus data. Thus, duplicate elimination and backup can be performed in advance using the label information of Prometheus, and only when a primary instance does not respond will the standby instance be queried, which can save half of the memory loss, avoid the duplicate elimination logic, and has better query performance; it well meets the requirements of high reliability, easy expansion, and low resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a schematic diagram of the architecture of a Prometheus cluster in an embodiment;

[0054] Figure 2 It is a schematic flowchart of the method for implementing a Prometheus cluster in an embodiment;

[0055] Figure 3 It is a schematic diagram of the principle of a native Prometheus-provided cluster in an embodiment;

[0056] Figure 4 Flow diagram of the implementation method of the Prometheus cluster in another embodiment;

[0057] Figure 5 Flow diagram of the configuration and task execution process of a Prometheus cluster in one embodiment;

[0058] Figure 6 Schematic diagram of the interaction process between the query component query and the Prometheus cluster in one embodiment;

[0059] Figure 7 Hierarchical architecture diagram of the primary cluster and the standby cluster in one embodiment;

[0060] Figure 8 Structural block diagram of the implementation device of the Prometheus cluster in one embodiment;

[0061] Figure 9 Structural block diagram of the implementation device of the Prometheus cluster in another embodiment;

[0062] Figure 10 Internal structure diagram of a computer device in one embodiment. Detailed implementation

[0063] In order to make the purpose, technical solutions and advantages of this application clearer, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0064] Exemplarily, such as Figure 1As shown in the figure, an architecture of a Prometheus cluster is provided, which may include: a Prometheus primary cluster and a Prometheus standby cluster. The configurations of each Prometheus in the primary and standby clusters are completed through configuration files, and query tasks for Prometheus are executed based on a query component to complete the query and aggregation of data of each Prometheus. In the Prometheus primary cluster, grouping can also be performed according to requirements, for example, divided into n groups (group 1 to group n, that is, [group-1] to [group-n]). Similarly, in the Prometheus standby cluster, grouping can also be performed according to requirements, for example, divided into n groups (group 1 to group n, that is, [group-1] to [group-n]). Exemplarily, [group-1] in the Prometheus primary cluster and [group-1] in the Prometheus standby cluster are in a primary-standby relationship, and the heartbeat signal is used to determine whether the primary Prometheus has an abnormality. When the primary Prometheus has an abnormality, the standby Prometheus resumes the original collection frequency and pre-computation frequency to execute collection and calculation tasks.

[0065] In an exemplary embodiment, as Figure 2 shown, a method for implementing a Prometheus cluster is provided. Taking the cluster in Figure 1 as an example for illustration, it includes the following steps 201 to step 203. Among them:

[0066] Step 201, configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the runtime environment variables.

[0067] Prometheus in this embodiment is an open-source monitoring tool that provides functions such as data collection, data aggregation and cleaning. Prometheus comes with a time series database for storing monitoring data, making data storage and query more convenient compared to other monitoring tools.

[0068] Exemplarily, as Figure 3 shown, a way to provide a cluster by native Prometheus is given. Briefly speaking, its principle is: through collection configuration and using the method of hash remainder for sharding, and using the federated query function of Prometheus itself to aggregate the sharded Prometheus. However, the way to provide a cluster by native Prometheus entirely relies on manual adjustment of collection configuration and federated deployment, with complex operation and maintenance and low performance.

[0069] To address the problems existing in the prior art, in this embodiment, the collection configuration template assigns tags to each Prometheus instance, so that the collection prediction allocation, layering, and primary / standby instance configuration can be implemented with the help of these tags. Among them, the first tag and the second tag are used to determine the collection target; the third tag is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances.

[0070] Exemplarily, to implement functions such as sharding, layering, and primary / standby, the Prometheus instance needs to know its own and cluster member information. Optionally, combining the template syntax and runtime environment variables, the Prometheus tags are planned through configuration.

[0071] Exemplarily, Prometheus instances are deployed on each target server, and the IP addresses of each target server are set as the environment variables during the runtime of the Prometheus instance. Optionally, taking the example of deploying two Prometheus instances on two servers, first, set the IP addresses of the servers to 10.117.19.18 and 10.117.19.19 respectively. Specify the tag template for Prometheus. After startup, the Prometheus instance on 10.117.19.18 is automatically tagged with group_name=p0 (group name) and group_priority=0 (group priority), and the Prometheus instance on 10.117.19.19 is automatically tagged with group_name=p1 and group_priority=1.

[0072] Exemplarily, the collection configuration template is loaded in a hot reload manner to generate the first tag, the second tag, and the third tag corresponding to each Prometheus instance. Among them, hot reload means that the configuration associated with the tag will also be adjusted along with the collection configuration template, so as to achieve the purpose of managing the Prometheus cluster with the collection configuration template and simplify the operation and maintenance.

[0073] Exemplarily, according to the Prometheus instances specified by the second tag, a hash ring is formed; multiple collection targets are generated according to the collection configuration template, and each collection target contains address information; the hash value is calculated according to the address information contained in the collection target, and the first tag corresponding to the closest hash value is found clockwise in the hash ring; the collection target is retained only when the first tag corresponding to the found hash value is the same as the first tag corresponding to the collection target.

[0074] Optionally, plan the first label (node_name) and the second label (node_member) for the Prometheus instance. Herein: node_name is the name of this Prometheus instance, and node_member is the name of all Prometheus instances in this cluster. Use the values of these two labels and the Prometheus collection target to implement consistent hashing allocation. The specific allocation process is as follows:

[0075] 1) Parse the management configuration template, and obtain the Prometheus label information corresponding to this instance according to the environment variables.

[0076] 2) Form a hash ring according to the members specified by the "node_member" label.

[0077] 3) The Prometheus collection configuration will finally generate individual collection targets, and each collection target will contain __address__ for specifying the remote data address. Calculate the hash value according to the value of __address__, and clockwise find the closest hash value in the hash ring generated in step 2), and find the node_name corresponding to this hash value.

[0078] 4) Compare whether the "node_name" label of itself is consistent with that in step 3). If it is consistent, keep this collection target; if it is inconsistent, discard this collection target.

[0079] In this embodiment, use the Prometheus label information generated by the configuration template and environment variables to perform consistent hashing allocation on the collection target, optimize the Prometheus cluster sharding, so as to achieve automatic reloading configuration and allocation, and simplify the operation and maintenance; compared with the native hash allocation, the consistent hashing allocation is more even, so the load of Prometheus will be more balanced.

[0080] Step 202, perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template.

[0081] In this embodiment, try to perform the pre-computation on the Prometheus instance locally as much as possible. For other pre-computations, process them hierarchically according to the data volume, and try to place the externally exposed data in one layer to save the most resources and achieve the best query performance. In addition, for the pre-computation across Prometheus instances, implement primary and standby management to facilitate further resource compression.

[0082] Exemplarily, the fourth label is configured to identify the Prometheus instances that need to be sharded and those that do not need to be sharded, and new fields are added to identify the pre-computation tasks for querying local data and the pre-computation tasks for querying remote data; for the Prometheus instances that need to be sharded, the consistent hashing allocation method is used to determine the pre-computation tasks to be executed; for the Prometheus instances that do not need to be sharded, each pre-computation task is executed.

[0083] In this embodiment, a fourth label (rule_shard) is planned for the Prometheus instance to identify whether consistent hashing allocation needs to be performed for pre-computation. When rule_shard = false, no sharding is required, that is, each pre-computation needs to be executed; when rule_shard = true, consistent hashing allocation needs to be performed. Among them, the method of consistent hashing allocation is the same as that of Prometheus collection sharding, which also uses the node_name and node_member labels to calculate the hash value of the pre-computation name to find the corresponding node_name.

[0084] In this embodiment, since Prometheus pre-computation can only query locally, it needs to be modified to enable pre-computation to support querying remote data sources and can be specified at the rule granularity. To achieve such a function, a new from_remote field is configured for each pre-computation. When from_remote = false, query locally; when from_remote = true, query remote data.

[0085] Optionally, a hierarchical label (level) is planned for each of the Prometheus instances; among them, all pre-computation queries will only match the Prometheus at the upper level. Such a planning method can reduce the amount of data queried. The data at the current level only depends on the upper level, allowing Prometheus to avoid unnecessary indexing. Theoretically, the higher the level, the greater the dimension of data aggregation and the smaller the amount of data.

[0086] Exemplarily, assume that 2 Prometheus are planned for data collection, namely 10.117.19.18 and 10.117.19.19. These two are used to collect the edge bandwidth of two provinces. The sharding is evenly distributed to these two Prometheus instances at the granularity of edge nodes and is at layer 0. 2 Prometheus are planned for pre-computation, namely 10.117.19.20 and 10.117.19.21, which aggregate the bandwidth at the province granularity and are at layer 1. After performing the pre-computation sharding configuration and pre-computation layer configuration on the Prometheus instances according to the pre-computation configuration template, the following effects can be achieved:

[0087] 1) The data collection of 10.117.19.18 and 10.117.19.19 can be distributed according to nodes. When rule_shard = false, the pre-computation is no longer distributed, and each Prometheus instance will perform the nodeBW pre-computation, and the data source for query is local.

[0088] 2) 10.117.19.20 and 10.117.19.21 only perform pre-computation. FJBW (project number) and ZJBW (project number) are respectively assigned to 10.117.19.20 and 10.117.19.21, and the hierarchical filtering conditions will be automatically added according to the hierarchy. The finally generated pre-computation query statements are: sum(nodeBw{province="A Province",__identify_level="0"}) and sum(nodeBw{province="B Province",__identify_level="0"}). The __identify prefix is used by the query component to accurately match the Prometheus instances that meet the conditions.

[0089] Optionally, the concept of a primary and standby cluster can also be introduced to reduce resource waste. For example, the same number of machines are planned for the primary and standby clusters. To ensure the consistency of the data collection and pre-computation of the consistent hashing distribution, the primary and standby of the Prometheus instances are determined during the planning. Add a third label (node_type) to specify the type of the Prometheus instance: when node_type = master, it is the primary instance; when node_type = slave, it is the standby instance.

[0090] Step 203, configure the query component to complete the query and aggregation of Prometheus data.

[0091] In this embodiment, only one instance of Prometheus acting as the primary and standby for each other is required to respond, thus reducing the data volume by half. The Prometheus instance to be queried can also be located according to the hierarchical information of Prometheus, saving resources.

[0092] In this embodiment, by connecting the self-developed query component (query) to the Prometheus remote read interface, pre-computation across Prometheus instances can be achieved, and the deployment and operation are simple. That is to say, the query component in this embodiment can implement the functions of raw data query and merging. Among them, promQL is still executed on Prometheus; before querying, Prometheus instances that meet the conditions can be filtered according to the Prometheus instance information, thus avoiding invalid queries.

[0093] Exemplarily, obtain query filtering conditions to determine Prometheus instances that meet the hierarchy; group according to the group name labels of the Prometheus instances, and sort the Prometheus instances in the same group in ascending order according to the priority within the group; execute the query task in a concurrent manner for multiple groups, and query data in order within each group; when a Prometheus instance in the same group successfully responds, stop querying the remaining Prometheus instances in the same group; aggregate the query data fed back by all groups to obtain the aggregated query data.

[0094] Exemplarily, plan a group_name, group_priority, and hierarchical label (level) for each Prometheus, and group, prioritize, and hierarchically classify the primary and standby Prometheus. The primary and standby Prometheus instances have the same group_name and group_priority (0 for the primary and 1 for the standby), and the smaller the number, the higher the priority. When querying, first obtain the __identify_level="xx" filtering condition, match the value of this label with the hierarchical label of Prometheus, select the Prometheus instance that meets the hierarchy, and then remove the __identify_level label from the query statement. Group by the group_name in the Prometheus instance information label, and then sort in ascending order using the group_priority label. Query data in order for each group_name, and execute multiple group_names concurrently. When the prometheus instances in a group (i.e., the same group_name, sequential query) successfully respond, stop querying this group. In this way, the standby will not be queried when the primary is normal, and the standby will continue to be queried when the primary is abnormal, thereby improving availability. After the Prometheus instances of all group_names return data, merge the data and return it to Prometheus. Prometheus then performs operations on the promQL statement based on the underlying data.

[0095] In the above method for implementing the Prometheus cluster, configure the first label, second label, and third label for each Prometheus instance according to the collection configuration template and runtime environment variables; wherein, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the type of the Prometheus instance includes a standby instance and a primary instance; thus, each Prometheus instance can be managed through labels, making the management of the Prometheus cluster simpler and easier to expand. Configure pre-computation sharding and pre-computation hierarchical configuration for the Prometheus instance according to the pre-computation configuration template; thereby abandoning the collection sharding ability of Prometheus itself, adding data sharding management and hierarchical management for pre-computation, which is convenient for implementing a highly available and high-performance cluster. Complete the query and aggregation of Prometheus data by configuring a query component. Thus, duplicate elimination and backup can be performed in advance using the label information of Prometheus. Only when a certain primary instance does not respond will the standby instance be queried, which can save half of the memory loss, avoid the duplicate elimination logic, and have better query performance; it well meets the requirements of high reliability, easy expansion, and low resource utilization.

[0096] In an exemplary embodiment, as Figure 4 shown, a method for implementing a Prometheus cluster is provided. Taking the cluster in Figure 1 as an example, the following steps 401 to 408 are included. Among them:

[0097] Step 401: Configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the runtime environment variables.

[0098] Step 402: Perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instances according to the pre-computation configuration template.

[0099] Step 403: Configure a query component to complete the query and aggregation of Prometheus data.

[0100] In this embodiment, for the specific implementation process and technical effects of steps 401 to 403, please refer to Figure 2 the relevant descriptions of steps 201 to 203 in the method embodiment shown, which will not be elaborated here.

[0101] Step 404: Specify the address of the standby instance by configuring the fifth label.

[0102] In this embodiment, the fifth label (heartbeat_addr) can be added to specify the address of the standby instance.

[0103] Step 405: The master instance sends a heartbeat signal to the address specified by the fifth label at a preset period.

[0104] In this embodiment, by using the addition of the third label (node_type) to specify the type of the Prometheus instance, when node_type = master, a heartbeat is sent to the address specified by heartbeat_addr once every preset period (for example, 1s).

[0105] Step 406: If the standby instance does not receive a heartbeat signal within a preset time, it is determined that the master instance has an exception.

[0106] In this embodiment, when node_type = slave, the Prometheus instance detects the received heartbeat time. If no heartbeat is received within a preset time (for example, more than 3s), it is considered that the master is abnormal, and operations such as restoring the original pre-computation frequency and writing data remotely (if configured) are performed. Stop the above operations after receiving a heartbeat again.

[0107] Step 407, when an exception occurs in the primary instance, the standby instance resumes its original pre-computed frequency.

[0108] Step 408, when no exception occurs in the primary instance, the standby instance executes a downsampling strategy.

[0109] Among them, the downsampling strategy includes: feeding back data at least once within the backtracking time window.

[0110] In this embodiment, the instant query of Prometheus can retrieve data within a specified backtracking time (default 5 minutes). Exemplarily, if there is a data point at exactly 11 o'clock, then this data point can be retrieved before 11:05. The standby Prometheus is only for providing hot data so that data can be immediately provided for querying after the primary Prometheus fails, achieving high availability. Therefore, when the primary Prometheus is normal, the standby Prometheus does not need to perform data collection and calculation according to the configured collection frequency and pre-computed frequency, and only needs to ensure that there is data within the specified backtracking time window.

[0111] Optionally, the implementation process of the downsampling strategy includes:

[0112] 1) Obtain the backtracking time configuration, which is generally the Prometheus startup parameter (delta-back).

[0113] 2) The standby Prometheus receives the heartbeat of the primary Prometheus at preset intervals; if the heartbeat of the primary Prometheus is normal, then set delta-back = Prometheus (the specified backtracking time); if the heartbeat of the primary Prometheus is abnormal, then set delta-back = 0.

[0114] 3) When delta-back = 0, the pre-computation is still executed regularly according to the configured frequency. When delta-back > 0, determine whether the last execution time and the current time are greater than delta-back. If so, execute; if not, skip this execution.

[0115] Exemplarily, Figure 5 A flowchart showing the configuration and task execution process of a Prometheus cluster is provided. Figure 6 A schematic diagram showing the interaction process between the query component query and the Prometheus cluster is provided. Exemplarily, Figure 7 A schematic diagram showing the hierarchical architecture of the primary cluster and the standby cluster is provided.

[0116] Optionally, in combination with Figures 5 - 7For the solution shown, when it is necessary to collect traffic information of 30 nodes in multiple provinces (where the dimension of each traffic information is: domain name, node IP, province, each domain name will appear on multiple nodes, and one province contains multiple nodes), obtain the bandwidth of each domain name and each province, and obtain the total bandwidth, the following plan can be made:

[0117] First, plan to use 6 Prometheus to collect the traffic information of these 30 nodes. Among them, 3 are a cluster, in a master-backup relationship, and are at the 0th layer.

[0118] The main cluster machines are: 10.110.10.10, 10.110.10.11, 10.110.10.12

[0119] The backup cluster machines are: 10.110.10.13, 10.110.10.14, 10.110.10.15

[0120] Among them: 10.110.10.10 and 10.110.10.13, 10.110.10.11 and 10.110.10.14, 10.110.10.12 and 10.110.10.15 are in a master-backup relationship, and their collected data is consistent.

[0121] Secondly, plan to use 4 Prometheus to calculate the bandwidth at the domain name (host) granularity and province (province) granularity, query the data of the above 6 Prometheus, and are at the 1st layer, depending on the 0th layer. Among them, 2 are a cluster, in a master-backup-master-backup relationship.

[0122] The main cluster machines are: 10.110.10.16, 10.110.10.17

[0123] The backup cluster machines are: 10.110.10.18, 10.110.10.19

[0124] Among them: 10.110.10.16 and 10.110.10.18, 10.110.10.17 and 10.110.10.19 are in a master-backup relationship.

[0125] Finally, plan to use 2 Prometheus to calculate the total bandwidth. On the basis of 2, aggregate all host bandwidths, and are at the 2nd layer, depending on the 1st layer. The two are in a master-backup relationship.

[0126] Master: 10.110.10.20

[0127] Backup: 10.110.10.21

[0128] Configure the above-mentioned Prometheus configuration management configuration template files, pre-computation configuration templates, deploy queries (specify the Prometheus address for the query, i.e., 10.110.10.10 to 10.110.10.21. The query will periodically obtain label information from Prometheus and update it), and deploy Prometheus.

[0129] Among them, deploying Prometheus includes adding the following parameters on the native basis:

[0130] --instance-template = instance.yml

[0131] --query-remote = query address

[0132] For 10.110.10.10 to 10.110.10.15, traffic data on the node needs to be collected, and the collection configuration needs to be specified. The other collection configurations are empty and only pre-computation is required.

[0133] Among them, the technical effects after the above configuration are as follows:

[0134] 1) In terms of collection, 10.110.10.10 to 10.110.10.15 form two clusters, which are in a master-backup relationship. The master and backup machines in each cluster perform consistent hashing based on node_name and node_member and are assigned the same collection targets.

[0135] 2) In terms of pre-computation, except for the 0th level (rule_shard = false is not allocated and full-scale pre-computation is performed), other levels also perform consistent hashing based on node_name and node_member. There are two machines in each cluster at the 1st level, and a total of two pre-computation rules need to be executed, with each machine assigned one.

[0136] 3) The master-backup relationship of the machines refers to node_type and hearbeat_addr in the configuration. Among them, the master (node_type = master) sends a heartbeat to the machine (backup) specified by hearbeat_addr every 1s.

[0137] 4) When the backup machine receives the heartbeat, it performs pre-computation at a frequency of 5m (the default backtracking time of Prometheus).

[0138] 5) Pre-computation. The 0th level (from_remote: false) only queries the local storage. Other levels query data from the query (from_remote: true). Among them, the pre-computation configuration above the 1st level generates a matching label __identify_level="${level}-1" according to the value of the level label -1.

[0139] 6) After receiving the query request, the query first extracts the __identify_level label and matches it with the prometheus label level (obtaining updates from prometheus at regular intervals). If they match, it is retained. At the same time, grouping and sorting are performed according to group_name and group_priority. The final result is that only the main prometheus will be queried. For example, when performing pre-computation queries at the 1st level, after the above matching, the query will only query data from the prometheus at 10.110.10.10~10.110.10.13 and merge and return it.

[0140] In this embodiment, the sharding, layering, and primary-backup of the Prometheus cluster are managed through configuration, and each instance of Prometheus is managed through labels. It is simpler and more scalable than the native Prometheus federated cluster. The custom query component can utilize the label information of Prometheus to perform deduplication and backup in advance, and only query the standby Prometheus when a certain main Prometheus does not respond, thus saving half of the memory loss, avoiding the deduplication logic, and having better query performance. It also enables the pre-computation of the standby Prometheus to achieve the automatic downsampling function through the heartbeat detection method of the primary sending to the standby, thereby further saving resources.

[0141] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0142] Based on the same inventive concept, an embodiment of the present application further provides an implementation device for a Prometheus cluster for implementing the above-mentioned implementation method of the Prometheus cluster. The implementation solution provided by this device for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the Prometheus cluster implementation device provided below can refer to the limitations on the implementation method of the Prometheus cluster in the above text, and will not be repeated here.

[0143] In an exemplary embodiment, as Figure 8 shown, an implementation device for a Prometheus cluster is provided, including: a collection configuration module 801, a pre-computation configuration module 802, and a query module 803, where:

[0144] The collection configuration module 801 is configured to configure the first label, the second label, and the third label for each Prometheus instance according to the collection configuration template and the runtime environment variables; wherein, the first label and the second label are used to determine the collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances;

[0145] The pre-computation configuration module 802 is configured to perform pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instance according to the pre-computation configuration template;

[0146] The query module 803 is configured to configure a query component to complete the query and aggregation of Prometheus data.

[0147] Exemplarily, the collection configuration module 801 is specifically configured to deploy Prometheus instances on each target server, and set the IP address of each target server as the runtime environment variable of the Prometheus instance; load the collection configuration template in a hot-loading manner to generate the first label, the second label, and the third label corresponding to each Prometheus instance; form a hash ring according to the Prometheus instance specified by the second label; generate multiple collection targets according to the collection configuration template, and each collection target includes address information; calculate the hash value according to the address information included in the collection target, and find the first label corresponding to the closest hash value in the hash ring in a clockwise direction; only retain the collection target when the first label corresponding to the found hash value is the same as the first label corresponding to the collection target.

[0148] Exemplarily, the pre-computation configuration module 802 is specifically configured to identify Prometheus instances that need to be sharded and Prometheus instances that do not need to be sharded by configuring a fourth tag, and to identify pre-computation tasks for querying local data and pre-computation tasks for querying remote data by adding new fields; for Prometheus instances that need to be sharded, a consistent hashing allocation method is used to determine the pre-computation tasks to be executed; for Prometheus instances that do not need to be sharded, each pre-computation task is executed.

[0149] Exemplarily, the pre-computation configuration module 802 is specifically configured to plan a hierarchical tag for each of the Prometheus instances; among them, all pre-computation queries will only match the Prometheus at the upper level.

[0150] Exemplarily, the query module 803 is specifically configured to obtain query filtering conditions and determine Prometheus instances that meet the hierarchy; group them according to the group name tags of the Prometheus instances, and sort the Prometheus instances in the same group in ascending order according to the priority within the group; execute the query tasks in a concurrent manner for multiple groups, and query data in order within each group; when a Prometheus instance in the same group successfully responds, stop querying data for the remaining Prometheus instances in the same group; aggregate the query data fed back by all groups to obtain the aggregated query data.

[0151] In this embodiment, according to the collection configuration template and the environment variables at runtime, a first tag, a second tag, and a third tag are configured for each Prometheus instance; wherein, the first tag and the second tag are used to determine the collection target; the third tag is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances; thus, each Prometheus instance can be managed through tags, making the management of the Prometheus cluster simpler and easier to expand. According to the pre-computation configuration template, pre-computation sharding configuration and pre-computation hierarchical configuration are performed on the Prometheus instances; thus, the self-collection sharding ability of Prometheus is abandoned, and new data sharding management and hierarchical management for pre-computation are added, facilitating the implementation of a highly available and high-performance cluster. The query and aggregation of Prometheus data are completed by configuring a query component. Thus, duplicate elimination and backup can be performed in advance using the tag information of Prometheus. Only when a certain primary instance does not respond will the standby instance be queried, which can save half of the memory loss, avoid the duplicate elimination logic, and has better query performance; it can well meet the requirements of high reliability, easy expansion, and low resource utilization.

[0152] In another exemplary embodiment, as Figure 9 shown, an implementation device of a Prometheus cluster is provided. Based on the Figure 8 device shown, it may further include:

[0153] An address specifying module 804, configured to specify the address of the standby instance by configuring a fifth tag;

[0154] A sending module 805, configured to send a heartbeat signal from the primary instance to the address specified by the fifth tag at a preset period;

[0155] A judging module 806, configured to judge that the primary instance is abnormal when the standby instance does not receive a heartbeat signal within a preset time;

[0156] A recovery module 807, configured to, when the primary instance is abnormal, the standby instance resumes its original pre-computation frequency;

[0157] A downsampling module 808, configured to, when the primary instance is not abnormal, the standby instance executes a downsampling strategy; the downsampling strategy includes: feeding back data at least once within a retrospective time window.

[0158] Each module in the above-mentioned implementation device of the Prometheus cluster can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0159] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for implementing a Prometheus cluster. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0160] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0161] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0162] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0163] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0164] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0165] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0166] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0167] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A method for implementing a Prometheus cluster, characterized in that, The method includes: Configuring a first label, a second label, and a third label for each Prometheus instance according to a collection configuration template and runtime environment variables; wherein, the first label and the second label are used to determine a collection target; the third label is used to identify the type of the Prometheus instance; the types of the Prometheus instances include standby instances and primary instances; Performing pre-computation sharding configuration and pre-computation hierarchical configuration on the Prometheus instances according to a pre-computation configuration template; Configuring a query component to complete the query and aggregation of Prometheus data.

2. The method according to claim 1, wherein The configuring a first label, a second label, and a third label for each Prometheus instance according to a collection configuration template and runtime environment variables includes: Deploying Prometheus instances on each target server and setting the IP address of each target server as the runtime environment variable of the Prometheus instance; Loading the collection configuration template in a hot-loading manner to generate a first label, a second label, and a third label corresponding to each Prometheus instance; Forming a hash ring according to the Prometheus instances specified by the second label; Generating a plurality of collection targets according to the collection configuration template, and each collection target includes address information; Calculating a hash value according to the address information included in the collection target, and clockwise finding the first label corresponding to the closest hash value in the hash ring; Only retaining the collection target when the first label corresponding to the found hash value is the same as the first label corresponding to the collection target.

3. The method according to claim 1, wherein The performing pre-computation sharding configuration on the Prometheus instances according to a pre-computation configuration template includes: Identifying the Prometheus instances that need to be sharded and the Prometheus instances that do not need to be sharded by configuring a fourth label, and identifying the pre-computation tasks for querying local data and the pre-computation tasks for querying remote data by adding new fields; For the Prometheus instances that need to be sharded, determining the pre-computation tasks to be executed by using a consistent hash allocation method; For the Prometheus instances that do not need to be sharded, executing each pre-computation task.

4. The method according to claim 1, characterized in that The performing pre-computation hierarchical configuration on the Prometheus instances according to a pre-computation configuration template includes: Planning a hierarchical label for each Prometheus instance; wherein, all pre-computation queries will only match the Prometheus instances of the upper level.

5. The method according to any one of claims 1 to 4, characterized in that The configuring a query component to complete the query and aggregation of Prometheus data includes: Obtaining a query filtering condition and determining the Prometheus instances that meet the level; Grouping according to the group name label of the Prometheus instances, and sorting the Prometheus instances in the same group in ascending order according to the priority within the group; Executing the query task in a manner of concurrent execution of multiple groups, and querying data in order within each group; When a Prometheus instance in the same group successfully responds, stop querying data from the remaining Prometheus instances in the same group; Aggregate the query data feedback from all groups to obtain the aggregated query data.

6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Specify the address of the standby instance by configuring the fifth label; The primary instance sends a heartbeat signal to the address specified by the fifth label at a preset period; If the standby instance does not receive the heartbeat signal within a preset time, it is determined that the primary instance has an abnormality; In the case where the primary instance has an abnormality, the standby instance resumes its original pre-computation frequency; In the case where the primary instance does not have an abnormality, the standby instance executes a downsampling strategy; the downsampling strategy includes: feeding back data at least once within the retrospective time window.

7. An implementation device for a Prometheus cluster, characterized in that The device includes: An acquisition configuration module, configured to configure the first label, the second label, and the third label for each Prometheus instance according to the acquisition configuration template and the runtime environment variables; wherein, the first label and the second label are used to determine the acquisition target; the third label is used to identify the type of the Prometheus instance; the type of the Prometheus instance includes a standby instance and a primary instance; A pre-computation configuration module, configured to perform pre-computation sharding configuration and pre-computation layering configuration on the Prometheus instance according to the pre-computation configuration template; A query module, configured to configure a query component to complete the query and aggregation of Prometheus data.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.