Private cloud monitoring system and method

By designing a private cloud monitoring system based on Prometheus and Grafana, the problem of dispersed monitoring and alarm information in the traditional cloud platform operation and maintenance monitoring system is solved, and global monitoring and efficient alarms of key components of the cloud platform are achieved, thereby improving system stability and utilization efficiency.

CN120238628APending Publication Date: 2025-07-01BEIJING INST OF REMOTE SENSING INFORMATION
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510367543.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The traditional cloud platform operation and maintenance monitoring system disperses monitoring and alarm information for various components of the cloud platform, which makes it impossible for users to grasp the status of the cloud platform from a global perspective, and there are problems such as many invalid alarm data and high alarm false alarm rate.

Method used

A private cloud monitoring system is designed, based on Prometheus and Grafana to realize the aggregation and display of global monitoring information of key components of the cloud platform, and improve the effectiveness of alarm through the combination of data acquisition module, data storage module, monitoring display module and alarm execution module.

Benefits of technology

The system concentrates monitoring and alarms of various components of the cloud platform, reduces invalid alarm data, reduces alarm false alarm rate, and improves the stability and utilization efficiency of the cloud platform system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238628A_ABST
    Figure CN120238628A_ABST
Patent Text Reader

Abstract

The invention discloses a private cloud monitoring system and method. The system comprises a data acquisition module, a data storage module, a monitoring display module and an alarm execution module. The data acquisition module is in data connection with the data storage module, the monitoring display module and the alarm execution module and is used for acquiring cloud platform data; the data storage module is in data connection with the data acquisition module and is used for storing cloud platform data; the monitoring display module is in data connection with the data acquisition module and is used for monitoring and displaying cloud platform data; and the alarm execution module is in data connection with the data acquisition module and is used for managing alarm information. According to the invention, monitoring and alarm of each component of the cloud platform are concentrated in the private cloud monitoring system, so that invalid alarm data of the cloud platform are reduced, and the false alarm rate is reduced. And the method can be popularized to the related fields of cloud platform management and application in practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud platform operation and maintenance monitoring, and particularly to a private cloud monitoring system and method. Background Art

[0002] With the development of cloud computing technology, it has become a trend to use a cloud platform as the basis for software development, operation, and management. The cloud platform provides infrastructure services such as computing, storage, and networking, as well as cloud component services such as databases, middleware, and big data. Software runs in the cloud platform and uses various services provided by the cloud platform, enabling users to avoid maintaining various hardware and software resources locally and having them unifiedly managed by the cloud platform. The cloud platform can also achieve capabilities such as on-demand resource expansion, elastic deployment, load balancing, data backup, and disaster recovery, greatly improving the utilization and management efficiency of hardware and software resources, as well as the reliability and scalability of software operation.

[0003] The cloud platform operation and maintenance monitoring system monitors and maintains the cloud platform, and can visually monitor the operating status and resource usage of each component of the cloud platform, and issue alarms for abnormal status to ensure the healthy operation of the cloud platform. There are many components in the cloud platform, and each component has its own management interface. The monitoring and alarming of each component in the traditional cloud platform operation and maintenance monitoring system are scattered in the respective management interfaces of the components, resulting in scattered monitoring and alarming information, making it impossible for users to grasp the current status of the cloud platform from a global perspective and discover problems in a timely manner. At the same time, there are problems such as a large amount of invalid alarm data and a high false alarm rate in the alarms of the traditional cloud platform, which cannot accurately reflect the true state of the cloud platform. To address the above problems, it is necessary to improve the cloud platform operation and maintenance monitoring system, provide global monitoring information and alarm information aggregation display for key components of the cloud platform, and improve the effectiveness of alarms. A private cloud is a cloud computing method for exclusive use by an enterprise or individual, providing the most effective control over data, security, and service quality. A private cloud is built for the exclusive use of a single customer, and the enterprise owns the infrastructure and can control the way of deploying applications on this infrastructure. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a private cloud monitoring system and method. Based on Prometheus and grafana, the system can provide global monitoring information for key components of the cloud platform, aggregate and display alarm information, and improve the effectiveness of alarms. The system overcomes the deficiencies of existing monitoring devices, can improve the stability and utilization efficiency of the cloud platform system, and has practical application significance in engineering practice.

[0005] To solve the above technical problems, in the first aspect of the embodiments of the present invention, a private cloud monitoring system is disclosed. The system includes a data collection module, a data storage module, a monitoring and display module, and an alarm execution module;

[0006] The data collection module is data-connected to the data storage module, the monitoring and display module, and the alarm execution module, and is used to collect cloud platform data;

[0007] The data storage module is data-connected to the data collection module and is used to store cloud platform data;

[0008] The monitoring and display module is data-connected to the data collection module and is used to monitor and display cloud platform data;

[0009] The alarm execution module is data-connected to the data collection module and is used to manage alarm information.

[0010] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the data collection module includes open public network access and enterprise local area network access;

[0011] The open public network access is used to obtain cloud platform data by using the cloud platform data interface and perform customized probe data collection for complex requirements in specific scenarios;

[0012] The enterprise local area network access is used to perform cross-network interaction monitoring data collection.

[0013] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the data storage module includes data sharding storage;

[0014] The data sharding storage includes dividing different collection tasks into different monitoring servers and summarizing the data through one monitoring server at the upper layer.

[0015] In the second aspect of the embodiments of the present invention, a private cloud monitoring method is disclosed. The method includes:

[0016] S1, using the data collection module to collect cloud platform data of the monitoring source;

[0017] S2, using the data storage module to store the cloud platform data;

[0018] S3, using the monitoring and display module to display the cloud platform data;

[0019] S4, using the alarm execution module to process the cloud platform data to obtain alarm information.

[0020] As an alternative implementation, in the second aspect of the embodiments of the present invention, the collection of cloud platform data of the monitoring source by using the data collection module includes:

[0021] S11. When public network access is enabled, combine the cloud platform data interface and the customized probe to collect the cloud platform data of the monitoring source;

[0022] S12. When accessing within the enterprise local area network, use the cross-network interaction monitoring method to collect the cloud platform data of the monitoring source.

[0023] As an alternative implementation, in the second aspect of the embodiments of the present invention, the collection of the cloud platform data of the monitoring source by combining the cloud platform data interface and the customized probe when public network access is enabled includes:

[0024] S111. Optimize the interface access frequency of the cloud platform data interface to obtain the optimized upper limit access volume;

[0025] The calculation method of the optimized upper limit access volume is:

[0026] Single user * per minute (60m) * 30 = 1800

[0027] Among them, the single user is the upper limit access volume of 30 times per second. The system currently needs to obtain more than 100 indicators, and the system automatically groups more than 100 indicators, with each group containing no more than 30 indicators;

[0028] S112. Based on the optimized upper limit access volume, use the method of timed polling to collect the cloud platform data of the monitoring source.

[0029] As an alternative implementation, in the second aspect of the embodiments of the present invention, the collection of the cloud platform data of the monitoring source by using the cross-network interaction monitoring method when accessing within the enterprise local area network includes:

[0030] S121. Deploy the cloud platform monitoring information system in two physically isolated networks respectively, and realize accessing the two cloud platform monitoring information systems with one address through mapping;

[0031] S122. Use the preset monitoring server to collect the cloud platform data of the two cloud platform monitoring information systems.

[0032] As an alternative implementation, in the second aspect of the embodiments of the present invention, the storage of the cloud platform data by using the data storage module includes:

[0033] S21. Use the data storage module to define the data storage structure;

[0034] S22. Based on the data storage structure, the cloud platform data is stored using a data sharding storage method.

[0035] As an alternative implementation, in the second aspect of the embodiments of the present invention, the use of the alarm execution module to process the cloud platform data to obtain alarm information includes:

[0036] S41. Process the cloud platform data to obtain monitoring data;

[0037] S42. Preset alarm rules; the alarm rules include alarm conditions, alarm levels, and alarm thresholds;

[0038] S43. Analyze the monitoring data according to the alarm rules. When it is determined that no alarm is triggered, cancel the existing alarm; when it is determined that an alarm is triggered, execute S44;

[0039] S44. Analyze the monitoring data to generate new alarm information or update existing alarm information, and save it in the database.

[0040] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0041] The monitoring and alarming of each component of the cloud platform in the present invention are concentrated in the cloud platform operation and maintenance monitoring device (private cloud monitoring system). The present invention adopts dual means of cloud platform data interface and customized probe data collection to customize the data storage structure; reduce the invalid alarm data of the cloud platform and reduce the alarm false alarm rate. It can be popularized in the fields related to cloud platform management and application in terms of practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 is a schematic structural diagram of a private cloud monitoring system disclosed in an embodiment of the present invention;

[0044] Figure 2 is a schematic flowchart of a private cloud monitoring method disclosed in an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of a data collection module disclosed in an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of data storage disclosed in an embodiment of the present invention;

[0047] Figure 5 It is a schematic diagram of the data structure disclosed in the embodiments of the present invention;

[0048] Figure 6 It is a schematic diagram of the sample data structure disclosed in the embodiments of the present invention;

[0049] Figure 7 It is a schematic diagram of the storage sharding mechanism disclosed in the embodiments of the present invention;

[0050] Figure 8 It is a schematic diagram of the monitoring and display disclosed in the embodiments of the present invention;

[0051] Figure 9 It is a schematic diagram of the alarm execution disclosed in the embodiments of the present invention;

[0052] Figure 10 It is a schematic diagram of the alarm analysis process disclosed in the embodiments of the present invention. Detailed implementation manners

[0053] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0054] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or equipment.

[0055] Referring to "embodiments" herein means that a specific feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0056] The present invention discloses a private cloud monitoring system and method. The system includes a data acquisition module, a data storage module, a monitoring and display module, and an alarm execution module; the data acquisition module is data-connected to the data storage module, the monitoring and display module, and the alarm execution module, and is used for acquiring cloud platform data; the data storage module is data-connected to the data acquisition module and is used for storing cloud platform data; the monitoring and display module is data-connected to the data acquisition module and is used for monitoring and displaying cloud platform data; the alarm execution module is data-connected to the data acquisition module and is used for managing alarm information. The present invention centralizes the monitoring and alarm of each component of the cloud platform in the private cloud monitoring system, reduces the invalid alarm data of the cloud platform, and reduces the alarm false alarm rate. In terms of practicality, it can be extended to the fields related to cloud platform management and application. The following will be described in detail respectively.

[0057] Embodiment 1

[0058] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a private cloud monitoring system disclosed in an embodiment of the present invention. Among them, Figure 1 the described private cloud monitoring system is applied to the technical field of cloud platform operation and maintenance monitoring, and the embodiments of the present invention do not make limitations. As Figure 1 shown, the private cloud monitoring system includes a data acquisition module, a data storage module, a monitoring and display module, and an alarm execution module;

[0059] The data acquisition module is data-connected to the data storage module, the monitoring and display module, and the alarm execution module, and is used for acquiring cloud platform data;

[0060] The data storage module is data-connected to the data acquisition module and is used for storing cloud platform data;

[0061] The monitoring and display module is data-connected to the data acquisition module and is used for monitoring and displaying cloud platform data;

[0062] The alarm execution module is data-connected to the data acquisition module and is used for managing alarm information.

[0063] Optionally, the data acquisition module includes open public network access and enterprise local area network access;

[0064] The open public network access is used for obtaining cloud platform data by using the cloud platform data interface and performing customized probe data acquisition for complex requirements in specific scenarios;

[0065] The enterprise local area network access is used for cross-network interaction monitoring data acquisition.

[0066] Optionally, the data storage module includes data sharding storage;

[0067] The data sharding storage includes dividing different collection tasks into different monitoring servers, and aggregating the data through one monitoring server at the upper layer.

[0068] Prometheus is open-source software written in Go language, mainly used for monitoring various data of servers. Its working principle is to require the target host to expose a specific port. Subsequently, the Prometheus Server accesses this port according to preset rules at specific times or time intervals, obtains the data, and stores the obtained data in the database for subsequent data analysis, visualization, and other operations. Grafana is a powerful open-source software platform with rich charts and high customizability, which can meet the diverse display needs of users. Through the built-in data source configuration template, Grafana can directly connect to Prometheus, conveniently utilize the data in different data sources, ignore the problem of data source storage, and display the data in various forms such as graphs and tables, thus meeting the various needs of users to the greatest extent.

[0069] It can be seen that the monitoring and alarming of each component of the cloud platform in the present invention are concentrated in the cloud platform operation and maintenance monitoring device (private cloud monitoring system). The present invention adopts dual means of cloud platform data interface and customized probe data collection, and customizes the data storage structure; reduces the invalid alarm data of the cloud platform and reduces the alarm false alarm rate. It can be popularized in the fields related to cloud platform management and application in terms of practicality.

[0070] Embodiment 2

[0071] Please refer to Figure 2 , Figure 2 which is a schematic flow diagram of a private cloud monitoring method disclosed in an embodiment of the present invention. Among them, Figure 2 The described private cloud monitoring method is applied to the technical field of cloud platform operation and maintenance monitoring, and the embodiments of the present invention do not make limitations. As Figure 2 shown, the private cloud monitoring method includes:

[0072] S1. Using the data collection module to collect cloud platform data of the monitoring source;

[0073] S2. Using the data storage module to store the cloud platform data;

[0074] S3. Using the monitoring display module to display the cloud platform data;

[0075] S4. Using the alarm execution module to process the cloud platform data to obtain alarm information.

[0076] Optionally, the using the data collection module to collect cloud platform data of the monitoring source includes:

[0077] S11. When accessing the public network, combine the cloud platform data interface and the customized probe to collect the cloud platform data of the monitoring source;

[0078] S12. When accessing the enterprise local area network, use the cross-network interaction monitoring method to collect the cloud platform data of the monitoring source.

[0079] Optionally, when accessing the public network, combining the cloud platform data interface and the customized probe to collect the cloud platform data of the monitoring source includes:

[0080] S111. Optimize the interface access frequency of the cloud platform data interface to obtain the optimized upper limit access volume;

[0081] The calculation method of the optimized upper limit access volume is:

[0082] Single user * per minute (60m) * 30 = 1800

[0083] Among them, the single user is the upper limit access volume of 30 times per second. The system currently needs to obtain more than 100 indicators, and the system automatically groups more than 100 indicators, with each group containing no more than 30 indicators;

[0084] S112. Based on the optimized upper limit access volume, use the timed polling method to collect the cloud platform data of the monitoring source.

[0085] Optionally, when accessing the enterprise local area network, using the cross-network interaction monitoring method to collect the cloud platform data of the monitoring source includes:

[0086] S121. Deploy the cloud platform monitoring information system in two physically isolated networks respectively, and realize that one address accesses the two cloud platform monitoring information systems through mapping;

[0087] S122. Use the preset monitoring server to collect the cloud platform data of the two cloud platform monitoring information systems.

[0088] Optionally, using the data storage module to store the cloud platform data includes:

[0089] S21. Use the data storage module to define the data storage structure;

[0090] S22. Based on the data storage structure, use the data sharding storage method to store the cloud platform data.

[0091] Optionally, using the alarm execution module to process the cloud platform data to obtain the alarm information includes:

[0092] S41. Process the cloud platform data to obtain monitoring data;

[0093] S42. Preset alarm rules; the alarm rules include alarm conditions, alarm levels, and alarm thresholds;

[0094] S43. Analyze the monitoring data according to the alarm rules. When it is judged that no alarm is triggered, cancel the existing alarm; when it is judged that an alarm is triggered, execute S44;

[0095] S44. Analyze the monitoring data, generate new alarm information or update existing alarm information, and save it in the database.

[0096] Among them, step S43 specifically includes:

[0097] Extract features from the monitoring data to obtain feature parameter information;

[0098] The method of feature extraction is:

[0099] S431. Process any sample point x in the monitoring data X i , where i = 1, 2,..., N. Select k neighbors around the sample point x i . There are p sample points in each neighbor. Calculate the distance between p sample points x q in each neighbor and any sample point x i where q = 1, 2,..., p, j = 1, 2,..., k, and N is the number of sample points in X;

[0100] The distance The calculation formula is:

[0101]

[0102] where D is the dimension of the sample point x i , x it is the t-th dimension of the sample point x i , x qtj is the t-th dimension of the q-th sample x qj in the j-th neighbor, The larger i , the higher the similarity between the sample point x q and x

[0103] S432. Sum up the distances to obtain the distance sum of any sample point x i

[0104] The distance sum The calculation formula is:

[0105]

[0106] S4333, process the said distance and to obtain the weight coefficient W of the neighbor sequence ij ;

[0107] The expression of the weight coefficient W of the neighbor sequence ij is:

[0108]

[0109] In the formula, W ij is the weight coefficient of the neighbor sequence between the sample point x i and its j-th neighbor, j = 1, 2,..., k, where k is the number of neighbors; is the distance between the sample point x i and its j-th neighbor;

[0110] S434, process the weight coefficient W of the neighbor sequence ij to obtain the weighted monitoring data feature information Y;

[0111] The expression of the i-th sample point y of the weighted monitoring data feature information Y i is:

[0112]

[0113] In the formula, α is a real number between 0 and 1, set by experiment, is the distance between the sample point x1 and the j-th neighbor, the distance between the sample point x N and the j-th neighbor;

[0114] S435, perform visualization processing on the weighted monitoring data feature information Y to obtain feature parameter information.

[0115] Use the feature parameter information to train a preset alarm model to obtain an optimized alarm model;

[0116] Use the optimized alarm model to classify the feature parameter information to be processed to obtain alarm information;

[0117] The alarm information includes triggering an alarm and not triggering an alarm.

[0118] Among them, the preset alarm model is a convolutional long short-term memory network (ConvLSTM) combined with an LSTM network. ConvLSTM introduces a convolutional operation on the basis of LSTM, improving the processing ability of time series data.

[0119] Model structure

[0120] Convolutional Long Short-Term Memory Network (ConvLSTM): By introducing convolutional operations into the LSTM unit, ConvLSTM can capture spatio-temporal features more effectively. Its calculation formula is as follows:

[0121] Forget gate:

[0122]

[0123] f t : Output of the forget gate, controlling the degree of information forgetting; σ: Sigmoid activation function, compressing the output value between 0 and 1; W xf : Input data; x t : Input data at the current moment is feature parameter information; W hf : Hidden state at the previous moment; h t-1 Convolution kernel weights related to the forget gate. W cf : Cell state; c t-1 : Cell state at the previous moment, storing the accumulation of past information. b f : Bias term of the forget gate, helping to adjust the output of the forget gate.

[0124] Input gate:

[0125]

[0126] i t : Output of the input gate, controlling the introduction of new information; W xi : Input data x t Convolution kernel weights related to the input gate; W hi : Hidden state at the previous moment; h t-1 Convolution kernel weights related to the input gate; W ci : Cell state c t-1 Weights related to the input gate; b i : Bias term of the input gate.

[0127] Output gate:

[0128]

[0129] o t : Output of the output gate, controlling the cell state; c t : How much information is passed to the hidden state; W xo : Input data; x t Convolution kernel weights related to the output gate; W ho : Hidden state at the previous moment; h t-1 Convolution kernel weights related to the output gate; W co : Cell state ct Weights related to the output gate; b o : Bias term of the output gate.

[0130] Cell state update:

[0131]

[0132] c t : Cell state at the current time step; tanh: Hyperbolic tangent function that compresses values between -1 and 1. W xc : Input data x t Convolution kernel weights related to cell state update. W hc : Hidden state at the previous time step; h t-1 Convolution kernel weights related to cell state update; b c : Bias term for cell state update.

[0133] Hidden state update:

[0134]

[0135] h t : Hidden state at the current time step, including physiological parameter features at the current time step.

[0136] o t : Output of the output gate, which determines how much information of the cell state c t is passed to the hidden state.

[0137] Embodiment III

[0138] A private cloud monitoring system and method in this embodiment are as follows:

[0139] Step 1: Data collection

[0140] Data collection is as Figure 3 shown.

[0141] 1. Determine the data source

[0142] In response to the increasing complexity of cloud platform products and the critical collection requirements for specific important instance operation information, this system adopts a dual-track data collection strategy that combines interfaces and probes. This strategy aims to ensure the comprehensiveness, efficiency, and deep customization of data collection through dual means to cope with the increasingly complex and changeable monitoring environment.

[0143] (1) Cloud platform interface integration:

[0144] The system deeply integrates the comprehensive interface services provided by the cloud platform, which can seamlessly connect and capture detailed data of each product instance. The cloud platform interface can provide strong support to ensure comprehensive coverage and real-time update of data.

[0145] (2) Customized probe data collection:

[0146] In response to complex requirements in specific scenarios, we introduced probe technology as a supplementary means. Probes are like intelligent detectors that can penetrate deep into the system and capture data that is difficult to obtain directly through standard interfaces or requires special processing. Through customized development, probes can accurately match business needs and achieve deep mining and customized capture of data.

[0147] 2. Cross-network interactive monitoring data collection

[0148] When the cloud platform monitoring information system involves cross-network interaction, it can be achieved through the following three methods:

[0149] (1) To ensure the physical isolation of the two networks, dedicated equipment must be used for controlled one-way data transmission to enable real-time synchronous data interaction between systems in the networks.

[0150] (2) To achieve the timeliness of real-time monitoring, cloud platform monitoring information systems are deployed in two physically isolated networks. Through mapping, one address is used to access the two monitoring systems.

[0151] (3) Due to the limitations of the network architecture (classic type networks and vpc type networks require proxy processing, and the network is blocked by default), Prometheus cannot directly pull this data. In this system, Pushgateway is used as an intermediary service, allowing these targets that cannot be directly captured to send monitoring data to Pushgateway by Push, and then Prometheus pulls data from Pushgateway. This effectively solves the need to collect and display data due to network isolation.

[0152] 3. Data interface breaks through current limiting measures

[0153] In this patent, monitoring data is mainly obtained through three methods. The interface method to obtain monitoring data mainly relies on the DescribeMetricLast interface provided by the cloud platform. The upper limit of API interface calls provided by the existing cloud platform is 10,000 times / s, and a single user cannot call more than 30 times / s. If the interface data call reaches or exceeds the upper limit, it will trigger an alarm or cause the interface access to be blocked, resulting in some monitoring data not being displayed normally and failing to meet monitoring expectations.

[0154] Based on the existing objective environment and on the premise of meeting the interface call limit, how to efficiently obtain all the required monitoring metric data (currently more than 100) to solve the problem of data access interface traffic throttling. This system adopts a strategy that combines batch query and timed polling. The query requirements for these metrics are dispersed into different time periods, and at the same time, the batch processing ability of the interface (if supported) is utilized to reduce the number of single requests. Specific implementation measures:

[0155] (1) Optimization of interface access frequency

[0156] Based on the current upper limit of 30 accesses per second for a single user of the monitoring metrics, when it is calculated at the minute level, the upper limit of the monitoring metric acquisition volume is:

[0157] Single user * per minute (60m) * 30 = 1800

[0158] And the system currently needs to obtain more than 100 metrics. The system automatically groups more than 100 metrics, with each group containing no more than 30 metrics (or adjusted according to the actual batch size supported by the interface), and fully utilizes the call quota per minute of a single user through this scheduling strategy.

[0159] (2) Timed polling

[0160] Regarding the adjustment of access frequency, it can temporarily meet the current business requirements. The monitoring metrics may continuously increase as the business grows. For general expansion requirements, the system adopts a timed polling strategy. After the system is deployed, it will query the changes of the monitoring metrics in real time. When the upper limit of metric acquisition is exceeded, at a certain time task point, it will sequentially call the DescribeMetricLast interface to query the monitoring metrics of each group according to the grouping strategy.

[0161] In summary, based on the predictable monitoring metric range of the current cloud platform, through the above two methods, it is completely possible to meet the monitoring requirements of all metric items.

[0162] Step 2: Data storage

[0163] Data storage is as Figure 4 shown.

[0164] 1. Logical model (custom data storage structure)

[0165] In Prometheus, time series data mainly uses key-value pairs. The key is the value that stores the actual measurement value as a number when representing the measurement value (Prometheus does not store the original information, such as log text, etc. It stores the metrics summarized over time). The key represents the monitoring metric, such as the CPU usage percentage or the memory usage.

[0166] The data structure is as Figure 5As shown in the figure. When there are multiple cores in the CPU, if more detailed information about metrics is required, it is necessary to use tags to specify the rate of the core CPU located at a certain IP. In the patent solution, to prevent a large amount of monitoring data from being lost, unify the data structure in a multi-cloud environment, and quickly find and locate problem machines or services. This system is extended and transformed based on general data, and the tags are extended to 11 types of metric items. By filtering metrics with tags, the cloud, host, IP, and service to be searched and located can be accurately retrieved. The schematic diagram of its sample data structure is as Figure 6 shown.

[0167] 2. Data sharding storage

[0168] To effectively solve the problem of storing a large amount of data, different types of collection tasks are divided into different Prometheus instances for execution, and functional sharding is performed. For example, one Prometheus is responsible for collecting node metric data, and another Prometheus is responsible for collecting monitoring metric data related to application services. Finally, at the upper layer, a Prometheus is used to summarize the data. Compared with the data sharding under the federation mechanism, in this patent, mainly static initialization and dynamic data splitting by HASH modulo are adopted. Its storage sharding mechanism is as Figure 7 shown.

[0169] When initializing the system, one or more default data push addresses need to be set as the benchmark points for data distribution. This ensures that even without specific product configuration, the system can smoothly push monitoring data to the preset target locations. This strategy is the static data push strategy.

[0170] When the system collects monitoring data and is ready to push it, the HASH modulo algorithm is used to obtain the push address of a specific product. If the product does not have a dedicated push address configured, it will automatically fallback to use the default address set. Then, based on the scale of the monitoring data (estimated by the product of the number of instances and the number of metric items), combined with the number of the push address set, this process ensures that the data can be evenly distributed to different storage locations, which not only reduces the pressure on a single storage node but also improves the access efficiency and security of the data.

[0171] 3. Federation mechanism

[0172] As the business of the cloud platform continues to expand and become more complex, when a single Prometheus instance cannot handle a large number of collection tasks, at this time, we can use the Prometheus federation cluster-based method to divide the monitoring tasks into different Prometheus instances.

[0173] The metrics corresponding to different local Prometheus instances are aggregated to form a global instance set for the entire monitoring platform, achieving the purpose of unified monitoring.

[0174] Monitoring and display are as Figure 8 shown. Alarm execution is as Figure 9 shown. The collection of monitoring alarm information on the cloud platform mainly obtains it through the cloud platform interface. Cloud platform alarms mainly classify and assemble alarm data through interface calls.

[0175] (1) Sources of alarm information

[0176] To ensure the real-time and accurate monitoring data, the sources of alarm information for this patent mainly include three aspects:

[0177] First, the collection of monitoring alarm information on the cloud platform mainly obtains it through the cloud platform interface. Cloud platform alarms mainly classify and assemble alarm data through interface calls, synchronously store it locally in Prometheus, and display it in Grafana through data source configuration.

[0178] Second, the Grafana alarm threshold mainly forms alarm information after comparing the information captured from the cloud platform with the configured threshold. Grafana alarms are configured through the alarm panel of Grafana, mainly including alarm topics, alarm calculation frequencies, alarm duration triggers, triggering conditions and thresholds, and specifying alarm channels. Through the configured established rules or custom push methods, the push and reception of alarm information are realized.

[0179] Third, the Alertmanager alarm mechanism defines boolean value expressions with the help of PromQL according to the rule configuration to generate alarm information. Once an alarm occurs, Prometheus generates alarm information and sends it to Alertmanager. Alertmanager performs alarm notifications periodically according to the custom alarm routing and interfaces with third-party platforms, such as SMS, email, DingTalk, etc.

[0180] Based on the above three different sources of monitoring data and combined with the pre-set alarm strategies, it fully meets the needs of specific users to receive alarm information for a certain item or a certain type of monitoring metrics.

[0181] (2) Alarm rule settings

[0182] Alarm rules define how to generate alarms. Alarm rules are stored in the database and can be managed such as added, modified, deleted, etc. on the management page, specifically including the following content.

[0183] Define alarm conditions: Determine which situations need to trigger an alarm, such as high system load, database connection failure, long response time of a specific service, etc., and determine the corresponding monitoring metrics, etc.

[0184] Set alarm levels: Define the levels of alarms, which are divided from high to low into: critical, major, minor, and reminder levels, to distinguish the importance and attention level of alarms.

[0185] Set alarm thresholds: Set the corresponding alarm thresholds for the monitoring metrics and alarm levels that require alarms. If the collected metric value exceeds the threshold, an alarm is triggered.

[0186] (3) Alarm data analysis

[0187] Alarm analysis implements a warning mechanism for abnormal situations, which is divided into four parts: alarm rule setting, collected data analysis, alarm generation, and alarm management. Its operation process is as Figure 10 shown.

[0188] 1) Alarm rule setting

[0189] Alarm rules define how to generate alarms. Alarm rules are stored in the database and can be added, modified, deleted, and other management operations can be performed on the management page. The specific contents are as follows.

[0190] Define alarm conditions: Determine which situations need to trigger an alarm, such as high system load, database connection failure, long response time of a specific service, etc., and determine the corresponding monitoring metrics, etc.

[0191] Set alarm levels: Define the levels of alarms, which are divided from high to low into: critical, major, minor, and reminder levels, to distinguish the importance and attention level of alarms.

[0192] Set alarm thresholds: Set the corresponding alarm thresholds for the monitoring metrics and alarm levels that require alarms. If the collected metric value exceeds the threshold, an alarm is triggered.

[0193] 2) Collected data analysis

[0194] According to the alarm rules, compare and analyze the collected metric data. If the metric value exceeds the alarm threshold, an alarm is triggered; otherwise, no alarm is issued.

[0195] The original collected data is stateless and the data volume is huge, which has no direct connection with the alarm rules. Both the alarm rules and the alarm levels are user-defined. Therefore, the comparison and analysis of the metric data and the alarm rules need to be retrieved from a large amount of data and also meet the fast matching with the rules. Therefore, the design of relevant data structures is very important.

[0196] The trigger alarm process is the most critical part of alarms. First, it is necessary to cache the built-in alarm rules to reduce the response time for rule hits and relieve the load pressure on the database. Second, for operations such as adding and updating custom alarm rules, all alarm rules need to be synchronously updated in the cache to improve the hit rate of alarm rules.

[0197] After the collected data matches the alarm rules, it is necessary to determine the subsequent data processing flow based on the alarm information of the current rules. First, it is necessary to perform logical operation comparison between the collected data and the thresholds of multiple matching alarm rules:

[0198] a) When none of the multiple rule thresholds are met, no alarm is generated for the current collected data in the alarm system. If there is an alarm in the current alarm system with an unrecovered status for this rule, then the current alarm is restored, the alarm status is modified to recovered, and the alarm end time is updated to the current time. If there is no alarm in the current alarm system with an unrecovered status for this rule, then the current collected data is not processed in the alarm system.

[0199] b) When at least one rule threshold is met and there is no alarm with an unrecovered status in the current alarm, a new alarm message is generated, and the first reported time of the alarm is updated to the current time, and the metric value is the currently collected value. When there is an alarm with an unrecovered status in the current alarm, it will enter the alarm escalation and alarm storm handling processes.

[0200] c) When the alarm rule matched by the metric value is the same as the current alarm level, duplicate removal processing for the alarm storm is performed, the alarm count of the current alarm is increased, and the most recent reported time is updated;

[0201] d) When the alarm matched by the metric value is higher than the current alarm level, the current alarm is escalated, the current alarm is automatically restored, and a new high-level alarm is generated;

[0202] e) When the alarm matched by the metric value is lower than the current alarm level, the current alarm is downgraded, the current alarm is automatically restored, and a new low-level alarm is generated.

[0203] 3) Alarm Generation

[0204] When an alarm is triggered, new alarm information is generated or existing alarm information is updated and saved in the database. The alarm information includes alarm source, alarm metric, alarm metric value, alarm level, triggered alarm rule, alarm time, etc., which is convenient for management personnel to understand the detailed situation of the alarm generation and conduct troubleshooting and traceability.

[0205] The alarm information is timely notified to the operation and maintenance personnel in the system in the form of message notifications.

[0206] 4) Alarm Management

[0207] After an alarm is generated, to avoid the invalidation of alarm information and alarm information storm problems, it is necessary to manage the alarm information, mainly including two aspects: alarm recovery and alarm storm handling.

[0208] 5) Alarm recovery

[0209] Alarm recovery refers to checking the validity of the generated alarm. If the problem indicated by the alarm no longer exists currently, this alarm is cancelled, mainly achieved through recovery detection and automatic recovery.

[0210] Recovery detection: The system periodically or real-time detects whether the problem has been solved. If the problem has been solved, the alarm is cancelled.

[0211] Automatic recovery: For some problems, the system attempts to automatically recover, such as restarting the service, releasing resources, etc.

[0212] Alarm storm handling

[0213] An alarm storm refers to a situation where a large number of alarms are generated in a short period, which may lead to low processing efficiency. The handling measures mainly include deduplication, suppression, priority sorting, and alarm upgrade and downgrade, etc.

[0214] Deduplication: Merge the same or similar alarms to reduce the number of alarms.

[0215] Suppression: Under specific conditions, temporarily suppress the sending of some alarms until the conditions change.

[0216] Priority sorting: Sort the alarms by priority to ensure that important alarms are processed first.

[0217] Upgrade and downgrade: When the current alarm level is not equal to the level of the matching alarm rule, the current alarm level will be modified based on the matching rule.

[0218] (4) Alarm suppression strategy

[0219] In view of the information bombing formed by a large number of alarm information in the cloud platform in this patent and the situation that key alarm information cannot be distinguished in time, this system mainly realizes alarm suppression through three methods: aggregation grouping, silence, and suppression.

[0220] 1) Aggregation grouping

[0221] Aggregation grouping combines a category of information into one notification. For example, for situations where the CPU instantaneous usage reaches a peak and then drops back to normal, the memory instantaneous usage is too high and then drops to normal, or the system experiences a temporary high usage during large query operations or big data calculations, etc., such information can be appropriately aggregated, grouped, and packaged to form a notification push. This effectively avoids the input of a large amount of invalid information, enabling users to always pay attention to key information and thus quickly locate problems.

[0222] 2) Silence

[0223] For known problems such as hardware failures, disk drops, insufficient configuration resources, power on / off, and major version iteration and upgrade of the cloud platform, or special application scenarios during maintenance, the system provides a silence suppression strategy, that is, it allows administrators to set a certain alarm silence period and ignore alarms within a certain time period for applicable alarm suppression.

[0224] 3) Suppression

[0225] The suppression strategy mainly aims at receiving a large number of alarm messages simultaneously. High-level alarms suppress low-level alarms. For example, when a physical server crashes, a high-level alarm will be triggered. At the same time, low-level alarms such as the processes running on this physical machine remaining alive will also be triggered. In such cases, the system will trigger high-level alarms through pre-set configuration conditions to ensure that users receive the most important alarm information in the first place.

[0226] It can be seen that the present invention has dual means of data collection, customized probe data collection, user-defined data storage structure, cross-network interaction monitoring data collection, docking of alarm notifications from third-party platforms, data interface breaking through current limiting measures, data sharding storage, and alarm analysis process for the monitoring large screen of the cloud platform. The monitoring and alarming of each component of the cloud platform are concentrated in the cloud platform operation and maintenance monitoring device. It reduces the invalid alarm data of the cloud platform and decreases the alarm false alarm rate. In terms of practicality, it can be extended to the fields related to cloud platform management and application.

[0227] The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0228] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each implementation can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0229] Finally, it should be noted that: what is disclosed in an embodiment of a private cloud monitoring system and method of the present invention is only a preferred embodiment of the present invention, and is only used to illustrate the technical solution of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A private cloud monitoring system, characterized in that: The system includes a data acquisition module, a data storage module, a monitoring display module and an alarm execution module; The data acquisition module is data-connected with the data storage module, the monitoring display module, and the alarm execution module to collect cloud platform data; The data storage module is data-connected to the data acquisition module and is used to store cloud platform data; The monitoring and display module is data-connected to the data acquisition module and is used to monitor and display the cloud platform data; The alarm execution module is data connected to the data acquisition module for managing alarm information.

2. The private cloud monitoring system according to claim 1, characterized in that: The data acquisition module includes an open public network access unit and an enterprise local area network access unit; The open public network access unit is used to obtain cloud platform data using the cloud platform data interface, and to perform customized probe data collection according to complex requirements in specific scenarios; The enterprise LAN access unit is used for cross-network interactive monitoring data collection.

3. The private cloud monitoring system according to claim 1, characterized in that: The data storage module includes data shard storage; The data sharding storage includes initialized static data splitting and HASH modulus dynamic data splitting; The initialized static data is divided into one or more default data push addresses set when the system is initialized, which are used as reference points for data distribution to push the cloud platform data to the preset target location; The dynamic data of the HASH modulus is split into: when the system collects cloud platform data and prepares to push it, the HASH modulus algorithm is used to obtain the push address of the cloud platform data. If the cloud platform data is not configured with a dedicated push address, it automatically falls back to using the default push address; When a single Prometheus cannot handle a large number of data collection tasks, the federation mechanism is used to divide different data collection tasks into different Prometheus, and a Prometheus is used at the upper layer to aggregate different Prometheus data.

4. A private cloud monitoring method, characterized in that: Applied to the private cloud monitoring system of claims 1 to 3, the method comprises: S1, using the data acquisition module to collect cloud platform data of the monitoring source; S2, storing the cloud platform data using a data storage module; S3, using a monitoring display module to display the cloud platform data; S4, using the alarm execution module to process the cloud platform data to obtain alarm information.

5. The private cloud monitoring method according to claim 4, characterized in that: The method of collecting the cloud platform data of the monitoring source by using the data collection module includes: S11, when the public network access is open, the cloud platform data interface and customized probes are combined to collect the cloud platform data of the monitoring source; S12, when accessing the enterprise LAN, collects cloud platform data of the monitoring source using a cross-network interactive monitoring method.

6. The private cloud monitoring method according to claim 5, characterized in that: When the public network access is opened, the cloud platform data interface and customized probe are combined to collect the cloud platform data of the monitoring source, including: S111, optimizing the interface access frequency of the cloud platform data interface to obtain an optimized upper limit access volume; The calculation method of the optimized upper limit access volume is: Single user*per minute (60m)*30=1800 The upper limit of the number of visits per second for a single user is 30. The system currently needs to obtain more than 100 indicators. The system automatically groups the more than 100 indicators into groups, with each group containing no more than 30 indicators. S112, based on the optimized upper limit access volume, the cloud platform data of the monitoring source is collected by using a timed polling method.

7. The private cloud monitoring method according to claim 5, characterized in that: When accessing the enterprise LAN, the cloud platform data of the monitoring source is collected by using a cross-network interactive monitoring method, including: S121, using a one-way network gate data exchange method to ensure the physical isolation of the two networks; S122, deploying cloud platform monitoring information systems in two physically isolated networks respectively, and using mapping to implement access to the two cloud platform monitoring information systems using one address; S123, using Pushgateway as an intermediary service, and using Prometheus to collect cloud platform data of the two cloud platform monitoring information systems.

8. The private cloud monitoring method according to claim 4, characterized in that: The storing of the cloud platform data by using a data storage module includes: S21, using the data storage module to define a data storage structure; S22, based on the data storage structure, using data sharding storage method to store the cloud platform data.

9. The private cloud monitoring method according to claim 4, characterized in that: The cloud platform data is processed by the alarm execution module to obtain alarm information, including: S41, processing the cloud platform data to obtain monitoring data; S42, preset alarm rules; the alarm rules include alarm conditions, alarm levels and alarm thresholds; S43, analyzing the monitoring data according to the alarm rule, and when it is determined that the alarm is not triggered, canceling the existing alarm; when it is determined that the alarm is triggered, executing S44; S44, analyzing the monitoring data, generating new alarm information or updating existing alarm information, and storing the information in a database.

Citation Information

Patent Citations

  • Equipment abstract in self-organization wireless local area network

    CN102857907A

  • Data collection method and data collection device

    CN106998259A

  • Private cloud monitoring method and device based on non-flat network, computer equipment and storage medium

    CN111459750A

  • Private cloud monitoring system and method

    CN115047817A

  • Large-trend security cloud operation data acquisition system

    CN119030783A