Cloud platform monitoring data storage and analysis system based on Greoptimum DB
By designing a cloud platform monitoring data storage analysis system based on GreptimeDB, the performance bottleneck problem of existing systems in high concurrency, large data volume and strong real-time environments is solved, and efficient collection, storage, cleaning and analysis of monitoring data is achieved, improving analysis efficiency and system performance.
Patent Information
- Application Number
- CN202411940755.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
The existing cloud platform monitoring data storage and analysis systems show performance bottlenecks in environments with high concurrency, large data volume and strong real-time performance, including problems such as data loss and abnormality, data inconsistency and duplication, and traditional statistical methods are difficult to adapt to the analysis tasks under large volumes and multi-source heterogeneous monitoring data.
A cloud platform monitoring data storage and analysis system based on GreptimeDB is designed. The system includes a data acquisition module, a metadata management module, a data transmission and storage module, a data cleaning and standardization module, and a data access and analysis module. Through asynchronous transmission, hot and cold data hierarchical storage, stream processing and batch processing, efficient collection, storage, cleaning and analysis of monitoring data is achieved.
Through data cleaning and standardized processing, the system improves the analysis efficiency of monitoring data, realizes high performance and intelligent analysis of real-time monitoring data, and solves the performance bottlenecks of existing systems in environments of high concurrency, large data volume and strong real-time.
Smart Images

Figure CN119938442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to a cloud platform monitoring data storage and analysis system. Background Art
[0002] With the rapid development of cloud computing technology, cloud platforms, as the core of modern information technology infrastructure, undertake big data processing and computing tasks. In order to ensure the stable and efficient operation of cloud platforms, it is particularly important to monitor the use of various cloud platform resources in real time and perform big data storage and analysis.
[0003] In the prior art, existing systems are mostly based on distributed databases and basic time series databases. Since cloud platform monitoring data has the characteristics of high concurrency, large data volume and strong real-time performance, these databases often show performance bottlenecks when storing massive monitoring data, including data missing and abnormal, data inconsistency and duplication, etc. In addition, in terms of data analysis, the existing platform monitoring data analysis method is based on traditional statistical methods, which is difficult to adapt to the analysis tasks under large-scale and multi-source heterogeneous monitoring data. Summary of the invention
[0004] The purpose of the present invention is to provide a cloud platform monitoring data storage and analysis system based on GreptimeDB to solve the above technical problems;
[0005] The cloud platform monitoring data storage and analysis system based on GreptimeDB includes:
[0006] The data collection module is used to collect monitoring data from the business nodes of the heterogeneous cloud platform;
[0007] A metadata management module, connected to the data acquisition module, for extracting, storing, integrating, associating, maintaining and updating metadata of the monitoring data;
[0008] A data transmission and storage module is connected to the metadata management module, stores the monitoring data into a database in combination with the metadata, and performs cold and hot data hierarchical storage of the monitoring data according to the generation time and access frequency in the metadata;
[0009] A data cleaning and standardization module, connected to the data transmission and storage module, cleans and standardizes the monitoring data stored in the database to obtain standardized data;
[0010] The data access and analysis module is connected to the data cleaning and standardization module and is used for performing stream processing and batch processing on the standardized data.
[0011] Preferably, the data acquisition module includes:
[0012] A data collection unit interacts with the service node through a data collection tool to obtain the monitoring data;
[0013] The asynchronous transmission unit is connected to the data acquisition unit and transmits the monitoring data asynchronously through a message queue.
[0014] Preferably, the metadata management module includes:
[0015] A metadata extraction unit, used for extracting and standardizing the metadata in the monitoring data;
[0016] A metadata query and retrieval unit, connected to the metadata extraction unit, for providing a graphical interface to display the relationship between the metadata, and searching the monitoring data and the corresponding metadata by keywords;
[0017] The metadata maintenance unit is connected to the metadata extraction unit and is used to store and integrate the metadata and update the metadata in real time as the monitoring data changes.
[0018] Preferably, the data transmission and storage module includes:
[0019] A data transmission unit, combining the metadata, and transmitting the monitoring data via a hypertext transfer protocol;
[0020] a hot and cold data identification unit, connected to the data transmission unit, for identifying hot and cold data according to the access frequency in the metadata, identifying the data as hot data when the access frequency is greater than a set threshold, and identifying the data as cold data when the access frequency is less than or equal to the set threshold;
[0021] A hierarchical management unit, connected to the hot and cold data identification unit, performs cold and hot data hierarchical storage management on the monitoring data;
[0022] A data storage unit, connected to the hierarchical management unit, for storing the monitoring data in the database, locally caching the hot data, and object storing the cold data;
[0023] The data protection unit is connected to the data transmission unit, and verifies the monitoring data in the data transmission unit through a data transmission guarantee mechanism. The data transmission unit transmits the verified monitoring data to the data storage unit.
[0024] Preferably, the data transmission guarantee mechanism includes:
[0025] Before data transmission, the monitoring data is hashed by a hash function to map the monitoring data to a fixed first hash value. Before the data storage unit receives the data, the hash calculation is performed again to obtain a second hash value. The first hash value and the second hash value are compared. When the first hash value is the same as the second hash value, the monitoring data is stored in the database. When the first hash value is different from the second hash value, the data transmission unit retransmits the monitoring data.
[0026] Preferably, the data transmission and storage module further includes:
[0027] The data encryption unit is connected to the data transmission unit, performs a handshake protocol before the monitoring data is transmitted through a security protocol, generates a session key, and encrypts the monitoring data through the session key.
[0028] Preferably, the data cleaning and standardization module includes:
[0029] A singular value detection unit, which identifies and corrects singular values in the monitoring data based on a preset data range and statistical model;
[0030] A missing value filling unit, connected to the singular value detection unit, fills the missing parts in the monitoring data based on the knowledge graph;
[0031] A duplicate data removal unit, connected to the missing value filling unit, screens out the same monitoring data through data comparison and deletes them;
[0032] A standardization unit is connected to the duplicate data removal unit, and performs dimension conversion on the monitoring data to obtain the standardized data.
[0033] Preferably, the data access and analysis module includes:
[0034] A stream processing unit, which performs time-based or data volume-based aggregation analysis on the real-time data in the standardized data through a distributed stream processing framework;
[0035] The batch processing unit decomposes the historical data in the standardized data into multiple data sets through the distributed system basic framework, processes them on different nodes respectively, and summarizes and reprocesses the processing results.
[0036] Preferably, the data access and analysis module further includes:
[0037] The anomaly monitoring unit calculates and visualizes the standardized data through a time series algorithm suite and a computing engine to generate a resource usage trend report.
[0038] Preferably, it also includes:
[0039] An alarm and report module, the alarm and report module is connected to the data access and analysis module, including:
[0040] An alarm judgment unit sets an indicator threshold according to business requirements, compares the standardized data with the indicator threshold, and generates an alarm message and sends it to an alarm contact person when the standardized data exceeds the set indicator threshold;
[0041] An alarm grading unit, connected to the alarm judgment unit, configured different alarm levels for different indicator thresholds from multiple dimensions;
[0042] The report generation unit is connected to the alarm judgment unit and the alarm classification unit, and performs comprehensive analysis on the standardized data to obtain a periodic monitoring report.
[0043] The beneficial effects of the present invention are: by adopting the above technical scheme, the collected monitoring data is cleaned to obtain standardized data, the analysis efficiency of the monitoring data is improved, and through stream processing and batch processing, high performance and intelligent analysis of real-time monitoring data is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a structural diagram of the cloud platform monitoring data storage and analysis system based on GreptimeDB of the present invention;
[0045] Figure 2 is a schematic diagram of a data acquisition module of the present invention;
[0046] Figure 3 is a schematic diagram of a metadata management module of the present invention;
[0047] Figure 4 is a schematic diagram of a data transmission and storage module of the present invention;
[0048] Figure 5 is a schematic diagram of a data cleaning and standardization module of the present invention;
[0049] Figure 6 is a schematic diagram of a data access and analysis module of the present invention;
[0050] Figure 7 is a schematic diagram of the alarm and report module of the present invention;
[0051] Figure 8 It is the technical roadmap of the cloud platform monitoring data storage and analysis system based on GreptimeDB of the present invention.
[0052] In the attached figure: 1. Data acquisition module; 11. Data acquisition unit; 12. Asynchronous transmission unit; 2. Metadata management module; 21. Metadata extraction unit; 22. Metadata query and retrieval unit; 23. Metadata maintenance unit; 3. Data transmission and storage module; 31. Data transmission unit; 32. Cold and hot data identification unit; 33. Hierarchical management unit; 34. Data storage unit; 35. Data protection unit; 36. Data encryption unit; 4. Data cleaning and standardization module; 41. Singular value detection unit; 42. Missing value filling unit; 43. Duplicate data removal unit; 44. Standardization unit; 5. Data access and analysis module; 51. Stream processing unit; 52. Batch processing unit; 53. Abnormal monitoring unit; 54. Timing algorithm suite; 55. Distributed computing framework; 56. Storage and computing separation suite; 57. Computing engine; 6. Alarm and report module; 61. Alarm judgment unit; 62. Alarm classification unit; 63. Report generation unit; 7. Heterogeneous cloud platform monitoring data source; 8. Data service layer; 9. Data management layer; 10. Data storage layer; 101. Database node; 102. Virtual machine node; 103. File node; 104. Monitoring management unit. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0055] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0056] Cloud platform monitoring data storage and analysis system based on GreptimeDB, such as Figure 1 , Figure 8 As shown, including
[0057] Data collection module 1, used to collect monitoring data from business nodes of heterogeneous cloud platforms;
[0058] The metadata management module 2 is connected to the data collection module 1 and is used to extract, store, integrate, associate, maintain and update the metadata of the monitoring data;
[0059] The data transmission and storage module 3 is connected to the metadata management module 2, stores the monitoring data into the database in combination with the metadata, and performs cold and hot data hierarchical storage of the monitoring data according to the generation time and access frequency in the metadata;
[0060] The data cleaning and standardization module 4 is connected to the data transmission and storage module 3 to clean and standardize the monitoring data stored in the database to obtain standardized data;
[0061] The data access and analysis module 5 is connected to the data cleaning and standardization module 4 and is used for stream processing and batch processing of the standardized data.
[0062] Specifically, the present invention provides a cloud platform monitoring data storage and analysis system based on GreptimeDB. The database in the data transmission and storage module 3 refers to GreptimeDB. GreptimeDB is a database designed for time series data storage and analysis. It has efficient data writing, query and storage capabilities, and can adapt well to the characteristics of cloud platform monitoring data, so as to achieve the purpose of fully utilizing monitoring information and ensuring the smooth and safe operation of the cloud platform; the data cleaning and standardization module 4 is used to clean the collected monitoring data, obtain standardized data, and improve the analysis efficiency of the monitoring data. The data access and analysis module 5 realizes high performance and intelligent analysis of real-time monitoring data through stream processing and batch processing.
[0063] In a preferred embodiment, referring to Figure 2 , the data acquisition module 1 includes,
[0064] The data collection unit 11 interacts with the service node through a data collection tool to obtain monitoring data;
[0065] The asynchronous transmission unit 12 is connected to the data collection unit 11 and transmits the monitoring data asynchronously through the message queue.
[0066] Specifically, the data collection tool is Telegraf, which is used to collect multi-source data and use the message queue RabbitMQ (Rabbit Message Queue) for high-concurrency, asynchronous transmission of large-scale data to ensure the stability and reliability of data transmission; the message queue is a way of communication between applications, mainly used to decouple the dependencies between applications and improve the scalability and flexibility of the system.
[0067] For example, the sender delivers the letter (message) to the post office (message queue), and the receiver (consumer) obtains the letter from the post office. The sender and receiver do not need to interact directly, so both parties can work independently without being affected by each other.
[0068] It provides a unified and transparent access interface for heterogeneous multi-source cloud platform data. The interface collects monitoring data from each node of the cloud platform, including but not limited to heterogeneous multi-source indicators such as CPU utilization, memory utilization, disk utilization, and network bandwidth occupancy.
[0069] The heterogeneous cloud platform monitoring data source 7 includes multiple business nodes (business node A, business node B, business node C...business node N), which represent cloud platform monitoring data from different business systems. These business systems run on the cloud platform, and the data generated needs to be uniformly stored, analyzed and managed to ensure the normal operation of each business system and the overall performance of the cloud platform.
[0070] Use the Telegraf tool, which can interact with each node of the cloud platform to obtain various monitoring data. Telegraf has a variety of plug-ins that can adapt to different types of data collection needs.
[0071] The message queue RabbitMQ is used to handle high concurrency and asynchronous transmission of large-scale data. When the amount of collected data is large, the data is put into the queue through RabbitMQ and transmitted in an asynchronous manner to avoid blockage during the data collection process and ensure the stability and reliability of data transmission.
[0072] Provide a unified and transparent access interface, which can shield the differences in hardware, software, etc. among different nodes of the cloud platform, making the collection process standardized for other modules.
[0073] In a preferred embodiment, referring to Figure 3 , the metadata management module 2 includes,
[0074] A metadata extraction unit 21, used to extract metadata from monitoring data;
[0075] The metadata query and retrieval unit 22 is connected to the metadata extraction unit 21 and is used to provide a graphical interface to display the relationship between metadata and search for monitoring data and corresponding metadata by keywords;
[0076] The metadata maintenance unit 23 is connected to the metadata extraction unit 21 and is used to store and integrate metadata and update metadata in real time as the monitoring data changes.
[0077] Specifically, the metadata management module 2 extracts and stores the collected metadata, which includes information such as data source, collection time, data structure, and data quality requirements.
[0078] By defining the metadata model, information such as monitoring data structure, field type, and data relationship is fully collected and managed, providing basic support for monitoring data analysis;
[0079] The metadata management module 2 integrates with the data analysis module to provide information such as field descriptions and relationship dependency descriptions for monitoring data analysis; integrates with the monitoring and alarm module to monitor data metadata changes in real time and warn of possible abnormal monitoring data; integrates with the data visualization module to provide metadata descriptions of dimensions and metrics for visualization reports. In the data service layer 8, data source services refer to services that provide access to various data sources to ensure that data can flow smoothly from the data source into the system for processing.
[0080] The data synchronization service is responsible for synchronizing data between different data storage nodes or systems to ensure data consistency and integrity.
[0081] Synchronization management refers to the management and monitoring of the data synchronization process, including synchronization strategy formulation, synchronization progress tracking, synchronization error handling, etc.
[0082] Report analysis is similar to report analysis in data management. It generates relevant reports based on the data in the data service process to assist management and decision-making.
[0083] In a preferred embodiment, referring to Figure 4 , the data transmission and storage module 3 includes,
[0084] The data transmission unit 31 transmits the monitoring data through the hypertext transfer protocol in combination with the metadata;
[0085] The hot and cold data identification unit 32 is connected to the data transmission unit 31 and is used to identify the hot and cold data according to the access frequency in the metadata. When the access frequency is greater than a set threshold, the data is identified as hot data, and when the access frequency is less than or equal to the set threshold, the data is identified as cold data.
[0086] The hierarchical management unit 33 is connected to the hot and cold data identification unit 32, and performs hot and cold data hierarchical storage management on the monitoring data; the data storage unit 34 is connected to the hierarchical management unit 33, and is used to store the monitoring data in the database, and use a high-performance storage medium such as local cache for hot data, and a low-cost storage medium such as object storage for cold data, and provide a unified data access interface;
[0087] The data protection unit 35 is connected to the data transmission unit 31 , and verifies the monitoring data in the data transmission unit 31 through a data transmission guarantee mechanism. The data transmission unit 31 transmits the verified monitoring data to the data storage unit 34 .
[0088] Specifically, the data is sent to GreptimeDB through a high-speed data transmission channel using HTTP (Hypertext Transfer Protocol). HTTP is a widely used network transmission protocol that can perform effective data transmission in different network environments.
[0089] A high-reliability data transmission guarantee mechanism is set up on the transmission channel, such as a data verification and retransmission mechanism, to ensure that data will not be lost or damaged during transmission.
[0090] When transmitting data, a confirmation mechanism is usually used. After the sender sends a data packet, it will wait for the receiver's confirmation message. If the confirmation message is not received within a certain period of time (timeout period), the sender will think that the data packet is lost or damaged, and then trigger the retransmission mechanism.
[0091] When the receiver finds data errors through the data verification mechanism, it will send a negative acknowledgment (NAK) message to the sender to inform the sender that the data is incorrect. After receiving the NAK, the sender will resend the data.
[0092] By analyzing the data access records, the access frequency of each data is counted. If a data is frequently accessed, it is marked as hot data; otherwise, it is cold data.
[0093] Monitor business operations based on the classification of hot and cold data.
[0094] For example, changes in hot data may reflect key operations of the business, while long-term inaccessibility of cold data may indicate idleness or potential problems with certain business functions.
[0095] Implement a hierarchical management strategy for hot and cold data, such as storing hot data in higher-performance storage media for quick and easy access; and using more economical storage methods for cold data.
[0096] In a preferred embodiment, the data transmission guarantee mechanism includes:
[0097] Before data transmission, the monitoring data is hashed by a hash function to map the monitoring data to a fixed first hash value. Before the data storage unit 34 receives the data, the hash calculation is performed again to obtain a second hash value. The first hash value and the second hash value are compared. When the first hash value is the same as the second hash value, the monitoring data is stored in the database. When the first hash value is different from the second hash value, the data transmission unit 31 retransmits the monitoring data.
[0098] Specifically, assuming that the data block to be transmitted is "1234567890", the sender uses SHA-256 to calculate its hash value:
[0099] "abcdef1234567890abcdef1234567890abcdef1234567890abcdef1234567890".
[0100] The sender sends the data "1234567890" and the hash value:
[0101] "abcdef1234567890abcdef1234567890abcdef1234567890abcdef1234567890abcdef1234567890" are sent to the receiver together.
[0102] After receiving the data, the receiver uses SHA-256 to calculate the hash value of the received data "1234567890" again.
[0103] If the calculated hash value is the same as the received hash value:
[0104] If “abcdef1234567890abcdef1234567890abcdef1234567890abcdef1234567890” are the same, the data is considered complete; otherwise, the data transmission is considered incorrect.
[0105] In a preferred embodiment, the data transmission and storage module 3 further includes:
[0106] The data encryption unit 36 is connected to the data transmission unit 31, performs a handshake protocol before the monitoring data is transmitted through a security protocol, generates a session key, and encrypts the monitoring data through the session key.
[0107] Specifically, the security protocol is SSL / TLS, which uses the data encryption algorithm SSL / TLS to encrypt the transmitted data to ensure the security of the data and prevent the data from being stolen or tampered with during transmission.
[0108] The handshake protocol is a handshake protocol that the client and server will perform before SSL / TLS data transmission. During the handshake process, the two parties will negotiate encryption algorithms, exchange keys and other information. For example, the two parties negotiate to use AES (Advanced Encryption Standard) as the encryption algorithm.
[0109] Authentication refers to the process of handshake to ensure that the identities of both parties are authentic. The server sends its digital certificate to the client, and the client verifies the digital certificate to confirm the server's identity.
[0110] In a preferred embodiment, referring to Figure 5 ,Data cleaning and standardization module 4 includes,
[0111] A singular value detection unit 41, which identifies and corrects singular values in the monitoring data based on a preset data range and statistical model;
[0112] A missing value filling unit 42 is connected to the singular value detection unit 41 and fills the missing parts in the monitoring data based on the knowledge graph;
[0113] The duplicate data removal unit 43 is connected to the missing value filling unit 42 to screen out the same monitoring data through data comparison and delete them;
[0114] The standardization unit 44 is connected to the duplicate data removal unit 43 to perform dimension conversion on the monitoring data to obtain standardized data.
[0115] Specifically, the data cleaning and standardization module 4 is used to unify the data source into a standard format, including automatic detection of singular values, filling in missing values based on the knowledge graph, removing duplicate data, unifying dimensions, etc.
[0116] The singular value detection unit 41 identifies data points that obviously deviate from the normal range by setting a reasonable data range and statistical model, and processes them, such as deleting or correcting them.
[0117] For example, when monitoring the CPU utilization data of multiple servers in the cloud platform, under normal circumstances, the CPU utilization of these servers during stable operation is between 10% and 80% (this is the set reasonable data range). By statistically analyzing the CPU utilization data collected over a period of time, a simple statistical model is constructed, such as calculating the mean and standard deviation. If it is found that the CPU utilization of a server at a certain moment reaches 95%, which is far beyond the normal range (deviating from the mean by multiple standard deviations), then this data point is determined to be a singular value.
[0118] At this point, the data cleaning and standardization module 4 can adopt two processing methods: if this is an obviously erroneous data, such as an abnormally high value caused by a failure of the acquisition equipment, then the data point is directly deleted; if it is not sure whether it is completely wrong, it will be corrected according to the data trend of the adjacent time period, such as correcting it to a value that has a smoother transition with the data of the previous and next time periods, such as 80%.
[0119] Fill missing values based on knowledge graphs and metadata, and use existing data relationships and rules in knowledge graphs to reasonably fill in missing parts in the data to make the data more complete.
[0120] In cloud platform monitoring, the server's CPU utilization, memory utilization, disk I / O and other data are collected simultaneously, and a knowledge graph about the relationship between these data has been constructed.
[0121] For example, the knowledge graph shows that when the CPU utilization of the server is continuously in a high load state (such as greater than 80%) for a period of time, the memory utilization will usually increase accordingly. Suppose at a certain moment, the collected memory utilization data is missing, but at this time the CPU utilization has been above 90% for 10 consecutive minutes. According to the rules in the knowledge graph, the data cleaning and standardization module 4 can refer to the change pattern of memory utilization in similar situations in the past and reasonably fill in the missing memory utilization value.
[0122] For example, in similar situations in the past, the average memory utilization rate would increase to about 70%, so the currently missing memory utilization value would be filled with 70% to make the data more complete for subsequent comprehensive analysis.
[0123] Remove duplicate data and use data comparison algorithms to find and delete identical data records to avoid data redundancy.
[0124] During the data collection process, some duplicate data records were collected due to failure of the collection equipment or network fluctuations.
[0125] For example, at a certain moment, the network bandwidth usage data of the cloud platform is collected. Due to network problems, three identical network bandwidth usage values (such as 50Mbps) are collected in the same second. The data cleaning and standardization module 4 can identify these three duplicate data records through the data comparison algorithm. Then, the module will delete two of them and keep only one, so as to avoid these duplicate data from interfering with subsequent data analysis, reduce data redundancy, and improve the efficiency and accuracy of data processing.
[0126] Unify dimensions and convert dimensions of data from different sources. For example, convert values in different units into standard units to facilitate subsequent data processing and analysis.
[0127] Cloud platform monitoring data comes from different data sources, which may use different units for the same type of data.
[0128] For example, the disk space usage of some servers is collected and recorded in GB (gigabytes), while that of other servers is in MB (megabytes).
[0129] In the data cleaning and standardization module 4, all disk space usage data will be uniformly converted to the same dimension, for example, GB. For data in MB, it is converted to GB through simple mathematical conversion (divided by 1024), so that the disk space usage of all servers can be easily compared and analyzed, ensuring the accuracy and consistency of subsequent calculations, statistics and analysis results based on these data.
[0130] The data management layer 9 includes:
[0131] Data cleaning: pre-processing the collected cloud platform monitoring data to remove noise, outliers, etc. to ensure the quality and accuracy of the data.
[0132] Backup management is responsible for the formulation and implementation of data backup strategies to ensure that data can be quickly restored when it is lost or damaged, thereby ensuring data security.
[0133] Report analysis: Generate various analysis reports based on cloud platform monitoring data, such as performance reports, fault reports, etc., to provide decision-making basis for operation and maintenance personnel and management personnel.
[0134] Permission management controls access rights to cloud platform monitoring data, ensuring that only authorized users or systems can access and operate relevant data, thereby ensuring data confidentiality.
[0135] Instance management: manage various instances in the cloud platform (such as virtual machine instances, application instances, etc.), including operations such as creating, starting, stopping, and deleting instances.
[0136] Fault management monitors faults that occur during the operation of the cloud platform, and promptly locates, diagnoses, and repairs faults to ensure the stable operation of the cloud platform.
[0137] Security management is responsible for the data security and system security of the cloud platform, including firewall settings, intrusion detection, encryption technology, etc., to prevent data leakage and malicious attacks.
[0138] In a preferred embodiment, referring to Figure 6 , the data access and analysis module 5 includes,
[0139] A stream processing unit 51 performs aggregation analysis based on time or data volume on real-time data in the standardized data through a distributed stream processing framework;
[0140] The batch processing unit 52 decomposes the historical data in the standardized data into multiple data sets through the distributed system basic framework, processes them on different nodes respectively, and summarizes and reprocesses the processing results;
[0141] The data access and analysis module 5 also includes:
[0142] The anomaly monitoring unit 53 calculates and visualizes the standardized data through the time series algorithm suite 54 and the computing engine 57 to generate a resource usage trend report.
[0143] Specifically, the stored data is processed in real time and in batches to meet different business needs, including real-time data analysis, long-term trend analysis of batch data, complex calculations, trend prediction and anomaly detection, etc.
[0144] For streaming data, use Apache Flink (a distributed stream processing framework with low latency, high throughput, and strong fault tolerance. It can efficiently process unbounded data streams, that is, data sources are continuously generated without a clear end boundary). Flink has the characteristics of low latency and high throughput, and can quickly process and analyze real-time incoming data to meet the needs of timely response to cloud platform monitoring data.
[0145] Based on the time-based aggregation analysis, the stream processing unit 51 uses Apache Flink as the distributed stream processing framework. In order to monitor whether the video freeze situation deteriorates in a short period of time, a sliding window is set to 5 minutes. Within these 5 minutes, the video freeze rate in the streaming data is aggregated and calculated, such as the average freeze rate, the maximum and minimum freeze rate, etc. Every minute, the window slides forward one minute to continuously update the freeze rate statistics within these 5 minutes. If it is found that the average freeze rate has risen from 2% to 5% in the past 5 minutes, and the maximum value has reached 10%, an alarm can be triggered in time to notify the technician to check the video source or network status to ensure the user's viewing experience.
[0146] Based on the aggregate analysis of data volume, for the indicator of the number of people watching online at the same time, an aggregate analysis is performed every time 1,000 data are reached (i.e., every time 1,000 new viewer connection records are added). The regional distribution of these 1,000 connection records is calculated, and the audience share in different regions is counted. In this way, the operators of the live broadcast platform can understand the changes in the geographical sources of the audience in real time, so as to adjust the live broadcast content or optimize the server deployment according to the preferences of the audience in different regions.
[0147] For batch processing of large amounts of historical data, the big data processing framework Hadoop (mainly composed of a distributed file system and a computing model) is used. Hadoop's distributed file system (HDFS, Hadoop Distributed File System) and parallel computing model (MapReduce) can effectively process large-scale data sets and complete long-term trend analysis and complex computing tasks.
[0148] The batch processing unit 52 uses Hadoop's MapReduce to decompose the standardized historical data of this month according to date, and the data of each day is regarded as a data set and distributed to different computing nodes for processing.
[0149] For example, computing node 1 processes data from number 1 to 10, computing node 2 processes data from number 11 to 20, and so on. Each node performs preliminary processing on the data it is responsible for, such as calculating the average number of simultaneous online viewers per day and the total video playback time.
[0150] After each computing node completes the initial processing, the results are aggregated together. Then further complex calculations are performed, such as comparing the average number of online viewers in different time periods (such as weekdays and weekends, daytime and nighttime) in the past month to analyze the periodic changes in user activity; or through correlation analysis, find out the potential relationship between the video freeze rate and factors such as the number of simultaneous online viewers and network bandwidth usage, and provide data support for the platform's resource optimization configuration, such as increasing server bandwidth resources in advance during periods of high user activity to reduce freezes.
[0151] Use the TensorFlow machine learning framework to predict trends and detect anomalies in monitoring data. TensorFlow can build various machine learning models to predict future data trends and identify anomalies in the data by learning from historical data.
[0152] TensorFlow is a machine learning framework in which data is represented in the form of tensors, and the calculation process of the model can be viewed as the flow of tensors in a computational graph.
[0153] Based on the time series algorithm suite 54 and the high-performance computing engine 57 (such as TCN, time convolution network), real-time calculation and visual automatic analysis are performed on the cloud platform monitoring data to generate resource usage trend reports. By processing and analyzing the time series data, abnormal behavior of the cloud platform is discovered.
[0154] In a preferred embodiment, it also includes:
[0155] Alarm and report module 6, alarm and report module 6 connects to data access and analysis module 5, refer to Figure 7 ,include,
[0156] The alarm judgment unit 61 sets an indicator threshold according to business requirements, compares the standardized data with the indicator threshold, and generates an alarm message and sends it to the alarm contact person when the standardized data exceeds the set indicator threshold;
[0157] The alarm classification unit 62 is connected to the alarm judgment unit 61 and configures different alarm levels for different indicator thresholds from multiple dimensions;
[0158] The report generation unit 63 is connected to the alarm judgment unit 61 and the alarm classification unit 62, and performs comprehensive analysis on the standardized data to obtain a periodic monitoring report.
[0159] Specifically, the alarm and report module 6 supports the setting of default and custom alarm rules. For example, if the CPU, memory or disk usage exceeds the set threshold, it will automatically generate alarm information and notify the set alarm contacts in real time through SMS, email and other methods, and provide solutions to ensure that the alarm problem is handled in a timely manner.
[0160] Take the cloud platform of an e-commerce enterprise as an example. During promotional activities, the business volume will increase significantly. In order to ensure the stable operation of the cloud platform, the operation and maintenance team sets the following alarm rules in the alarm and report module based on historical experience and business needs:
[0161] CPU usage threshold: The CPU usage threshold for normal operation is set to 70% (this is the default threshold set based on the stable CPU load under normal business volume). However, considering the possible traffic peak during promotional activities, a custom temporary threshold is set to 85%. When the CPU usage exceeds this value, an alarm message is automatically generated and the set alarm contacts are notified in real time through SMS, email and other methods.
[0162] Memory usage threshold: Under normal circumstances, the memory usage threshold is set to 80%. If the memory usage remains above 90% for 15 consecutive minutes, an alarm is triggered.
[0163] Disk usage threshold: Since e-commerce platforms store a large amount of product images, order data, etc., disk space is relatively tight, so the disk usage threshold is set to 85%. If the disk usage exceeds 95%, an alarm will be issued immediately.
[0164] During a promotion, the traffic on the cloud platform suddenly increased. Monitoring data showed that the CPU usage of a key server rose rapidly from 70% to 90% within 10 minutes, exceeding the set temporary threshold of 85%. At this time, the alarm and report module 6 immediately and automatically generates an alarm message, including the specific identification of the server, the current value of the CPU usage, the situation that exceeds the threshold, and the business scope that may be affected (for example, if the server is responsible for processing product search requests, it may cause the search response to slow down). The module sends the alarm information to relevant members of the operation and maintenance team, including engineers responsible for server maintenance, system administrators, and heads of business departments, through pre-set SMS and email notification methods. At the same time, based on the built-in knowledge base and previous operation and maintenance experience, the system provides possible solutions for this alarm:
[0165] (1) It is recommended to check whether there are abnormal processes on the server that occupy too much CPU resources. If so, try to terminate these processes.
[0166] (2) Consider temporarily increasing the server's resource allocation, such as allocating some other non-critical business resources to the server to relieve CPU pressure.
[0167] (3) Analyze the current business traffic to see if there is any malicious attack or unreasonable traffic peak. If so, take appropriate traffic restriction or protection measures.
[0168] Supports alarm classification, supports configuring different alarm levels for different thresholds, and supports providing multi-dimensional settings such as indicator thresholds, statistical cycles, duration cycles, and alarm frequencies. According to different thresholds and business impacts, the alarm and report module 6 classifies alarms:
[0169] Low-level alarms are defined as low-level alarms when the CPU usage is between 85% and 90% and lasts no longer than 30 minutes. In this case, the system only sends an email to notify the operation and maintenance personnel to pay attention, but no emergency measures are required immediately. The operation and maintenance personnel can check the server status at a convenient time and analyze whether there are potential problems.
[0170] Medium-level alarm: If the CPU usage exceeds 90% and lasts between 30 minutes and 1 hour, or the memory usage is between 90%-95% for more than 20 minutes, it is considered a medium-level alarm. At this time, in addition to sending email and SMS notifications, the system will automatically repeat the alarm information every 15 minutes until the problem is resolved. At the same time, the operation and maintenance personnel are required to take corresponding measures within 1 hour and report the situation to their superiors.
[0171] High-level alarms are triggered when serious situations occur, such as disk usage exceeding 95%, or CPU usage exceeding 95% for more than 1 hour. The system will immediately notify all key members of the operation and maintenance team, including technical supervisors and senior company leaders, via SMS and phone calls. Relevant personnel are required to put aside other work and immediately solve the problem with all their efforts to avoid long-term business interruption.
[0172] Combined with the analysis engine analysis interface, it generates periodic monitoring reports.
[0173] Combined with the analysis interface of the data access and analysis engine, the alarm and report module 6 generates a periodic monitoring report of the cloud platform every day:
[0174] Resource usage trend chart: The report uses a line chart to show the changing trends of the CPU, memory, and disk usage of the cloud platform over the past week. Operation and maintenance personnel can intuitively see the peak and trough periods of resource usage, and whether there are abnormal fluctuations.
[0175] For example, during a promotion, CPU and memory usage increased significantly, and then gradually returned to normal levels after the promotion ended.
[0176] Business performance indicator statistics: statistics on key business performance indicators of e-commerce platforms, such as product search response time, order processing success rate, page loading speed, etc. Through the statistical data and change trends of these indicators, the business department can understand the user experience, and the technical department can analyze the relationship between these performance indicators and cloud platform resource usage in order to carry out targeted optimization.
[0177] The alarm event summary lists all alarm events that occurred in the past day, including the time, level, cause, processing results and other information of the alarm. This helps the operation and maintenance team to conduct post-event analysis, summarize lessons learned, optimize alarm rules and solutions, and improve the ability to deal with similar problems in the future.
[0178] The data storage 10 comprises,
[0179] Database node 101, namely GreptimeDB node, represents a specific database instance for storing data and is responsible for storing cloud platform monitoring data.
[0180] The virtual machine node 102 refers to a data storage node related to a virtual machine running in the cloud platform. Virtual machines are widely used in the cloud platform, and their operating data needs to be stored and analyzed.
[0181] The file node 103 is a node for storing file data. There will be a large number of file operations on the cloud platform, such as log files, configuration files, etc. These file data need to be effectively stored and managed.
[0182] The monitoring management unit 104 is responsible for monitoring each node of the data storage, ensuring the stability and security of the data storage, and timely discovering and handling problems that may occur in the storage nodes.
[0183] The cloud platform monitoring data in the data access and analysis module 5 is data collected directly from the cloud platform, including but not limited to indicators such as CPU utilization, memory utilization, network bandwidth, and disk I / O. These data reflect the operating status of the cloud platform.
[0184] Information basic data serves as the basic supporting data for the operation of the cloud platform, including the configuration information, user information, permission information, etc. of the cloud platform. These data are important bases for the normal operation and management of the cloud platform.
[0185] Historical data is data generated during the operation of the cloud platform, which is very helpful for analyzing the long-term operation trends and performance changes of the cloud platform.
[0186] The time series algorithm suite 54 is used to process time series data. For example, cloud platform monitoring data is usually generated in chronological order. The time series algorithm suite 54 can help analyze the characteristics of data such as change trends and periodicity over time.
[0187] The distributed computing framework 55 provides a method for distributed data processing on multiple computers (nodes), which can improve the efficiency and speed of data processing and is suitable for processing large-scale cloud platform monitoring data.
[0188] The storage and computing separation suite 56 separates the data storage and computing functions, which can flexibly expand storage and computing resources, optimize the configuration according to actual needs, and improve the overall performance of the system.
[0189] The high-performance computing engine 57 provides powerful computing power for data processing, accelerates the data analysis and calculation process, and ensures that results can be obtained quickly when processing large amounts of cloud platform monitoring data.
[0190] Metadata storage is responsible for storing metadata of monitoring data, such as the source, format, definition and other information of the data, to facilitate data management and query.
[0191] The intelligent prediction and alarm operation and maintenance algorithm analyzes the cloud platform monitoring data, predicts possible problems or abnormal situations in the future, and generates corresponding alarm information to help operation and maintenance personnel take timely measures.
[0192] In an embodiment of a system monitoring service, a monitoring task is first set up, and a data collection module 1 is deployed on each node of the cloud platform to collect system resources such as CPU, memory, and disk at regular intervals. The collected data is securely transmitted to the core storage and analysis engine, such as checking the changes in CPU, memory, and disk utilization in the past week, and finding system performance bottlenecks. Custom thresholds are set, and the system automatically sends alarm notifications and automatically generates monitoring data logs to facilitate problem location and analysis afterwards.
[0193] Under the huge challenges of increasingly complex business logic and cloud platform operation and maintenance, it can provide an efficient and easy-to-use automated cloud platform monitoring and operation and maintenance platform. This application allows operation and maintenance personnel to break away from the difficult-to-sort out complex dependencies between applications and focus on handling abnormal alarm situations.
[0194] At the same time, the solution provided by this application reduces the burden on cloud platform operation and maintenance personnel to sort out call links, troubleshoot and locate abnormal clusters. Cloud platform operation and maintenance personnel do not have to sort out complex interface and database call relationships, which reduces the difficulty of cloud platform operation and maintenance management, saves human resources, and improves operation and maintenance efficiency.
[0195] In society, the solution proposed in this application can improve the public security and monitoring capabilities of the cloud platform, improve the reliability of cloud services, and shorten the average fault recovery time to less than 10 minutes, effectively improving service stability.
[0196] Economically, the solution proposed in this application saves manual intervention by operation and maintenance personnel through automated data collection and alarms, saving the platform an average of more than 20% of labor costs and economic losses caused by cloud service interruptions.
[0197] The cloud platform monitoring data storage and analysis system designed based on GreptimeDB in the present invention can specifically solve the cloud platform monitoring data storage and analysis problems, provide massive elastic storage space, reliable and efficient analysis and computing environment and high-speed information retrieval; provide a time series algorithm suite 54, a distributed computing framework 55 and a high-performance computing engine 57; provide a storage and computing separation suite 56, a visual analysis tool; provide adaptive data compression and data encryption algorithms; provide automatic security audit logs and intelligent alarm algorithms. It realizes automatic extraction of unified interface information, automatic storage of interface information, analysis and alarm solutions, provides more efficient and stable solutions, gives full play to the value of monitoring data, and ensures the smooth operation of the cloud platform.
[0198] The above description is only a preferred embodiment of the present invention, and does not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A cloud platform monitoring data storage and analysis system based on GreptimeDB, characterized in that: include, The data collection module is used to collect monitoring data from the business nodes of the heterogeneous cloud platform; A metadata management module, connected to the data acquisition module, for extracting, storing, integrating, associating, maintaining and updating metadata of the monitoring data; A data transmission and storage module is connected to the metadata management module, stores the monitoring data into a database in combination with the metadata, and performs cold and hot data hierarchical storage of the monitoring data according to the generation time and access frequency in the metadata; A data cleaning and standardization module, connected to the data transmission and storage module, cleans and standardizes the monitoring data stored in the database to obtain standardized data; The data access and analysis module is connected to the data cleaning and standardization module and is used for performing stream processing and batch processing on the standardized data.
2. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1 is characterized in that: The data acquisition module includes: A data collection unit interacts with the service node through a data collection tool to obtain the monitoring data; The asynchronous transmission unit is connected to the data acquisition unit and transmits the monitoring data asynchronously through a message queue.
3. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1 is characterized in that: The metadata management module includes: A metadata extraction unit, used for extracting and standardizing the metadata in the monitoring data; A metadata query and retrieval unit, connected to the metadata extraction unit, for providing a graphical interface to display the relationship between the metadata, and searching the monitoring data and the corresponding metadata by keywords; The metadata maintenance unit is connected to the metadata extraction unit and is used to store and integrate the metadata and update the metadata in real time as the monitoring data changes.
4. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1 is characterized in that: The data transmission and storage module includes: A data transmission unit, combining the metadata, and transmitting the monitoring data via a hypertext transfer protocol; a hot and cold data identification unit, connected to the data transmission unit, for identifying hot and cold data according to the access frequency in the metadata, identifying the data as hot data when the access frequency is greater than a set threshold, and identifying the data as cold data when the access frequency is less than or equal to the set threshold; A hierarchical management unit, connected to the hot and cold data identification unit, performs cold and hot data hierarchical storage management on the monitoring data; A data storage unit, connected to the hierarchical management unit, for storing the monitoring data in the database, locally caching the hot data, and object storing the cold data; The data protection unit is connected to the data transmission unit, and verifies the monitoring data in the data transmission unit through a data transmission guarantee mechanism. The data transmission unit transmits the verified monitoring data to the data storage unit.
5. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 4 is characterized in that: The data transmission guarantee mechanism includes: Before data transmission, the monitoring data is hashed by a hash function to map the monitoring data to a fixed first hash value. Before the data storage unit receives the data, the hash calculation is performed again to obtain a second hash value. The first hash value and the second hash value are compared. When the first hash value is the same as the second hash value, the monitoring data is stored in the database. When the first hash value is different from the second hash value, the data transmission unit retransmits the monitoring data.
6. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 4 is characterized in that: The data transmission and storage module also includes: The data encryption unit is connected to the data transmission unit, performs a handshake protocol before the monitoring data is transmitted through a security protocol, generates a session key, and encrypts the monitoring data through the session key.
7. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1 is characterized in that: The data cleaning and standardization module includes: A singular value detection unit, which identifies and corrects singular values in the monitoring data based on a preset data range and statistical model; A missing value filling unit, connected to the singular value detection unit, fills the missing parts in the monitoring data based on the knowledge graph; A duplicate data removal unit, connected to the missing value filling unit, screens out the same monitoring data through data comparison and deletes them; A standardization unit is connected to the duplicate data removal unit, and performs dimension conversion on the monitoring data to obtain the standardized data.
8. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1 is characterized in that: The data access and analysis module includes: A stream processing unit, which performs time-based or data volume-based aggregation analysis on the real-time data in the standardized data through a distributed stream processing framework; The batch processing unit decomposes the historical data in the standardized data into multiple data sets through the distributed system basic framework, processes them on different nodes respectively, and summarizes and reprocesses the processing results.
9. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 8, characterized in that: The data access and analysis module also includes: The anomaly monitoring unit calculates and visualizes the standardized data through a time series algorithm suite and a computing engine to generate a resource usage trend report.
10. The cloud platform monitoring data storage and analysis system based on GreptimeDB according to claim 1, characterized in that: Also includes, An alarm and report module, the alarm and report module is connected to the data access and analysis module, including: An alarm judgment unit sets an indicator threshold according to business requirements, compares the standardized data with the indicator threshold, and generates an alarm message and sends it to an alarm contact person when the standardized data exceeds the set indicator threshold; An alarm grading unit, connected to the alarm judgment unit, configured different alarm levels for different indicator thresholds from multiple dimensions; The report generation unit is connected to the alarm judgment unit and the alarm classification unit, and performs comprehensive analysis on the standardized data to obtain a periodic monitoring report.
Citation Information
Patent Citations
Multi-cloud heterogeneous resource unified monitoring method and device based on cloud native
CN116594836A
Cloud platform monitoring service system and method based on big data analysis
CN119046081A