Visual operation and maintenance management system based on distributed cloud platform
The visual operation and maintenance management system of the distributed cloud platform solves the performance bottlenecks and reliability problems of traditional centralized systems, realizes efficient data processing and real-time monitoring, improves system stability and fault early warning capabilities, and is suitable for enterprise-level operation and maintenance management.
Patent Information
- Application Number
- CN202511349991.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional centralized visualization systems suffer from high single-point failure risk, poor computing power scalability, difficulty in cross-platform adaptation, and lack of real-time rendering support for dynamic data streams, leading to performance bottlenecks and reliability issues.
A visual operation and maintenance management system based on a distributed cloud platform is adopted, including an infrastructure layer, a data acquisition layer, a data processing layer, a data storage layer, a visualization service layer, and a user interaction layer. It utilizes Kubernetes to orchestrate and manage a microservice architecture, combined with a high-availability cluster of Kafka, Elasticsearch, InfluxDB, and MySQL, to achieve real-time data acquisition, processing, and storage. It also generates diverse charts through ECharts and Grafana for real-time monitoring and alerts.
It improves the stability and scalability of the system, enables high-concurrency collection and processing of massive amounts of operation and maintenance data, supports monitoring and early warning at the second or even millisecond level, provides rich chart interaction functions, enhances fault early warning capabilities, and is suitable for complex enterprise-level operation and maintenance scenarios.
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, specifically a visual operation and maintenance management system based on a distributed cloud platform. Background Technology
[0002] Current visualization systems suffer from pain points such as high risk of single point of failure, poor scalability of computing power, and difficulty in cross-platform adaptation. Traditional centralized architectures are prone to latency when dealing with high-concurrency data requests, and existing visualization tools lack support for real-time rendering of dynamic data streams.
[0003] The performance bottlenecks and reliability issues of traditional centralized systems are technical problems that need to be addressed. Summary of the Invention
[0004] The technical objective of this invention is to address the above-mentioned shortcomings by providing a visualized operation and maintenance management system based on a distributed cloud platform, thereby resolving the performance bottlenecks and reliability issues of traditional centralized systems.
[0005] In a first aspect, the present invention provides a visual operation and maintenance management system based on a distributed cloud platform, comprising an infrastructure layer and a data acquisition layer, a data processing layer, a data storage layer, a visualization service layer, and a user interaction layer deployed on the infrastructure layer; The infrastructure layer is a distributed cloud platform environment that includes computing nodes, storage nodes, and network nodes, and supports containerized deployment and elastic scaling. The data acquisition layer is used to collect logs, metrics and event data from multi-source heterogeneous systems in real time as monitoring data, and store the monitoring data in the Kafka cluster. The data processing layer is used to preprocess the monitoring data in the Kafka cluster and transmit the preprocessed monitoring data to the data storage layer. The data storage layer is used to store preprocessed monitoring data through storage components based on a hot-temperature-cold tiering strategy. The visualization service layer is used to receive front-end requests, read relevant data from the storage layer based on the front-end requests, call the analysis engine to analyze the relevant data, generate diverse charts based on the analysis results, monitor specified key indicators in real time based on the front-end requests, and trigger multi-channel early warning notifications based on the monitoring results. The user interaction layer provides a visual interface including web and mobile interfaces. Through the visual interface, it provides users with monitoring services and permission management services. The monitoring service allows users to customize the monitoring data in the database, and the permission management service restricts users' access permissions to the monitoring data in the database.
[0006] As a preferred option, the infrastructure layer, data acquisition layer, data processing and storage layer, visualization service layer, and user interaction layer are orchestrated and managed by Kubernetes.
[0007] Preferably, the data acquisition layer is used to perform the following operations: For heterogeneous systems including servers, container instances, and network devices, metrics are collected based on a lightweight collection agent service deployed in the heterogeneous system. The collected metrics include CPU utilization, memory usage, disk I / O throughput, and network traffic. Logs and event data are collected in real time from application logs, slow query records in the database, and distributed API call chains as business data. During the data collection process, a lightweight data collection agent service preprocesses the collected metrics and business data. This preprocessing includes timestamp standardization, metric normalization, and log structuring. The preprocessed metrics and business data are then transmitted to a Kafka cluster via a two-way authenticated TLS 1.3 encrypted channel. The Kafka cluster is deployed with multiple nodes and configured with multiple partitions and replication factors for each topic. The consumer group mode allows multiple subsystems, including monitoring and analysis and alarm engines, to consume the same data stream in parallel, achieving data multiplexing. Metrics and business data are retained in the Kafka cluster for a predetermined period.
[0008] As a preferred option, the data processing layer adopts a hybrid stream and batch processing architecture to perform the following operations: Perform data cleaning on monitoring data in the Kafka cluster. Data cleaning includes filtering null records, correcting format errors, and filling in missing fields. For the monitoring data after cleaning, perform a standardized transformation; For the standardized monitoring data, Flink is called to perform multi-dimensional statistics by using KeyedProcessFunction, combining business dimensions with sliding time windows, and based on the 3-sigma anomaly detection algorithm, to identify anomaly patterns including sudden increases in indicators and periodic deviations, and to obtain anomaly identification results. Based on the anomaly identification results, high-frequency metrics are written to the Elasticsearch cluster, and a cold and hot node separation architecture is adopted. Low-frequency metrics are stored in the InfluxDB sharded cluster, and automatic downsampling is achieved through the RP strategy. Data is aggregated monthly and permanently stored. For the stored data, Spark on Kubernetes is used to schedule complex analyses on terabyte-level historical data in InfluxDB, including capacity prediction based on the ARIMA model, root cause localization through association rule mining, and resource utilization clustering analysis. The analysis results are written to the star schema fact table of the data warehouse via JDBC, while generating dimension table update records. Finally, trend comparison analysis across time periods is achieved through Superset.
[0009] As a preferred option, the storage components in the data storage layer include Elasticsearch, InfluxDB, object storage system, and MySQL high-availability cluster. Elasticsearch, InfluxDB, object storage system, and MySQL high-availability cluster are connected through a unified data bus, supporting on-demand calls and cross-layer queries. For preprocessed monitoring data, hot data is written to Elasticsearch, supporting millisecond-level full-text search, complex aggregation analysis, and real-time visualization queries. Warm data is migrated to InfluxDB, which provides storage and query services based on time-series compression algorithms and query capabilities. Cold data is automatically archived to an object storage system through lifecycle management strategies. Structured information is stored in a MySQL high-availability cluster, providing fast index access. The structured information includes metadata, user permission configurations, and dashboard definitions.
[0010] As a preferred option, the visualization service layer consists of independent microservices. After receiving a front-end request, it calls the data query interface to obtain preprocessed monitoring data from the database and uses analysis engines including ECharts and Grafana to generate diverse charts, including line charts, bar charts, topology charts, and heatmaps. The visualization service layer supports custom dashboards, multi-chart linkage, drill-down analysis, and time range filtering. It also integrates an alarm engine that monitors key indicators in the database in real time based on Prometheus Alertmanager or custom rules, triggering multi-channel alert notifications, including emails and SMS messages.
[0011] As a preferred option, the user interaction layer provides a responsive web interface developed based on Vue.js or React, and is adapted for both PC and mobile devices.
[0012] Preferably, the system provides a logging service to record operation records and generate logs for the infrastructure layer, data acquisition layer, data processing layer, data storage layer, visualization service layer, and user interaction layer. All logs are collected and stored in the ELK stack.
[0013] Preferably, all services provided in the system are configured with health checks and automatic restart strategies, and the system supports canary releases and version rollbacks.
[0014] The visual operation and maintenance management system based on a distributed cloud platform of the present invention has the following advantages: 1. Built on a distributed cloud platform, it adopts a microservices and containerized architecture, possessing excellent elastic scalability and high availability. Computing, storage, and network resources can be dynamically allocated, supporting horizontal scaling, effectively handling the high-concurrency collection and processing needs of massive operational data, and improving the system's stability and fault tolerance; 2. By introducing the Kafka message queue, data collection and processing are decoupled. Combined with the Flink stream processing engine, low-latency, high-throughput real-time processing of multi-source heterogeneous data such as logs, metrics, and events is achieved. The integrated stream and batch processing architecture takes into account both real-time monitoring and historical analysis needs, significantly improving data processing efficiency and response speed, and meeting the monitoring and early warning requirements at the second or even millisecond level. 3. A hot-warm-cold tiered storage strategy is adopted, storing data in Elasticsearch, InfluxDB and object storage respectively. This not only ensures fast query performance for frequently accessed data, but also reduces long-term storage costs, achieving a balance between performance and economy. Metadata and user configuration information are managed through a relational database to ensure data consistency and security. 4. In terms of visualization, it provides a rich variety of chart types and interactive functions, supports custom dashboards, multi-chart linkage, drill-down analysis and real-time refresh, enabling operation and maintenance personnel to intuitively grasp the system's operating status, quickly locate anomalies and bottlenecks, and the alarm engine supports flexible rule configuration and multi-channel notification, enhancing fault early warning and response capabilities. 5. Supports multi-tenant management and fine-grained access control, suitable for complex enterprise-level operation and maintenance scenarios. The overall architecture is modular and loosely coupled, facilitating maintenance and upgrades, and improving the system's scalability and manageability. Detailed Implementation
[0015] The present invention will be further described below with reference to specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments are not intended to limit the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0016] This invention provides a visual operation and maintenance management system based on a distributed cloud platform to solve the technical problems of performance bottlenecks and reliability in traditional centralized systems.
[0017] Example: The present invention provides a visual operation and maintenance management system based on a distributed cloud platform, comprising an infrastructure layer and a data acquisition layer, a data processing layer, a data storage layer, a visualization service layer, and a user interaction layer deployed on the infrastructure layer.
[0018] The infrastructure layer is a distributed cloud platform environment that includes compute nodes, storage nodes, and network nodes, supporting containerized deployment and elastic scaling.
[0019] The infrastructure layer, data acquisition layer, data processing and storage layer, visualization service layer, and user interaction layer are orchestrated and managed by Kubernetes.
[0020] The system is deployed in a distributed cloud platform environment, adopting a microservice architecture. Each module is deployed via Docker containers and orchestrated and managed by Kubernetes to achieve elastic resource scheduling and high availability. The system runs on a cloud infrastructure consisting of multiple compute nodes, storage nodes, and network nodes, supporting horizontal scaling and automatic failover.
[0021] The data acquisition layer is used to collect logs, metrics, and event data from multi-source heterogeneous systems in real time as monitoring data, and store the monitoring data in a Kafka cluster.
[0022] In a specific implementation, the data acquisition layer is used to perform the following operations: (1) For heterogeneous systems including servers, container instances and network devices, metrics are collected based on the lightweight collection agent service deployed in the heterogeneous system. The collected metrics include CPU utilization, memory usage, disk I / O throughput and network traffic. Logs and event data are collected in real time from application logs, slow query records in the database and distributed API call chains as business data. (2) During the collection process, the lightweight collection agent service preprocesses the collected metrics and business data. Through preprocessing, timestamp standardization, metric normalization and log structuring operations are performed. The preprocessed metrics and business data are then transmitted to the Kafka cluster through a two-way authenticated TLS 1.3 encrypted channel. The Kafka cluster is deployed with multiple nodes and configured with multiple partitions and multiple replication factors for each topic. The consumer group mode allows multiple subsystems, including monitoring analysis and alarm engines, to consume the same data stream in parallel, realizing multiplexing of data. The metrics and business data are retained in the Kafka cluster for a predetermined period of time.
[0023] In this embodiment, during the data acquisition phase, the monitoring system deploys lightweight acquisition agents on each monitored server, container instance, and network device. These agents employ a resource-optimized design, achieving full-dimensional data acquisition with less than 1% CPU overhead. This includes system-level metrics such as CPU utilization, memory usage, disk I / O throughput, and network traffic, while simultaneously capturing application logs (e.g., Nginx access logs, Java GC logs), slow database query records (SQL statements exceeding 500ms), and distributed API call chains (cross-service calls traversed by TraceID). During acquisition, the agents preprocess the raw data, including timestamp standardization (unifying to UTC+8 timezone), metric normalization (e.g., converting memory units to MB), and log structuring (parsing text logs into JSON format). The processed data is transmitted through a two-way authenticated TLS 1.3 encrypted channel to ensure security during public network transmission. The data is ultimately written to a highly available Kafka cluster, which is deployed with 3 nodes, configured with 6 partitions and 2 replication factors for each topic, and a single cluster can support a write throughput of 200,000 TPS. Kafka's ISR (In-Sync Replicas) mechanism, combined with the min.insync.replicas=1 configuration, ensures both data reliability and write performance. Consumer Group mode allows multiple subsystems, such as monitoring and analysis and alerting engines, to consume the same data stream in parallel, achieving data multiplexing. Data is retained in Kafka for 7 days, providing a buffer window for possible backtracking analysis.
[0024] The data processing layer is used to preprocess the monitoring data in the Kafka cluster and transmit the preprocessed monitoring data to the data storage layer.
[0025] The data processing layer adopts a hybrid stream and batch processing architecture to perform the following operations: (1) Perform data cleaning on the monitoring data in the Kafka cluster. Data cleaning includes filtering out null records, fixing format errors, and filling in missing fields. (2) Perform standardization transformation on the cleaned monitoring data; (3) For the standardized and transformed monitoring data, Flink is called to perform multi-dimensional statistics by using KeyedProcessFunction and combining business dimensions with sliding time windows. Based on the 3-sigma anomaly detection algorithm, anomaly patterns including sudden increases in indicators and period deviations are identified to obtain anomaly identification results. (4) Based on the anomaly identification results, high-frequency indicators are written to the Elasticsearch cluster. A cold and hot node separation architecture is adopted, and low-frequency indicators are stored in the InfluxDB sharded cluster. Automatic downsampling is achieved through the RP strategy, and data is aggregated monthly and permanently stored. (5) For the stored data, Spark on Kubernetes scheduling is used to perform complex analysis on the TB-level historical data in InfluxDB, including capacity prediction based on the ARIMA model, root cause localization by association rule mining, and resource utilization clustering analysis. The analysis results are written to the star schema fact table of the data warehouse via JDBC, and dimension table update records are generated at the same time. Finally, trend comparison analysis across time periods is achieved through Superset.
[0026] In this embodiment, data processing adopts a hybrid "stream and batch" architecture design, balancing real-time performance with the deep analytical capabilities of batch processing. The real-time stream processing module is built on Apache Flink and deployed in a distributed cluster mode. The processing flow includes a multi-level pipeline: first, data cleaning is performed (filtering null records, correcting format errors, and completing missing fields), and then standardization transformation is performed (converting CPU utilization to percentage, network bandwidth to Mbps, and aligning timestamps to 5-second granularity). At the aggregation layer, Flink uses KeyedProcessFunction to perform multi-dimensional statistics based on business dimensions (host IP / container ID / service name) combined with sliding time windows (1 minute / 5 minutes / 15 minutes), while running a 3-sigma-based anomaly detection algorithm to identify abnormal patterns such as sudden increases in metrics (e.g., CPU instantaneously exceeding 90%) or periodic deviations (e.g., daily fluctuations greater than 200%). The processing results are output in two paths: high-frequency metrics are written to an Elasticsearch cluster using a hot / cold node separation architecture, with hot nodes (SSD storage) retaining data from the most recent 7 days to support millisecond-level response times for the real-time dashboard; low-frequency metrics are stored in a sharded InfluxDB cluster, with automatic downsampling achieved through a Retention Policy (RP) strategy. Raw data is retained for 30 days, and monthly aggregated data is permanently saved. The batch processing module uses Spark on Kubernetes scheduling, starting offline jobs daily at midnight to perform complex analyses on terabytes of historical data in InfluxDB, including capacity prediction based on the ARIMA model, root cause localization through association rule mining (e.g., the causal relationship between increased disk IOPS and slow MySQL queries), and resource utilization clustering analysis. The analysis results are written to the star schema fact table of the data warehouse via JDBC, simultaneously generating dimension table update records. Finally, trend comparison analysis across time periods is achieved through Superset. The system ensures consistency in the definition of metrics for both streaming and batch processing through unified metadata management, with errors controlled within 0.1%.
[0027] The data storage layer is used to store preprocessed monitoring data through storage components based on a hot-temperature-cold tiered strategy.
[0028] In practice, the data storage layer includes Elasticsearch, InfluxDB, an object storage system, and a high-availability MySQL cluster. These components are interconnected via a unified data bus, supporting on-demand access and cross-layer queries. For preprocessed monitoring data, hot data is written to Elasticsearch, supporting millisecond-level full-text search, complex aggregation analysis, and real-time visualization queries. Warm data is migrated to InfluxDB, providing storage and query services based on time-series compression algorithms and query capabilities. Cold data is automatically archived to the object storage system through lifecycle management strategies. Structured information is stored in the high-availability MySQL cluster, providing fast index access. This structured information includes metadata, user permission configurations, and dashboard definitions.
[0029] In this embodiment, data storage employs a tiered strategy to achieve an optimal balance between performance and cost. Hot data (the most recent 24 hours) is written to Elasticsearch, supporting millisecond-level full-text search, complex aggregation analysis, and real-time visualization queries. Warm data (1 week to 3 months) is migrated to InfluxDB, leveraging its efficient time-series compression algorithm and query capabilities to balance processing performance and storage costs. Cold data (over 3 months) is automatically archived to an object storage system (such as Amazon S3 or Alibaba Cloud OSS) through a lifecycle management strategy, reducing long-term storage overhead. Structured information such as metadata, user permission configurations, and dashboard definitions is stored in a highly available MySQL cluster, ensuring data consistency, transaction integrity, and fast index access. All storage components are interconnected through a unified data bus, supporting on-demand calls and cross-tier queries, improving the overall maintainability and scalability of the system.
[0030] The visualization service layer is used to receive front-end requests, read relevant data from the storage layer based on the front-end requests, call the analysis engine to analyze the relevant data, generate diverse charts based on the analysis results, monitor specified key indicators in real time based on the front-end requests, and trigger multi-channel early warning notifications based on the monitoring results.
[0031] The visualization service layer consists of independent microservices. After receiving requests from the front end, it calls the data query interface to obtain preprocessed monitoring data from the database. It then uses analysis engines, including ECharts and Grafana, to generate diverse charts, such as line charts, bar charts, topology maps, and heatmaps. The visualization service layer supports custom dashboards, multi-chart linkage, drill-down analysis, and time range filtering. It also integrates an alarm engine that monitors key indicators in the database in real time based on Prometheus Alertmanager or custom rules, triggering multi-channel alert notifications, including emails and SMS messages.
[0032] In this embodiment, the visualization service layer consists of independent microservices. After receiving requests from the front end, it calls the data query interface to obtain aggregated results and uses engines such as ECharts and Grafana to generate diverse charts such as line charts, bar charts, topology maps, and heatmaps. It supports custom dashboards, multi-chart linkage, drill-down analysis, and time range filtering. The system integrates an alarm engine, based on Prometheus Alertmanager or custom rules, to monitor key indicators (such as CPU > 90% or a sudden increase in error rate) in real time, triggering notifications via email, SMS, and other channels.
[0033] The user interaction layer provides a visual interface, including web and mobile interfaces, to provide users with monitoring and permission management services. The monitoring service allows users to customize monitoring of data in the database, and the permission management service restricts users' access permissions to the monitoring data in the database.
[0034] In practice, the user interaction layer provides a responsive web interface developed based on Vue.js or React, and is adapted for both PC and mobile devices.
[0035] In this embodiment, the user interaction layer provides a responsive web interface, developed based on Vue.js or React, and adapted for both PC and mobile devices. After logging in, users can view the default monitoring view, as well as customize dashboards, set alarm thresholds, and view historical events and processing records. The system supports multi-tenancy and RBAC access control, with different roles (administrators, operations personnel, and visitors) having different data access and operation permissions.
[0036] In this embodiment, the system provides a logging service to record the operation records of the infrastructure layer, data acquisition layer, data processing layer, data storage layer, visualization service layer, and user interaction layer, and generate logs. All logs are collected into the ELK stack.
[0037] In this embodiment, all services provided by the system are configured with health checks and automatic restart strategies, and the system supports canary releases and version rollbacks.
[0038] This embodiment of the system fully leverages the elastic resources and distributed architecture advantages of cloud computing, integrating stream and batch data processing capabilities to achieve efficient collection, low-latency processing, hierarchical storage, and intelligent analysis of massive amounts of operational data. Through an intuitive and interactive visual interface, it helps operations and maintenance personnel quickly grasp system status, identify potential risks, and improve fault response efficiency. The system not only needs to possess high availability and scalability but also support multi-tenant management, access control, and security auditing. It is suitable for various complex scenarios such as cloud computing, edge computing, and the industrial internet, driving the evolution of operations and maintenance management towards automation and intelligence.
[0039] The above provides a detailed description of a visual operation and maintenance management system based on a distributed cloud platform provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A visualized operation and maintenance management system based on a distributed cloud platform, characterized in that, It includes the infrastructure layer and the data acquisition layer, data processing layer, data storage layer, visualization service layer and user interaction layer deployed on the infrastructure layer; The infrastructure layer is a distributed cloud platform environment that includes computing nodes, storage nodes, and network nodes, and supports containerized deployment and elastic scaling. The data acquisition layer is used to collect logs, metrics and event data from multi-source heterogeneous systems in real time as monitoring data, and store the monitoring data in the Kafka cluster. The data processing layer is used to preprocess the monitoring data in the Kafka cluster and transmit the preprocessed monitoring data to the data storage layer. The data storage layer is used to store preprocessed monitoring data through storage components based on a hot-temperature-cold tiering strategy. The visualization service layer is used to receive front-end requests, read relevant data from the storage layer based on the front-end requests, call the analysis engine to analyze the relevant data, generate diverse charts based on the analysis results, monitor specified key indicators in real time based on the front-end requests, and trigger multi-channel early warning notifications based on the monitoring results. The user interaction layer provides a visual interface including web and mobile interfaces. Through the visual interface, it provides users with monitoring services and permission management services. The monitoring service allows users to customize the monitoring data in the database, and the permission management service restricts users' access permissions to the monitoring data in the database.
2. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The infrastructure layer, data acquisition layer, data processing and storage layer, visualization service layer, and user interaction layer are orchestrated and managed by Kubernetes.
3. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The data acquisition layer is used to perform the following operations: For heterogeneous systems including servers, container instances, and network devices, metrics are collected based on a lightweight collection agent service deployed in the heterogeneous system. The collected metrics include CPU utilization, memory usage, disk I / O throughput, and network traffic. Logs and event data are collected in real time from application logs, slow query records in the database, and distributed API call chains as business data. During the data collection process, a lightweight data collection agent service preprocesses the collected metrics and business data. This preprocessing includes timestamp standardization, metric normalization, and log structuring. The preprocessed metrics and business data are then transmitted to a Kafka cluster via a two-way authenticated TLS 1.3 encrypted channel. The Kafka cluster is deployed with multiple nodes and configured with multiple partitions and replication factors for each topic. The consumer group mode allows multiple subsystems, including monitoring and analysis and alarm engines, to consume the same data stream in parallel, achieving data multiplexing. Metrics and business data are retained in the Kafka cluster for a predetermined period.
4. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The data processing layer adopts a hybrid stream and batch processing architecture to perform the following operations: Perform data cleaning on monitoring data in the Kafka cluster. Data cleaning includes filtering null records, correcting format errors, and filling in missing fields. For the monitoring data after cleaning, perform a standardized transformation; For the standardized monitoring data, Flink is called to perform multi-dimensional statistics by using KeyedProcessFunction, combining business dimensions with sliding time windows, and based on the 3-sigma anomaly detection algorithm, to identify anomaly patterns including sudden increases in indicators and periodic deviations, and to obtain anomaly identification results. Based on the anomaly identification results, high-frequency metrics are written to the Elasticsearch cluster, and a cold and hot node separation architecture is adopted. Low-frequency metrics are stored in the InfluxDB sharded cluster, and automatic downsampling is achieved through the RP strategy. Data is aggregated monthly and permanently stored. For the stored data, Spark on Kubernetes is used to schedule complex analyses on terabyte-level historical data in InfluxDB, including capacity prediction based on the ARIMA model, root cause localization through association rule mining, and resource utilization clustering analysis. The analysis results are written to the star schema fact table of the data warehouse via JDBC, while generating dimension table update records. Finally, trend comparison analysis across time periods is achieved through Superset.
5. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The storage components in the data storage layer include Elasticsearch, InfluxDB, object storage system, and MySQL high-availability cluster. Elasticsearch, InfluxDB, object storage system, and MySQL high-availability cluster are connected through a unified data bus, supporting on-demand access and cross-layer queries. For preprocessed monitoring data, hot data is written to Elasticsearch, supporting millisecond-level full-text search, complex aggregation analysis, and real-time visualization queries. Warm data is migrated to InfluxDB, which provides storage and query services based on time-series compression algorithms and query capabilities. Cold data is automatically archived to an object storage system through lifecycle management strategies. Structured information is stored in a MySQL high-availability cluster, providing fast index access. The structured information includes metadata, user permission configurations, and dashboard definitions.
6. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The visualization service layer consists of independent microservices. After receiving front-end requests, it calls the data query interface to obtain pre-processed monitoring data from the database and uses analysis engines including ECharts and Grafana to generate diverse charts, including line charts, bar charts, topology charts, and heatmaps. The visualization service layer supports custom dashboards, multi-chart linkage, drill-down analysis, and time range filtering. It also integrates an alarm engine that monitors key indicators in the database in real time based on Prometheus Alertmanager or custom rules, triggering multi-channel alert notifications, including emails and SMS messages.
7. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The user interaction layer provides a responsive web interface based on Vue.js or React, and is adapted for both PC and mobile devices.
8. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, The system provides a logging service to record operations and generate logs for the infrastructure layer, data acquisition layer, data processing layer, data storage layer, visualization service layer, and user interaction layer. All logs are collected and stored in the ELK stack.
9. The visualized operation and maintenance management system based on a distributed cloud platform according to claim 1, characterized in that, All services provided in the system are configured with health checks and automatic restart policies, and the system supports canary releases and version rollbacks.
Citation Information
Patent Citations
Multi-source heterogeneous data real-time processing system and method based on Flink stream computing technology
CN110245158A
Centralized visualization method and system for operation and maintenance data of cloud computing center
CN110928740A
Internet of Things data real-time calculation service system and construction method
CN115203184A
Supply chain management system for research and development of electronic materials
CN120563008A
Method and system for realizing large-scale node monitoring based on ELK Stack technology
CN120582967A