Cloud operation scheduling system and method based on real-time data stream

Through a cloud operation scheduling system based on real-time data flow, using Apache Kafka and Flink/Spark Streaming for real-time data processing and multi-dimensional analysis, the problem of slow response in traditional systems is solved, and efficient and real-time operation management and optimization are achieved.

CN120335947APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338599.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional cloud operation management and scheduling systems cannot respond to system dynamic changes in real time, resulting in slow response, difficult to meet the needs of high concurrency and high real-time, and there are performance bottlenecks.

Method used

The cloud operation scheduling system based on real-time data flow is adopted, and Apache Kafka is used as the message bus, combined with Apache Flink or Spark Streaming for streaming calculations, data cleaning, conversion and aggregation is performed, and multi-dimensional data analysis is performed through Apache Spark MLlib to generate visual reports or automatically trigger operation processes.

Benefits of technology

Real-time improvement of the system is achieved, can quickly respond to emergencies, improve operation and maintenance efficiency and intelligent management, and provide intuitive decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335947A_ABST
    Figure CN120335947A_ABST
Patent Text Reader

Abstract

The invention provides a cloud operation scheduling system and method based on real-time data streams, belongs to the technical field of distributed computing, and introduces a real-time data stream processing technology to enhance real-time information processing capability, develop a multi-dimensional data analysis module and realize deep mining of operation and maintenance data. According to the system, the real-time performance of the system is improved by efficiently processing and analyzing real-time data, and meanwhile, an administrator is assisted in discovering and solving potential problems and optimizing operation management by analyzing multi-dimensional data. The method is mainly used for operation management and scheduling in the distributed cloud computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of distributed computing technology, big data processing, and cloud computing operation, and particularly to a cloud operation scheduling system and method based on real-time data streams. Background Art

[0002] In a distributed cloud computing environment, traditional operation management and scheduling systems often rely on static data analysis and processing methods. These systems can usually only process historical data and cannot respond to dynamic changes in the system in real time, resulting in slow responses when dealing with emergencies or abnormal situations and making it difficult to achieve efficient and real-time operation management. In addition, traditional systems often have performance bottlenecks when processing large-scale data and are difficult to meet the high concurrency and high real-time requirements of modern cloud computing environments. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a cloud operation scheduling system based on real-time data streams, which can simplify user operations in an intelligent manner, improve operation and maintenance efficiency, and enhance the stability and user experience of the system.

[0004] The technical solution of the present invention is as follows:

[0005] A cloud operation scheduling system based on real-time data streams, comprising:

[0006] A data acquisition layer for receiving data streams from different sources and transmitting them to the subsequent processing layer through a message bus.

[0007] A data processing layer for cleaning, transforming, and aggregating the data streams.

[0008] A data analysis layer for performing multi-dimensional analysis based on the processed data.

[0009] A decision support layer for generating visualization reports or alarms according to the analysis results and providing operation suggestions or automatically triggering predefined operation processes.

[0010] Wherein

[0011] The data acquisition layer uses Apache Kafka as the message bus, and the data processing layer uses Apache Flink or Spark Streaming for stream computing.

[0012] The data processing layer uses Apache Flink or Spark Streaming as the stream computing framework to perform real-time processing on the data in Kafka.

[0013] The data analysis layer includes a machine learning model built using Apache Spark MLlib to identify patterns, predict trends, or detect anomalies.

[0014] The decision support layer includes visualization tools to display the analysis results, and the visualization tools are selected from Elasticsearch Kibana or Tableau.

[0015] Furthermore,

[0016] It also includes a data storage layer for storing massive data, where the data storage layer uses HDFS for storage and uses Cassandra or MongoDB as a NoSQL database.

[0017] In addition, the present invention also provides a method for cloud operation scheduling, including: configuring an Apache Kafka cluster to receive real-time data streams; using Apache Flink or Spark Streaming to perform real-time processing on the data streams; applying machine learning algorithms or statistical methods to perform multi-dimensional analysis on the processed data; generating a visualization report or an alarm according to the analysis results, and providing operation suggestions or automatically triggering a predefined operation process.

[0018] Furthermore,

[0019] The specific steps are as follows:

[0020] Step 1: Configure an Apache Kafka cluster;

[0021] Step 2: Push the raw data generated by servers, applications, and network devices into the Kafka topic through an API interface or a log collector;

[0022] Step 3: Define functions or operators for data cleaning, transformation, and aggregation;

[0023] Step 4: Store the processed data in a distributed cache or a database;

[0024] Step 5: Build a machine learning model using the Apache Spark MLlib library;

[0025] Step 6: Analyze the data through the model to identify patterns, predict trends, or detect anomalies;

[0026] Step 7: Use data visualization tools to display the analysis results;

[0027] Step 8: Set an alarm mechanism to automatically notify when an anomaly is detected or a key metric reaches a threshold.

[0028] The beneficial effects of the present invention are

[0029] The present invention can not only achieve fast processing and analysis of real-time information, significantly improve the real-time performance of the system, and quickly respond to and handle emergencies, but also help administrators discover potential problems and optimization points through multi-dimensional data analysis, improving the management efficiency and intelligent level of the distributed cloud operation system. The data visualization technology provides intuitive decision-making support for administrators, helping to make more scientific and reasonable decisions. Description of the Drawings

[0030] Appendix Figure 1 It is an architecture diagram of a real-time data stream processing system;

[0031] Appendix Figure 2 It is a data processing flow chart;

[0032] Appendix Figure 3 It is a multi-dimensional data analysis flow chart. Detailed Implementation Manner

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] As shown in the appendix Figure 1 The present invention mainly includes the following main components: a data acquisition layer (1), a data processing layer (2), a data analysis layer (3), a decision support layer (4), scalability and reliability (5).

[0035] Among them

[0036] 1. Data acquisition layer (1)

[0037] Data sources: including logs of various cloud services, performance metrics, user behavior data, etc.

[0038] Data access: Using Apache Kafka as a message bus, responsible for collecting data from multiple data sources and aggregating it into a unified data stream.

[0039] 2. Data processing layer (2)

[0040] Real-time processing engine: Using Apache Flink for real-time data processing, including operations such as data cleaning, transformation, and aggregation.

[0041] Storage and Caching: Use databases such as Elasticsearch or Cassandra to store the processed data for easy querying and further analysis.

[0042] 3. Data Analysis Layer (3)

[0043] Multi-dimensional Analysis: Develop specialized analysis modules to support time series analysis, spatial analysis, business logic analysis, etc.

[0044] Machine Learning Model: Integrate machine learning algorithms for advanced analysis tasks such as pattern recognition and predictive analysis.

[0045] 4. Decision Support Layer (4)

[0046] Visualization: Use tools such as Grafana or Kibana to display the analysis results in the form of charts.

[0047] Strategy Formulation: Based on the analysis results, the system automatically generates or recommends operation and maintenance strategies.

[0048] Automated Execution: Combine with the DevOps tool chain to achieve automated execution of operation and maintenance operations.

[0049] 5. Scalability and Reliability (5)

[0050] Horizontal Scalability: The system design considers horizontal scalability to adapt to the growing data volume and user requirements.

[0051] Fault Tolerance Mechanism: Adopt means such as replicas and failovers to ensure the high availability of the system.

[0052] The present invention has the following characteristics:

[0053] 1. Real-time Information Processing Capability

[0054] To ensure that the system can respond to various events in a timely manner and make decisions, the present invention adopts real-time data stream processing technologies. For example, Apache Kafka is used as a message bus to capture data streams from different sources, and streaming computing frameworks such as Apache Flink or Spark Streaming are used for real-time processing and analysis. These technologies can significantly reduce the time delay from data generation to availability, thereby greatly improving the response speed of the system. Specifically, Apache Kafka provides a high-throughput message queue service that can quickly collect data from multiple data sources; while Apache Flink and Spark Streaming can process these data streams within milliseconds, completing operations such as data cleaning, transformation, and aggregation to prepare data for subsequent analysis work. Multi-dimensional data analysis

[0055] 2. Multi-dimensional Data Analysis

[0056] In addition to the basic data processing functions, the present invention particularly emphasizes the importance of multi-dimensional data analysis. In a cloud computing environment, operation and maintenance data usually contains a large amount of spatio-temporal information and business logic associations. Relying solely on single-dimensional data analysis often fails to reveal the essence of problems. The present invention can deeply analyze operation and maintenance data from multiple perspectives such as time series, geographical location distribution, and business processes. For example, through time series analysis of historical data, future resource demand trends can be predicted; through spatial analysis of geographical location data, performance bottlenecks in specific regions can be discovered; through analysis of business logic, the root causes of failures can be identified. This comprehensive data insight not only helps operation and maintenance personnel identify the root causes of problems faster but also provides more accurate solutions, thereby effectively reducing the failure rate and improving the stability and availability of the system.

[0057] The present invention realizes efficient data processing and analysis through real-time data stream processing technology. The data processing process is as shown in Figure 2 the following figure.

[0058] 1. Real-time data stream processing

[0059] The system uses Apache Kafka as the message bus to ensure efficient data transmission from generation to consumption. Apache Kafka utilizes its distributed architecture and persistent message queue to provide guarantee for real-time data transmission. The following are the specific implementation steps:

[0060] Step 1: Configure the Apache Kafka cluster, including setting up broker nodes, topics, and partitions.

[0061] Step 2: Push the raw data generated by servers, applications, network devices, etc. into the Kafka topic through the API interface or log collector.

[0062] 2. Data processing layer

[0063] Use Apache Flink or Spark Streaming as the streaming computing framework to perform real-time processing on the data in Kafka. The specific implementation is as follows:

[0064] Step 3: Define functions or operators for data cleaning, transformation, and aggregation, such as using Flink's Map, Filter, FlatMap, etc. operators.

[0065] Step 4: Store the processed data in a distributed cache or database, such as Redis or MySQL, for subsequent analysis.

[0066] 3. Data analysis layer

[0067] Based on the processed data, apply machine learning algorithms or statistical methods for in-depth analysis. The following are the specific implementation steps:

[0068] Step Five: Use the Apache Spark MLlib library to build machine learning models, such as classification, regression, or clustering models.

[0069] Step Six: Analyze the data through the model to identify patterns, predict trends, or detect anomalies.

[0070] 4. Decision Support Layer

[0071] Based on the analysis results, generate visual reports or alerts, provide operation suggestions, or automatically trigger predefined operation processes. The specific implementation is as follows:

[0072] Step Seven: Use data visualization tools (such as Elasticsearch Kibana or Tableau) to display the analysis results.

[0073] Step Eight: Set up an alert mechanism to automatically notify the operation and maintenance personnel when anomalies are detected or key metrics reach the threshold.

[0074] 5. Key Technology Selection

[0075] Real-time data stream processing: Use Apache Kafka as the message bus and Apache Flink or Spark Streaming for stream computing.

[0076] Data analysis tools: Use components of the Hadoop ecosystem (such as Hive, Pig) for batch data analysis, and Apache Spark MLlib or TensorFlow for machine learning model training for predictive analysis.

[0077] Data storage: Use HDFS for massive data storage, and Cassandra or MongoDB as NoSQL databases. Elasticsearch - for indexing and searching the processed data for quick query.

[0078] A sudden increase in the server load in the cloud data center may lead to a decline in system performance. The present invention captures the real-time data stream of the server load through Apache Kafka. The real-time data processing module quickly analyzes the captured data and identifies the trend of the load increase. The system automatically schedules resources, such as adding servers or optimizing load distribution, according to the analysis results to cope with emergencies.

[0079] Through real-time processing and scheduling, the system successfully avoids performance degradation and ensures the normal operation of the service.

[0080] The specific analysis process of multi-dimensional data analysis is as attached Figure 3 as shown below.

[0081] 1. Multi-dimensional data analysis

[0082] Step 1: Define multi-dimensional analysis metrics, such as time dimension, space dimension, business logic dimension, etc.

[0083] Step 2: Group, aggregate, and statistically analyze the data according to the defined metrics.

[0084] Step 3: Display the multi-dimensional analysis results through visualization tools to help operation and maintenance personnel quickly locate problems.

[0085] 2. Selection of key technologies

[0086] Analysis tools: Use Hive and Pig in the Hadoop ecosystem for batch data analysis, and Apache Spark MLlib for advanced analysis.

[0087] Visualization tools: Elasticsearch Kibana or Tableau. Grafana - provides a flexible data visualization interface to facilitate operation and maintenance personnel to understand the analysis results.

[0088] Automated execution: Ansible or Jenkins - supports the execution of automated operation and maintenance scripts.

[0089] Administrators need to understand the overall operation status of the cloud data center and potential performance bottlenecks. Multi-dimensional data analysis deeply mines historical and real-time operation and maintenance data. The analysis found that the processing time of a certain type of service request is relatively long, resulting in a decline in user experience. Based on the analysis results, the administrator adjusted the processing flow of service requests and optimized resource allocation. After optimization, the processing time of service requests was significantly shortened, and the user experience was improved.

[0090] The above description is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A cloud operation scheduling system based on real-time data streams, characterized in that it includes: A data collection layer for receiving data streams from different sources and transmitting them to the subsequent processing layer through a message bus; A data processing layer for cleaning, transforming, and aggregating the data streams; A data analysis layer for performing multi-dimensional analysis based on the processed data; A decision support layer for generating visual reports or alerts based on the analysis results and providing operation suggestions or automatically triggering predefined operation processes.

2. The system according to claim 1, characterized in that the data collection layer uses Apache Kafka as the message bus, and the data processing layer uses Apache Flink or Spark Streaming for stream computing.

3. The system according to claim 1, characterized in that the data processing layer uses Apache Flink or Spark Streaming as the stream computing framework to perform real-time processing on the data in Kafka.

4. The system according to claim 1, characterized in that the data analysis layer includes a machine learning model built using Apache Spark MLlib to identify patterns, predict trends, or detect anomalies.

5. The system according to claim 1, characterized in that the decision support layer includes visualization tools to display the analysis results, and the visualization tools are selected from Elasticsearch Kibana or Tableau.

6. The system according to claim 1, characterized in that it further includes a data storage layer for storing massive data, where the data storage layer uses HDFS for storage and uses Cassandra or MongoDB as a NoSQL database.

7. A method for cloud operation scheduling, characterized in that, The method includes: Configuring an Apache Kafka cluster to receive real-time data streams; using Apache Flink or Spark Streaming to perform real-time processing on the data streams; performing multi-dimensional analysis on the processed data; generating visual reports or alerts based on the analysis results, and providing operation suggestions or automatically triggering predefined operation processes.

8. The method according to claim 7, characterized in that the specific steps are as follows: Step 1: Configure an Apache Kafka cluster; Step 2: Push the raw data generated by servers, applications, and network devices into the Kafka topic through an API interface or a log collector; Step 3: Define functions or operators for data cleaning, transformation, and aggregation; Step 4: Store the processed data in a distributed cache or database; Step 5: Build a machine learning model using the Apache Spark MLlib library; Step 6: Analyze the data through the model to identify patterns, predict trends, or detect anomalies; Step 7: Use data visualization tools to display the analysis results; Step 8: Set up an alert mechanism to automatically notify when an anomaly is detected or a key metric reaches a threshold.