Two-dimensional layering method for realizing intelligent operation and maintenance of cloud computing environment

By employing a two-dimensional layered approach and intelligent operation and maintenance methods, the problem of low efficiency in traditional operation and maintenance in cloud computing environments has been solved, enabling rapid response and self-repair, thereby improving operation and maintenance efficiency and system stability.

CN121967513APending Publication Date: 2026-05-01NANJING HAOMINGYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING HAOMINGYU INFORMATION TECHNOLOGY CO LTD
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional manual operation and maintenance methods cannot meet the rapidly changing needs of cloud computing environments, resulting in low operation and maintenance efficiency and difficulty in quickly responding to and predicting potential problems.

Method used

A two-dimensional layered approach is adopted, including data collection, storage, analysis, monitoring and automated repair in the vertical dimension, and the division of labor in the strategy, management and operation layers in the horizontal dimension. Combined with monitoring tools, machine learning and automation scripts, intelligent operation and maintenance can be achieved.

Benefits of technology

It improves operational efficiency, enables rapid response and fault repair, reduces manual intervention, enhances the system's self-healing and predictive capabilities, and ensures system stability and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967513A_ABST
    Figure CN121967513A_ABST
Patent Text Reader

Abstract

The invention discloses a two-dimensional layering method for realizing intelligent operation and maintenance of a cloud computing environment, and the method comprises the steps: building an automatic operation and maintenance frame, and dividing the frame into two dimensions, namely a transverse dimension and a longitudinal dimension. According to the two-dimensional layering method for realizing intelligent operation and maintenance of the cloud computing environment, automatic scripts are compiled, repair operation is automatically executed according to specific warning conditions and events, such as service restarting, resource allocation adjustment and construction of a system with a self-healing capability, and automatic self-repair is performed according to a defined strategy when a fault is monitored, so that the two-dimensional layering method has remarkable advantages, and the practicability is high. The system can quickly respond, immediately start repair at the moment of a fault, shorten the system recovery time, reduce the influence of business interruption, quickly restart the fault service, improve the efficiency, save the manual troubleshooting and processing time and enable operation and maintenance personnel to concentrate on more important tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to IT operations and maintenance automation, specifically a two-dimensional layered method for achieving intelligent operations and maintenance in cloud computing environments. Background Technology

[0002] Cloud computing technology is a computing method based on the Internet that distributes computing tasks across numerous distributed computers. It is elastic and scalable, capable of rapidly adjusting computing resources according to demand, and features on-demand self-service, allowing users to independently obtain the resources they need. It integrates resources through resource pooling for efficient allocation and management, and also supports extensive network access, allowing users to access the network anytime, anywhere using various network devices. Its application scenarios include enterprise data storage and backup, big data processing, and website and application hosting. However, with the popularization and application of cloud computing technology, enterprise IT infrastructure has become increasingly large and complex, and traditional manual operation and maintenance methods can no longer meet the rapidly changing needs. Therefore, this invention provides a two-dimensional layered method for realizing intelligent operation and maintenance of cloud computing environments. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a two-dimensional layered method for intelligent operation and maintenance of cloud computing environments, solving the problem that manual operation and maintenance of IT infrastructure can no longer meet rapidly changing needs.

[0004] To achieve the above objectives, this invention provides the following technical solution: a two-dimensional layered method for intelligent operation and maintenance of a cloud computing environment, which constructs an automated operation and maintenance framework and divides the framework into two dimensions, characterized by a horizontal dimension and a vertical dimension, wherein the vertical dimension specifically includes the following steps: S1. Data Collection: Automatically acquire system and application data using monitoring and logging tools; S2. Data storage: The acquired data is stored in the database; S3. Data Analysis: Real-time data analysis to identify anomalies, using methods such as trend analysis and correlation analysis to proactively discover potential problems and risks; S4. Monitoring and Alerts: Set thresholds to automatically trigger alarms and detect problems; S5. Response and Automated Repair: Notify operations and maintenance personnel and execute automated repair measures; S6. Visualization and Feedback: Provide intuitive data displays and establish a feedback mechanism.

[0005] Preferably, the horizontal dimension is specifically divided into the following levels according to function: 1) Strategy layer: Formulating and implementing strategies and allocating resources; 2) Management level, responsible for managing physical hardware resources, including servers, storage devices, and network devices; 3) Operational layer: Ensure the normal operation of the hardware and perform regular maintenance and upgrades.

[0006] Preferably, in step S1, monitoring tools are used to collect system and application performance metrics in real time, including CPU utilization, memory usage, disk I / O, and network traffic. Furthermore, a log management system is used to collect application logs, system logs, and network logs to obtain detailed information about the system's operating status. Then, APIs or SDKs are used to obtain operating data from different systems, services, and devices.

[0007] Preferably, in step S2, the collected data is stored in a high-performance data warehouse to support subsequent analysis and queries. Before data storage, data cleaning and preprocessing are performed to remove duplicate, erroneous, and invalid data.

[0008] Preferably, in step S3, stream processing technology is used to analyze real-time data streams, identify abnormal situations or potential faults in a timely manner, and a predictive model is established by applying machine learning algorithms. By learning from historical data, future performance trends and potential problems are predicted, such as anomaly detection models and fault prediction models.

[0009] Preferably, step S4 sets thresholds for key performance indicators. When system indicators exceed or fall below these thresholds, a warning is triggered. Furthermore, a machine learning-based intelligent warning system is used to reduce false alarms, improve the accuracy and intuitiveness of warnings, and enable maintenance personnel to respond to problems quickly.

[0010] Preferably, in step S5, an automated script is written to automatically perform repair operations based on specific warning conditions and events, such as restarting services, adjusting resource allocation, and building a self-healing system that automatically repairs itself when a fault is detected according to a defined strategy.

[0011] Preferably, in step S6, the system operation status is visualized through a dashboard, providing an intuitive data display so that maintenance personnel can quickly understand the system status and establish a feedback mechanism to continuously monitor the repair effect and key indicators. Beneficial effects

[0012] This invention provides a two-dimensional layered method for intelligent operation and maintenance of cloud computing environments. Compared with existing technologies, it has the following advantages: (1) This two-dimensional layered method for realizing intelligent operation and maintenance of cloud computing environment automatically performs repair operations based on specific warning conditions and events by writing automated scripts, such as restarting services, adjusting resource allocation, and building a system with self-healing capabilities. It automatically repairs itself when a fault is detected according to the defined strategy. It has significant advantages. It can respond quickly and start repair immediately at the moment of the fault, shorten the system recovery time, reduce the impact of business interruption, and at the same time, it can quickly restart faulty services, improve efficiency, save manual troubleshooting and processing time, and allow operation and maintenance personnel to focus on more important tasks.

[0013] (2) This two-dimensional layered method for realizing intelligent operation and maintenance of cloud computing environment triggers warnings by setting thresholds for key performance indicators and using a machine learning-based intelligent warning system to promptly detect abnormal system indicators, prevent problems from worsening, reduce false alarms, prevent operation and maintenance personnel from being disturbed, focus on the real problems, improve the accuracy of warnings, help them quickly locate key issues, enhance intuitiveness, facilitate understanding of problems, ensure that operation and maintenance personnel respond quickly, and reduce losses.

[0014] (3) This two-dimensional layered method for realizing intelligent operation and maintenance of cloud computing environment can predict future performance trends and potential problems in advance by learning from historical data through machine learning models. This gives the operation and maintenance team enough time to prepare response strategies and avoids being caught off guard when problems occur. For example, it can predict that the server may fail due to increased load in the next week and expand resources in advance. Attached Figure Description

[0015] Figure 1 This is a flowchart of the deployment method of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Please see Figure 1 This invention provides a technical solution: a two-dimensional layered method for intelligent operation and maintenance of cloud computing environments, which constructs an automated operation and maintenance framework and divides the framework into two dimensions: a horizontal dimension and a vertical dimension, wherein the vertical dimension specifically includes the following steps: S1. Data Collection: Automatically acquire system and application data using monitoring and logging tools; S2. Data storage: The acquired data is stored in the database; S3. Data Analysis: Real-time data analysis to identify anomalies, using methods such as trend analysis and correlation analysis to proactively discover potential problems and risks; S4. Monitoring and Alerts: Set thresholds to automatically trigger alarms and detect problems; S5. Response and Automated Repair: Notify operations and maintenance personnel and execute automated repair measures; S6. Visualization and Feedback: Provide intuitive data displays and establish a feedback mechanism.

[0018] In this embodiment of the invention, the horizontal dimension is specifically divided into the following levels according to function: 1) Strategy layer: Formulating and implementing strategies and allocating resources; 2) Management level, responsible for managing physical hardware resources, including servers, storage devices, and network devices; 3) Operational layer: Ensure the normal operation of the hardware and perform regular maintenance and upgrades.

[0019] By clearly defining responsibilities, the strategy layer formulates and implements strategies and allocates resources, the management layer is responsible for managing physical hardware resources, and the operations layer ensures the normal operation of the hardware and performs regular maintenance and upgrades. This improves efficiency and professionalism, allowing personnel at each level to focus on their areas of expertise. It also enhances system stability, ensuring that hardware management and maintenance work is carried out systematically. Furthermore, it allows for flexibility in adapting to changes; each layer can independently adjust to respond to business or technological changes, reducing risks, facilitating problem identification and resolution, and promoting collaboration while ensuring smooth information flow.

[0020] In this embodiment of the invention, in S1, monitoring tools are used to collect system and application performance metrics in real time, including CPU utilization, memory usage, disk I / O, and network traffic. In addition, application logs, system logs, and network logs are collected through a log management system to obtain detailed information on the system's operating status. Then, APIs or SDKs are used to obtain operating data from different systems, services, and devices.

[0021] In this embodiment of the invention, the collected data is stored in a high-performance data warehouse in S2 to support subsequent analysis and querying. Before data storage, data cleaning and preprocessing are performed to remove duplicate, erroneous and invalid data.

[0022] Data cleaning removes duplicate, erroneous, and invalid data, ensuring that subsequent analysis is based on accurate and valuable data. This avoids erroneous analytical conclusions caused by incorrect data, improves the reliability of decision-making, and reduces the need for data storage, thereby lowering storage costs.

[0023] In this embodiment of the invention, stream processing technology is used in S3 to analyze real-time data streams, identify abnormal situations or potential faults in a timely manner, and establish a predictive model by applying machine learning algorithms. By learning from historical data, future performance trends and potential problems are predicted, such as anomaly detection models and fault prediction models.

[0024] Machine learning models can predict future performance trends and potential problems by learning from historical data. This gives operations teams enough time to prepare response strategies and avoids chaos when problems occur. For example, they can predict that servers may fail due to increased load in the next week and expand resources in advance.

[0025] In this embodiment of the invention, S4 sets thresholds for key performance indicators. When the system indicators exceed or fall below these thresholds, a warning is triggered. Furthermore, by utilizing a machine learning-based intelligent warning system, false alarms are reduced, and the accuracy and intuitiveness of warnings are improved, enabling maintenance personnel to respond to problems quickly.

[0026] By setting thresholds for key performance indicators to trigger warnings and utilizing a machine learning-based intelligent warning system, abnormal system indicators can be detected in a timely manner, preventing problems from escalating, reducing false alarms, allowing maintenance personnel to focus on the real issues without being distracted, improving the accuracy of warnings, helping them quickly locate key issues, enhancing intuitiveness, facilitating problem understanding, ensuring rapid response by maintenance personnel, and reducing losses.

[0027] In this embodiment of the invention, in step S5, an automated script is written to automatically perform repair operations based on specific warning conditions and events, such as restarting services and adjusting resource allocation. Furthermore, a system with self-healing capabilities is built, which automatically performs self-repair when a fault is detected according to a defined strategy, thereby minimizing the need for manual intervention.

[0028] By writing automated scripts, repair operations can be automatically performed based on specific warning conditions and events, such as restarting services, adjusting resource allocation, and building self-healing systems. These systems automatically repair themselves when faults are detected according to defined strategies, offering significant advantages. They can respond quickly, initiating repairs immediately upon a fault, shortening system recovery time, reducing business interruption, and rapidly restarting faulty services, improving efficiency, saving manual troubleshooting and processing time, and allowing operations and maintenance personnel to focus on more important tasks.

[0029] In this embodiment of the invention, in step S6, the system operation status is visualized through a dashboard, providing an intuitive data display that enables maintenance personnel to quickly understand the system status and establish a feedback mechanism to continuously monitor the repair effect and key indicators, and continuously optimize the data collection, analysis and response process according to the actual situation.

[0030] By visualizing the system's operational status through dashboards, maintenance personnel can quickly understand the status, establish a feedback mechanism to continuously monitor the repair effects and key indicators, enable efficient decision-making, improve problem-solving efficiency, continuously optimize processes, enhance risk management, and optimize resource allocation.

[0031] In summary, by analyzing the impact of passenger experience and wayfinding information on wayfinding behavior at the micro level of hubs, we can effectively reveal the influence of single-point information and flow layout on passenger wayfinding. This can help avoid improper hub layout and signage placement methods that reduce passenger wayfinding experience and efficiency. By comparing similar information to achieve control variable analysis, we can promote the humanized and intelligent design of hub information layout.

[0032] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.

[0033] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A two-dimensional layered method for realizing intelligent operation and maintenance in a cloud computing environment, comprising building an automated operation and maintenance framework and dividing the framework into two dimensions, characterized in that: The horizontal dimension and the vertical dimension, where the vertical dimension specifically includes the following steps: S1. Data Collection: Automatically acquire system and application data using monitoring and logging tools; S2. Data storage: The acquired data is stored in the database; S3. Data Analysis: Analyze data in real time, identify anomalies, and discover potential problems and risks in advance; S4. Monitoring and Alerts: Set thresholds to automatically trigger alarms and detect problems; S5. Response and Automated Repair: Notify operations and maintenance personnel and execute automated repair measures; S6. Visualization and Feedback: Provide data visualization and establish a feedback mechanism.

2. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: The horizontal dimension is specifically divided into the following levels according to function: 1) Strategy layer: Formulating and implementing strategies and allocating resources; 2) Management level, responsible for managing physical hardware resources, including servers, storage devices, and network devices; 3) Operational layer: Ensure the normal operation of the hardware and perform regular maintenance and upgrades.

3. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: In step S1, monitoring tools are used to collect system and application performance metrics in real time, including CPU utilization, memory usage, disk I / O, and network traffic. In addition, application logs, system logs, and network logs are collected through a log management system to obtain detailed information on the system's operating status. Then, APIs or SDKs are used to obtain operating data from the system.

4. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: In step S2, the collected data is stored in a data warehouse to support subsequent analysis and queries. Before data storage, data cleaning and preprocessing are performed to remove duplicate, erroneous, and invalid data.

5. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: In S3, data stream processing technology is used to analyze data in real time, identify abnormal situations or potential faults, and a predictive model is built by applying machine learning algorithms. By learning from historical data, future performance trends and potential problems are predicted.

6. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: The S4 sets threshold indicators, triggers warnings when system indicators exceed or fall below these thresholds, and uses intelligent monitoring tools to learn from false alarms.

7. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: The S5 program writes automated scripts that automatically perform repair operations when an event occurs, and builds a self-healing system that automatically repairs itself when a fault is detected.

8. The two-dimensional layered method for realizing intelligent operation and maintenance of a cloud computing environment according to claim 1, characterized in that: The S6 uses a dashboard to visualize the system's operation, providing data display and monitoring the repair effect.