Cloud platform operation and maintenance method and system, electronic equipment and storage medium
Through reinforcement learning and multi-model analysis, data collection strategies are dynamically adjusted to generate multi-objective optimization decisions, solving the problems of time-consuming fault location and low resource utilization in hybrid cloud architecture operation and maintenance, and achieving efficient and stable cloud platform operation and maintenance.
Patent Information
- Application Number
- CN202510776814.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies in hybrid cloud architecture operations and maintenance take a long time to locate faults and have low resource prediction accuracy, making it difficult to meet the needs of dynamically changing business operations, resulting in low operation and maintenance efficiency and low resource utilization.
Through the reinforcement learning model, the data sampling frequency and the number of concurrent requests are dynamically adjusted. Deep learning, pre-trained language models, graph neural networks and evolutionary algorithms are combined to generate multi-objective optimization decisions, realize global optimization decisions and closed-loop operation and maintenance, and use the distributed gradient enhancement library to build a machine learning model for data storage management.
It improves operation and maintenance efficiency, shortens fault location time, improves resource utilization, ensures the stability and reliability of the cloud platform, and forms an efficient closed-loop operation and maintenance process.
Smart Images

Figure CN120670205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-cloud operation and maintenance technology, and in particular to a cloud platform operation and maintenance method, system, electronic device and storage medium. Background Art
[0002] Currently, most enterprises' big data application service platforms are built based on cloud resources provided by various Internet cloud vendors. Due to considerations such as cost control and disaster recovery, the architectural design is mostly reflected in the form of hybrid cloud architecture.
[0003] For hybrid cloud architecture operations and maintenance, existing fault location technology relies on single-modal analysis, resulting in a cumbersome process. It takes an average of 35 minutes from fault discovery to intervention, which is inefficient. Resource prediction accuracy is low, making it impossible to accurately predict resource needs and meet the needs of dynamically changing business operations, limiting the foresight and proactive nature of operations and maintenance. Summary of the Invention
[0004] The purpose of the present invention is to provide a cloud platform operation and maintenance method, system, electronic device and storage medium, which dynamically collect multimodal data through reinforcement learning, fuse multi-model analysis to generate optimization decisions and perform repairs, realize global optimization decision-making and closed-loop operation and maintenance, and improve efficiency and stability.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention discloses a cloud platform operation and maintenance method, which includes:
[0007] Collect multimodal data from several cloud platforms;
[0008] A multi-objective optimization decision is generated by analyzing a decision model, which includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; the evolutionary algorithm model combines resource demand, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision;
[0009] Perform repair operations based on multi-objective optimization decisions.
[0010] Furthermore, after collecting multimodal data from several cloud platforms, the data sampling frequency and the number of concurrent requests are dynamically adjusted based on the reinforcement learning model.
[0011] Furthermore, the method further includes: obtaining a strategy execution result after executing the repair operation according to the multi-objective optimization decision;
[0012] Use successful strategy execution results as positive samples to enhance and optimize analytical decision-making models;
[0013] The failed strategy execution results are used as negative samples to fine-tune the parameters of the analysis and decision model.
[0014] Furthermore, the log data of the repair operation is divided into audit storage data and direct storage data according to preset rules;
[0015] Combining Drools rules with the isolation forest algorithm, it identifies whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored.
[0016] Furthermore, the data to be stored are stored in the first area by category, wherein the data to be stored include the collected multimodal data, the directly stored data in the log data of the repair operation, and the audit storage data after identification;
[0017] A distributed gradient boosting library is used to build a machine learning model to dynamically evaluate the value of the data to be stored, and the data to be stored is stored in layers in the second area based on the value evaluation results.
[0018] Furthermore, the loss function of the machine learning model is
[0019]
[0020] Where N is the number of sample data, y i is the true value evaluation result of the i-th sample data according to the scoring formula, is the prediction value evaluation result of the i-th sample data, ω i is the weight coefficient of the i-th sample data.
[0021] Furthermore, by configuring a unified resource adapter collector, access adaptation of various resources on multiple different cloud platforms can be achieved to collect multimodal data from several cloud platforms; based on the preset rule library, the field names of heterogeneous indicators on different cloud platforms are uniformly converted into standardized names, and the multimodal data are standardized through preset formulas; then the standardized multimodal data is stored in the stream processing pipeline for downstream analysis and decision-making models to obtain and process the data in real time.
[0022] In a second aspect, the present invention discloses a cloud platform operation and maintenance system, which includes:
[0023] Data acquisition module, used to collect multimodal data from several cloud platforms;
[0024] An analysis and decision-making module is used to generate multi-objective optimization decisions through an analysis and decision-making model, which includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; the evolutionary algorithm model combines resource requirements, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision;
[0025] The operation and maintenance management module is used to perform repair operations based on multi-objective optimization decisions.
[0026] Furthermore, it also includes a log audit module and a data storage module. The log audit module is used to divide the log data of the repair operation into audit storage data and direct storage data according to preset rules; combining Drools rules and the isolation forest algorithm to identify whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored; the data storage module is used to store the data to be stored in the first area according to category, and the data to be stored includes the collected multimodal data, and the directly stored data and the identified audit storage data in the log data of the repair operation; and a distributed gradient enhancement library is used to build a machine learning model for dynamically evaluating the value of the data to be stored, and storing the data to be stored in layers in the second area according to the value evaluation results.
[0027] In a third aspect, the present invention discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned cloud platform operation and maintenance method when executed by a processor.
[0028] The present invention has the following unexpected beneficial effects:
[0029] 1. The cloud platform operation and maintenance method described in this invention utilizes a reinforcement learning model to dynamically adjust data sampling frequency and concurrent request counts based on the cloud platform's real-time load and data variations. This intelligently optimizes data collection strategies, avoiding resource waste caused by over-collection and the omission of critical information due to insufficient sampling. This ensures that collected data is both accurate and efficient, providing a reliable data foundation for subsequent analysis and decision-making. A deep learning model predicts future resource demand ranges, helping operation and maintenance personnel plan resource allocation in advance, avoiding service delays caused by insufficient resources or cost waste caused by redundant resources, thereby improving resource utilization. A pre-trained language model extracts entity and semantic features from log alarm data, rapidly matching historical fault databases to generate a recommended list of similar cases, significantly shortening fault location time and reducing troubleshooting difficulty for operation and maintenance personnel. A graph neural network model quantifies service dependency weights and calculates fault propagation probability, enabling the prediction of fault spread trends and enabling operation and maintenance personnel to proactively implement isolation or protection measures to reduce the scope of fault impact. An evolutionary algorithm model is then employed to integrate multi-source information to generate multi-objective optimization decisions, balancing multiple objectives such as resource allocation and fault remediation, achieving global optimization for operation and maintenance decisions and improving the overall stability and reliability of the cloud platform.
[0030] 2. The cloud platform operation and maintenance method described in the present invention quickly executes repair operations based on multi-objective optimization decisions, forming a closed-loop operation and maintenance process of "data collection-analysis and decision-making-repair execution", which greatly shortens the time for handling operation and maintenance problems, improves operation and maintenance efficiency, reduces business interruption time, and ensures cloud platform service continuity and user experience.
[0031] 3. The cloud platform operation and maintenance system of the present invention adopts a modular design, which facilitates the subsequent upgrade and optimization of individual modules or the introduction of new modules to adapt to the continuous changes in cloud platform operation and maintenance scenarios and technological evolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the implementation or prior art description. Obviously, the drawings described below are only some embodiments of the present invention.
[0033] Figure 1 A flow chart of the cloud platform operation and maintenance method according to an embodiment of the present invention is shown.
[0034] Figure 2 A structural diagram of an implementation scheme of a cloud platform operation and maintenance system according to an embodiment of the present invention is shown.
[0035] Figure 3 A structural diagram of another implementation of the cloud platform operation and maintenance system according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0036] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0037] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0038] In one embodiment, the present invention discloses a cloud platform operation and maintenance method, see Figure 1 As shown, the method includes:
[0039] Collect multimodal data from several cloud platforms;
[0040] A multi-objective optimization decision is generated by analyzing a decision model, which includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; the evolutionary algorithm model combines resource demand, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision;
[0041] Perform repair operations based on multi-objective optimization decisions.
[0042] The cloud platform operation and maintenance method described in the present invention generates multi-objective optimization decisions by analyzing the decision model. Specifically, it predicts the future resource demand range through a deep learning model, helping operation and maintenance personnel to plan resource allocation in advance, avoiding business stalls due to insufficient resources, or cost waste caused by resource redundancy, and improving resource utilization. The pre-trained language model extracts entity and semantic features from the log alarm data, quickly matches the historical fault library to generate a recommendation list of similar cases, greatly shortening the fault location time and reducing the difficulty of troubleshooting for operation and maintenance personnel. The graph neural network model is used to quantify the service dependency weights and calculate the fault propagation probability, which can predict the fault diffusion trend and help operation and maintenance personnel take isolation or protection measures in advance to reduce the scope of fault impact. The evolutionary algorithm model is used to integrate multi-source information to generate multi-objective optimization decisions, balance multiple goals such as resource allocation and fault repair, achieve the global optimization of operation and maintenance decisions, and improve the overall stability and reliability of the cloud platform.
[0043] The cloud platform operation and maintenance method described in the present invention quickly executes repair operations based on multi-objective optimization decisions, forming a closed-loop operation and maintenance process of "data collection-analysis and decision-making-repair execution", which greatly shortens the time for handling operation and maintenance problems, improves operation and maintenance efficiency, reduces business interruption time, and ensures cloud platform service continuity and user experience.
[0044] As a preferred embodiment of the present invention, after collecting multimodal data from several cloud platforms, the data sampling frequency and the number of concurrent requests are dynamically adjusted based on the reinforcement learning model.
[0045] Existing technologies lack dynamic, intelligent throttling designs, resulting in time-consuming analysis and processing of operational issues, a large daily workload, low platform resource utilization, and difficulty in flexibly allocating resources based on actual needs, leading to resource waste or shortages. The cloud platform operation and maintenance method described in the present invention utilizes a reinforcement learning model to dynamically adjust data sampling frequency and the number of concurrent requests based on the cloud platform's real-time load and data variation characteristics, intelligently optimizing data collection strategies. This avoids resource waste caused by excessive collection and prevents the omission of key information due to insufficient sampling, ensuring that the collected data is both accurate and efficient, providing a reliable data foundation for subsequent analysis and decision-making.
[0046] Exemplarily, the state space of the reinforcement learning model includes basic indicators (such as API quotas, average response delays, etc.) and dynamic features (such as real-time sampling frequency, real-time concurrent request numbers), and the constraints include: hard thresholds (current limiting times, etc.) adjusted by the reference strategy and soft prediction results based on LSTM (quota usage trends in the next 5 minutes, etc.). For the adjustment of the real-time acquisition frequency, the processing methods in the action space include adjusting the sampling frequency according to discrete actions of 1 second, 5 seconds or 20 seconds, or continuously adjusting the sampling frequency in the range of 1 second to 60 seconds; for the adjustment of the number of concurrent requests, the processing methods in the action space include directly outputting the optimized value of the number of concurrent requests, and adjusting the number of real-time concurrent requests based on the optimized value. By dynamically adjusting the sampling frequency and the number of concurrent requests, intelligent throttling driven by the reinforcement learning model is realized, which solves the difficulty that traditional flow control based on fixed rules (such as the upper limit of the number of requests per second) cannot adapt to dynamic load fluctuations. The multi-objective optimization reward function of the reinforcement learning model is defined as:
[0047] Where S success is the success rate of API calls, is the first weight coefficient to ensure stability; R error is the request error rate, which is the ratio of the number of request failures to the total number of requests. is the second weight coefficient to suppress abnormal requests; U quota is the quota utilization rate, It is the third weight coefficient to control the usage amount.
[0048] For example,
[0049] As a preferred embodiment of the present invention, the cloud platform operation and maintenance method described in the present invention also includes: obtaining the strategy execution result after executing the repair operation based on the multi-objective optimization decision; using the successful strategy execution result as a positive sample to enhance the optimization analysis decision model; using the failed strategy execution result as a negative sample to fine-tune the parameters of the analysis decision model.
[0050] This preferred implementation collects the results of strategy execution and converts them into positive and negative samples, forming a closed-loop evolutionary mechanism of "decision-making-execution-feedback-optimization". Successful strategies serve as positive samples, which can strengthen the model's memory of effective decision-making patterns. For example, when the resource expansion strategy generated by the evolutionary algorithm successfully alleviates the peak pressure of the business, the model will increase the weight of such strategies. Failed strategies serve as negative samples, and decision deviations can be corrected through parameter fine-tuning. For example, when the fault isolation strategy fails to effectively block the propagation, the graph neural network will recalibrate the service dependency weights to avoid the recurrence of similar errors. This mechanism breaks the limitations of traditional static models, enabling analytical decision-making models to continuously evolve as business scenarios change, and adapt to the dynamic and complex operating environment of the cloud platform.
[0051] Using real execution results as training data avoids the subjectivity and lag inherent in manual optimization. For example, traditional O&M models rely on experts to manually adjust neural network parameters. This approach, however, automatically triggers a gradient descent algorithm based on failure samples, accurately pinpointing parameter deviations. Furthermore, the accumulation of positive examples accelerates the model's learning of optimal decision paths.
[0052] For example, the actual SLA, cost, abnormal recovery time and other policy execution results are recorded, and manual correction marks are made (such as the expansion quantity adjusted by the operation and maintenance personnel). The LoRA (Low-Rank Adaptation) technology is used to update only some parameters of the analysis and decision-making model for low-cost fine-tuning. Incremental training is performed once an hour, and the full model is updated weekly. The policy weights are dynamically adjusted based on the reinforcement learning reward function.
[0053] Specifically, the reinforcement learning reward function is defined as R = 0.7 × SLA compliance rate - 0.3 × cost overrun rate. The strategy library is maintained to dynamically eliminate inefficient strategies. Inefficient strategies are automatically removed if they fail to meet the standards for three consecutive times.
[0054] The existing technology operation and maintenance strategy library has slow iterations and an update cycle of up to weeks, making it difficult to quickly respond to complex and changing operation and maintenance environments and business needs. At the same time, it is impossible to keep up with the development trend of large models to improve the algorithm in a timely manner, resulting in operation and maintenance technology lagging behind the development speed of the industry. Exemplarily, the deep learning model described in the present invention is a Transformer-based time series prediction model, which is used to capture the long-term dependencies of indicators such as CPU and memory, and output a resource demand prediction range including a confidence interval. For example, based on the time series data of historical resource indicators (CPU, memory, network, etc.) at the input end, combined with the overall status of the current big data cluster task, a Transformer time series prediction model with sparse attention is used to predict the resource demand output confidence interval for the next hour (such as the CPU utilization prediction value ±5%).
[0055] The pre-trained language model is DeBERTa, which is used to parse the semantics of alarm logs, such as "Kafka cluster disk usage exceeds 90%." The DeBERTa model extracts entities (resource type, indicator name, value) and semantic features (urgency, related services), matches the historical fault library, and generates a recommendation list of similar cases.
[0056] The graph neural network model is Graph Transformer, which quantifies the service dependency weights and calculates the fault propagation probability.
[0057] The evolutionary algorithm model is the NSGA-II algorithm, which combines resource requirements, a list of recommended similar cases, and the probability of fault propagation to generate a Pareto optimal solution set, and selects the final strategy based on business priorities, thus generating a multi-objective optimization decision.
[0058] Furthermore, executing repair operations based on multi-objective optimization decisions specifically involves pushing the multi-objective optimization decisions to designated personnel (i.e., operations personnel with appropriate authority in the corresponding field), and executing repair operations based on the operational instructions entered by the operations personnel. By pre-defining a variety of standard operations, such as expansion, migration, restart, and configuration changes, and confirming the operational instructions, the system automatically calls the uniformly packaged APIs of various cloud platforms to execute repair operations, such as EMR node expansion, component configuration parameter updates, and node restarts. Health checks are also performed after each stage of the operation.
[0059] As a preferred embodiment of the present invention, the log data of the repair operation is divided into audit storage data and direct storage data according to preset rules; combining Drools rules with the isolation forest algorithm, it is identified whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored.
[0060] This preferred implementation divides the log data of repair operations into audit storage data and direct storage data through preset rules, thereby achieving differentiated data management. Direct storage data covers routine operation and maintenance operation records, with fast storage as the main method to ensure data writing efficiency; audit storage data focuses on key operations (such as permission changes, core service restarts), and performs more rigorous processing and storage. This classification method reduces redundant storage and reduces storage costs, while enabling operation and maintenance personnel to quickly locate the required information based on the data type, improving retrieval efficiency. For example, when troubleshooting security incidents, relevant operation records can be retrieved directly from the audit storage data without traversing massive logs, shortening the time to locate the problem.
[0061] The operation and maintenance procedures of existing technologies lack a complete traceability mechanism. Manual auditing is arduous and time-consuming, making it difficult to achieve transparent and standardized management of operation and maintenance operations. There are compliance risks, which also increase the management costs and potential risks of enterprises. This preferred implementation combines Drools rules with the isolation forest algorithm to identify anomalies in audit storage data and builds a multi-level risk detection system. Drools rules can quickly capture abnormal behaviors of known patterns based on preset logic in operation and maintenance scenarios (such as "deletion operations on multiple core databases by the same account in a short period of time" is an anomaly); the isolation forest algorithm uses machine learning to automatically discover outliers in the data and mine potential unknown anomalies. The combination of the two enables comprehensive monitoring of audit storage data.
[0062] Identifying and alerting abnormalities in audited storage data provides strict security auditing for cloud platform operations. When an abnormal operation is identified, the generated alert not only notifies operations personnel to address it promptly but also serves as audit evidence, facilitating tracing the source of the operation and clarifying accountability. This helps standardize operations procedures and ensures compliance with enterprise security policies and compliance requirements. For example, if abnormal operations such as batch resource creation during off-hours are identified, high-risk logs trigger corporate WeChat and email notifications for the corresponding personnel, providing contextual information.
[0063] The coordinated operation of log data classification and anomaly detection mechanisms ensures the rational allocation of cloud platform resources. Rapid processing of directly stored data reduces system resource usage and ensures the operational performance of cloud platform services. Focused detection and anomaly handling of audited stored data strengthens security without impacting business performance.
[0064] It's important to note that as log data accumulates during anomaly identification, Drools rules and the Isolation Forest algorithm model can be further optimized. For example, the rule base can be adjusted based on newly discovered anomaly patterns, and the Isolation Forest algorithm can be retrained using newly discovered anomaly samples to more accurately identify abnormal behavior. This data-driven model optimization mechanism continuously improves anomaly detection capabilities over time, enabling it to better cope with the complex and ever-changing operating environment of cloud platforms and increasingly sophisticated security threats, forming a virtuous cycle of "detection-optimization-redetection."
[0065] As a preferred embodiment of the present invention, the data to be stored are stored in the first area according to categories, and the data to be stored include the collected multimodal data, the directly stored data in the log data of the repair operation, and the audit storage data after identification; a distributed gradient boosting library is used to construct a machine learning model for dynamically evaluating the value of the data to be stored, and the data to be stored are stored in layers in the second area according to the value evaluation results.
[0066] This preferred implementation adopts a multi-AZ (Availability Zone) multi-copy strategy, combined with a hierarchical storage architecture, to provide differentiated protection for data of different values. High-value real-time time series data and critical log data can be stored synchronously in multiple availability zones through a multi-copy strategy to ensure that the data is not lost and can be accessed at any time; although historical archived data is of lower value, the reliability of long-term storage can still be guaranteed through a copy strategy. This strategy works in conjunction with a dynamic value assessment mechanism to prioritize the disaster recovery capabilities of core data when resources are limited, while reducing overall disaster recovery costs and improving the resilience of the cloud platform to extreme situations such as natural disasters and hardware failures.
[0067] This preferred implementation stores the data to be stored in the first area by category, that is, adopting a differentiated storage solution. For example: time series data relies on the InfluxDB cluster, and is stored by time slicing combined with the Gorilla compression algorithm to ensure that monitoring indicators collected at high frequency can be efficiently written and queried; log data is passed through the Elasticsearch cluster, using the inverted index and BM25 algorithm to achieve fast retrieval, meeting the needs of instant analysis of massive logs during troubleshooting; historical archive data is stored in object storage (such as AWS S3, Alibaba Cloud OSS, Huawei Cloud OBS, Tencent Cloud COS, etc.), and is compressed and stored in Parquet columnar format, taking into account long-term data preservation and space occupancy.
[0068] Traditional hot and cold data stratification relies on fixed rules, which often leads to high-performance storage media being occupied by low-value data. However, the present invention uses a machine learning model built through a distributed gradient boosting library to analyze the access frequency of data in real time. access 、Business Importance business , error value f comliance Cost f cost The multi-dimensional characteristics of data are used to dynamically quantify the value of data. Specifically, the value score V is calculated as follows:
[0069] V=α·f access +β·f business +γ·f comliance -λ·f cost , α, β, γ, and λ are the weight coefficients corresponding to access frequency, business importance, error value, and cost, respectively.
[0070] For example, time series data with a value score greater than 0.8 enters the memory optimization layer, using time series data storage + SSD local disk storage media. At the same time, based on the actual distribution of the current value-weighted data volume, real-time monitoring and scheduling of resources are performed (for example, if high-value data is greater than 50%, refer to the resource expansion process) to achieve deep integration between models.
[0071] Furthermore, a query parser converts SQL queries into the native query language of each storage engine, enabling cross-engine data retrieval. This, combined with data tiered storage and value-oriented strategies, significantly improves data acquisition efficiency. For example, when dealing with sudden outages, operations and maintenance personnel can quickly retrieve real-time time series data from InfluxDB, log records from Elasticsearch, and historical audit data from object storage using unified SQL statements, eliminating the need to switch operations between multiple platforms. High-value data is stored at an easily accessible level, further shortening data retrieval time. Compared to traditional unordered storage and multi-system independent query modes, this can improve problem location efficiency and effectively ensure cloud platform service continuity.
[0072] Furthermore, the machine learning model's assessment of data value is not only used to adjust storage policies, but also provides in-depth support for overall cloud platform operation and maintenance decisions. By analyzing the value distribution of data in different storage engines, key business links and potential risks can be identified. For example, if the value of a certain type of log data in an Elasticsearch cluster continues to rise, it may indicate frequent anomalies in the corresponding business module and require close monitoring. If historical data in object storage is of extremely low value and occupies a large amount of space, automatic archiving or cleanup mechanisms can be triggered. This intelligent management approach, which deeply integrates data storage with business operations, drives operations and maintenance from passive response to active optimization, improving the refined operations of the cloud platform.
[0073] Furthermore, machine learning models built using a distributed gradient boosting library are highly scalable, seamlessly adapting to upgrades and business changes in multimodal data storage architectures. Whether it's new time series metrics and log types generated by newly added business modules, or new data formats introduced in object storage, the model can quickly learn and update its evaluation logic. The multi-engine storage architecture can also be flexibly expanded or replaced based on business needs (e.g., switching from InfluxDB to other time series databases). This feature ensures that the cloud platform's data storage management remains efficient and stable despite technological iterations and business expansion, avoiding operational bottlenecks caused by rigid storage architectures.
[0074] As a preferred embodiment of the present invention, the loss function of the machine learning model is
[0075]
[0076] Where N is the number of sample data, y i is the true value evaluation result of the i-th sample data according to the scoring formula, is the prediction value evaluation result of the i-th sample data, ω i is the weight coefficient of the i-th sample data.
[0077] The loss function evaluates the result y based on the true value of the sample i To set the weight coefficient ω i When y i > threshold, ω i =3, which means that for high values (here y i >Threshold definition. For example, when the threshold is set at 0.7, the machine learning model will give greater weight to data samples during training. This means that the machine learning model will pay more attention to the prediction errors of these high-value data samples. By increasing the penalty for these prediction deviations, the machine learning model will be encouraged to more accurately learn and predict the value assessment results of high-value data, thereby improving the accuracy of high-value data value assessment and ensuring that important data is properly treated and stored.
[0078] For y i Data samples ≤ threshold, weight coefficient ω i = 1, giving lower weight to high-value data samples. This setting doesn't ignore low-value data; rather, it balances the impact of data samples of varying value on model training while ensuring the accuracy of high-value data evaluation. This prevents the machine learning model from focusing too much on prediction errors from low-value data, which can affect learning and prediction performance on high-value data. Overall, it optimizes the machine learning model's ability to evaluate data of varying value, allowing it to more rationally allocate learning resources.
[0079] This loss function is designed to meet the actual business needs for data value assessment. In business scenarios such as cloud platform data storage, high-value data, such as key business metrics and important system logs, often has a more critical impact on business decisions and system stability. Using this loss function, machine learning models can more effectively capture the characteristics of high-value data, providing a more reliable basis for subsequent operations such as data tiered storage based on value assessment results. This allows storage resources to be more rationally allocated to high-value data, improving overall business operational efficiency and data management.
[0080] As a preferred embodiment of the present invention, by configuring a unified resource adapter collector, access adaptation of various resources on multiple different cloud platforms is achieved to collect multimodal data from several cloud platforms; based on the preset rule library, the field names of heterogeneous indicators of different cloud platforms are uniformly converted into standardized names, and the multimodal data are standardized through preset formulas; then the standardized multimodal data is stored in the stream processing pipeline for downstream analysis and decision models to acquire and process the data in real time.
[0081] This preferred implementation effectively addresses the difficulty of accessing resources from different cloud platforms by configuring a unified resource adapter collector. Different cloud platforms differ in interface specifications and data formats. The resource adapter collector acts like a universal converter, adapting and integrating various resources across multiple cloud platforms. This setup allows resources from platforms like AWS, Alibaba Cloud, and Tencent Cloud to be uniformly collected, breaking down data barriers between cloud platforms and enabling multimodal data aggregation across cloud platforms, providing a comprehensive data foundation for subsequent data processing and analysis.
[0082] Based on a pre-set rule base, the field names of heterogeneous indicators across different cloud platforms are uniformly converted to standardized names, and multimodal data is standardized, significantly improving data standardization. In an environment where multiple cloud platforms coexist, indicator naming and data formats often vary, making data analysis and integration extremely difficult. Through standardization, the previously disorganized heterogeneous data becomes organized, enabling data from different cloud platforms to be compared and analyzed under the same standards. This improves data usability and understandability, while reducing the difficulty and cost of data processing.
[0083] Standardized multimodal data is stored in a stream processing pipeline, allowing downstream analytical decision-making models to access and process the data in real time. This real-time data flow mechanism enables analytical decision-making models to obtain the latest data promptly. Data timeliness is crucial in scenarios such as cloud platform management, such as real-time monitoring of resource usage and performance metrics on cloud platforms. Through the stream processing pipeline, analytical decision-making models can analyze data in real time, promptly identify potential issues, and make decisions, such as automatically adjusting resource allocation and providing fault warnings. This improves the management efficiency and stability of the cloud platform and ensures the continued stable operation of the business.
[0084] In one embodiment, the present invention discloses a cloud platform operation and maintenance system, see Figure 2 As shown, the system includes a data acquisition module, an analysis and decision-making module, and an operation and maintenance management module.
[0085] The data acquisition module is used to collect multimodal data from several cloud platforms.
[0086] The analysis and decision-making module is used to generate multi-objective optimization decisions through an analysis and decision-making model. The analysis and decision-making model includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; and the evolutionary algorithm model combines resource demand, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision.
[0087] The operation and maintenance management module is used to perform repair operations based on multi-objective optimization decisions.
[0088] As a preferred embodiment of the present invention, see Figure 3 As shown, the cloud platform operation and maintenance system also includes a log audit module, which is used to divide the log data of the repair operation into audit storage data and direct storage data according to preset rules; combining Drools rules and the isolation forest algorithm to identify whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored.
[0089] As a preferred embodiment of the present invention, see Figure 3 As shown, the cloud platform operation and maintenance system also includes a data storage module for storing the data to be stored in the first area according to categories. The data to be stored includes collected multimodal data, directly stored data in the log data of the repair operation, and audit storage data after identification; and a distributed gradient enhancement library is used to build a machine learning model for dynamically evaluating the value of the data to be stored, and storing the data to be stored in layers in the second area according to the value evaluation results.
[0090] In one embodiment, the present invention discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned cloud platform operation and maintenance method when executed by a processor.
[0091] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or modification made by those skilled in the art based on the present invention is within the protection scope of the present invention.
Claims
1. A cloud platform operation and maintenance method, characterized in that: include: Collect multimodal data from several cloud platforms; A multi-objective optimization decision is generated by analyzing a decision model, which includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; the evolutionary algorithm model combines resource demand, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision; Perform repair operations based on multi-objective optimization decisions.
2. The cloud platform operation and maintenance method according to claim 1, characterized in that: After collecting multimodal data from several cloud platforms, the data sampling frequency and number of concurrent requests are dynamically adjusted based on the reinforcement learning model.
3. The cloud platform operation and maintenance method according to claim 1, characterized in that: Also includes: After executing the repair operation based on the multi-objective optimization decision, obtain the strategy execution results; Use successful strategy execution results as positive samples to enhance and optimize analytical decision-making models; The failed strategy execution results are used as negative samples to fine-tune the parameters of the analysis and decision model.
4. The cloud platform operation and maintenance method according to claim 1, characterized in that: Divide the log data of the repair operation into audit storage data and direct storage data according to preset rules; Combining Drools rules with the isolation forest algorithm, it identifies whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored.
5. The cloud platform operation and maintenance method according to claim 4, characterized in that: Storing the data to be stored in the first area by category, the data to be stored including the collected multimodal data, directly stored data in the log data of the repair operation, and the audit storage data after identification; A distributed gradient boosting library is used to build a machine learning model to dynamically evaluate the value of the data to be stored, and the data to be stored is stored in layers in the second area based on the value evaluation results.
6. The cloud platform operation and maintenance method according to claim 5, characterized in that: The loss function of the machine learning model is Where N is the number of sample data, y i is the true value evaluation result of the i-th sample data according to the scoring formula, is the prediction value evaluation result of the i-th sample data, ω i is the weight coefficient of the i-th sample data.
7. The cloud platform operation and maintenance method according to claim 1, characterized in that: By configuring a unified resource adapter collector, we can achieve access adaptation of various resources on multiple cloud platforms and collect multimodal data from several cloud platforms. Based on the preset rule base, the field names of heterogeneous indicators on different cloud platforms are uniformly converted into standardized names, and multimodal data are standardized using preset formulas; the standardized multimodal data is then stored in the stream processing pipeline for downstream analysis and decision-making models to acquire and process the data in real time.
8. A cloud platform operation and maintenance system, characterized in that: include: Data acquisition module, used to collect multimodal data from several cloud platforms; An analysis and decision-making module is used to generate multi-objective optimization decisions through an analysis and decision-making model, which includes a deep learning model, a pre-trained language model, a neural network model, and an evolutionary algorithm model. The deep learning model is used to predict the resource demand range for a preset time period in the future; the pre-trained language model is used to extract entity and semantic features from the log alarm data of multimodal data, and a similar case recommendation list is generated by matching the historical fault library; the graph neural network model is used to quantify the service dependency weights in the log alarm data and calculate the fault propagation probability; the evolutionary algorithm model combines resource requirements, the similar case recommendation list, and the fault propagation probability to generate a multi-objective optimization decision; The operation and maintenance management module is used to perform repair operations based on multi-objective optimization decisions.
9. The cloud platform operation and maintenance system according to claim 8, characterized in that: The system also includes a log audit module and a data storage module. The log audit module is used to divide the log data of the repair operation into audit storage data and direct storage data according to preset rules; combine Drools rules with the isolation forest algorithm to identify whether the audit storage data is abnormal. If the audit storage data is identified as abnormal, a prompt message is generated and stored; if the audit storage data is identified as normal, it is directly stored; The data storage module is used to store the data to be stored in the first area according to categories, wherein the data to be stored includes the collected multimodal data, the directly stored data in the log data of the repair operation, and the audit storage data after identification; A distributed gradient boosting library is used to build a machine learning model to dynamically evaluate the value of the data to be stored, and the data to be stored is stored in layers in the second area based on the value evaluation results.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the cloud platform operation and maintenance method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Cross-platform data security processing and intelligent operation and maintenance management system based on multi-cloud service
CN121334151A
Multi-level data storage and mixed query method and system, storage medium and equipment
CN121542305A