Multi-cloud management method, system and device and computer readable storage medium

By collecting multi-cloud data and using algorithms such as deep learning, reinforcement learning and random forests to formulate management strategies, the efficiency and accuracy of resource management in multi-cloud environments are solved, and efficient and accurate multi-cloud management is achieved.

CN120144308APending Publication Date: 2025-06-13山东浪潮数据库技术有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510300728.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing multi-cloud management methods are not efficient, accurate and adaptable enough, and it is difficult to deal with dynamically changing resource requirements and failures in real time.

Method used

By collecting multi-cloud data from multi-cloud environments, using deep learning algorithms to analyze resources, security, performance and cost, combining reinforcement learning algorithms and random forests to formulate resource management strategies and security management strategies, and using multi-objective optimization functions and support vector machines for dynamic resource management.

Benefits of technology

Accurate analysis of resources, costs, performance and security in multi-cloud environments is achieved, and management strategies are formulated and resource allocation is dynamically adjusted through AI intelligent driving to ensure the efficient operation of business applications, improve resource utilization, and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144308A_ABST
    Figure CN120144308A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-cloud management method, system and device and a computer readable storage medium, and is applied to the technical field of cloud computing, and the method comprises the steps: collecting multi-cloud data from a multi-cloud environment, the multi-cloud data comprising resource data, security data, configuration and operation state data and cost data; based on the multi-cloud data, utilizing a deep learning algorithm to analyze resources, security, performance and cost to obtain various analysis conditions; making a resource management strategy and a safety management strategy by using a reinforcement learning algorithm and a random forest based on each analysis condition; and based on the resource management strategy and the security management strategy, dynamically managing the computing resources, the network resources and the storage resources by using a multi-objective optimization function and a support vector machine. Various algorithms (such as a deep learning algorithm, a reinforcement learning algorithm, a random forest and a support vector machine) are utilized to realize efficient management of resources in the multi-cloud environment, the resource utilization rate is improved, and resource waste is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and particularly relates to a multi-cloud management method, system, device and computer-readable storage medium. Background Art

[0002] A multi-cloud environment refers to an enterprise using cloud services from multiple cloud service providers, or simultaneously using multiple cloud architecture modes such as public clouds, private clouds, and hybrid clouds. The multi-cloud environment brings many advantages to enterprises, but at the same time, it also brings a series of management problems. Current multi-cloud management methods have certain limitations. For example, traditional multi-cloud management methods mainly rely on manual configuration and rule-based automated scripts. For complex multi-cloud environments, manual configuration is prone to errors and cannot respond in real time to dynamic resource requirements and fault situations. Although rule-based automated scripts can improve efficiency to a certain extent, due to the lack of learning and adaptive capabilities, it is also difficult to cope with the changing conditions and complex scenarios in the multi-cloud environment.

[0003] Therefore, how to provide an efficient, accurate, and adaptive multi-cloud management method is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a multi-cloud management method, system, device and computer-readable storage medium, which solves the problem that the existing multi-cloud management method is not efficient, accurate and adaptive enough.

[0005] To solve the above technical problem, the present invention provides a multi-cloud management method, including:

[0006] Collect multi-cloud data from the multi-cloud environment, where the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data;

[0007] Based on the multi-cloud data, use deep learning algorithms to analyze resources, security, performance, and cost respectively to obtain various analysis situations;

[0008] Based on the various analysis situations, use reinforcement learning algorithms and random forests to formulate resource management strategies and security management strategies;

[0009] Based on the resource management strategy and security management strategy, use multi-objective optimization functions and support vector machines to dynamically manage computing resources, network resources, and storage resources respectively.

[0010] Optionally, based on the multi-cloud data, using deep learning algorithms to analyze resources, security, performance, and cost respectively to obtain various analysis situations, including:

[0011] Based on the above-mentioned resource data, use a long short-term memory network model for resource analysis to obtain the resource usage analysis situation; the architecture of the long short-term memory network model includes an input layer, a hidden layer, and an output layer; the input features at each time step in the input layer correspond to the resource usage metrics at a specific moment, and the hidden layer includes multiple input gates, forget gates, and output gates; the early stopping method is introduced in the training process of the long short-term memory network model, and the Adam optimizer is used to adjust the learning rate;

[0012] Based on the above-mentioned security data, use an autoencoder to obtain the reconstructed data, and input the security data and the reconstructed data into a discriminator for security analysis to obtain the security analysis situation;

[0013] Organize the configuration and running status data into transactions, scan all the transactions to count the occurrence frequencies of each item, form frequent item sets with the items that meet the minimum support threshold according to the frequencies, obtain the correlation between cloud services according to the frequent item sets, and obtain the performance analysis situation according to the correlation;

[0014] Based on the above-mentioned cost data, use a linear regression model for cost analysis to obtain the cost analysis situation; the linear regression model includes a computing cost linear regression model, a storage cost linear regression model, and a network cost linear regression model.

[0015] Optionally, based on the above-mentioned various analysis situations, use a reinforcement learning algorithm and a random forest to formulate a resource management strategy and a security management strategy, including:

[0016] Abstract the multi-cloud environment as a Markov decision process, introduce a neural network, and obtain a resource allocation strategy by using the neural network to approximate the Q-value function; the state space in the Markov decision process includes the resource usage analysis situation and the cost analysis situation of each cloud service, the action space is resource allocation and scheduling operations, and the reward function is the improvement of resource utilization rate, the reduction of cost, and the satisfaction degree of business requirements;

[0017] Based on the security analysis situation, use a random forest for security decision-making. When the security decision determines that there is a security risk, then use a rule engine to trigger corresponding rules according to the risk type and relevant data of the security risk to generate a security management strategy; the random forest consists of trained classification and regression decision trees.

[0018] Optionally, based on the resource management strategy and the security management strategy, use a multi-objective optimization function and a support vector machine to dynamically manage computing resources, network resources, and storage resources respectively, including:

[0019] Construct the multi-objective optimization function based on cost and resource utilization rate, and solve the multi-objective optimization function using a linear programming algorithm based on the constraint conditions, and configure the virtual machines according to the solution results;

[0020] Based on the resource management policy, adjust the storage capacity, set the data backup frequency, and set the storage type according to the storage capacity in different cloud storage systems;

[0021] Based on the resource management policy and the security management policy, use the support vector machine regression algorithm and the radial basis function to adjust the network bandwidth allocation and firewall rules.

[0022] Optionally, after analyzing resources, security, performance, and cost respectively using deep learning algorithms based on the multi-cloud data and obtaining the analysis results for each item, it further includes:

[0023] Score each cloud service provider based on the analysis results for each item;

[0024] Based on the goals of maximizing resource utilization rate, minimizing cost, maximizing performance, and maximizing security level, calculate the comprehensive scores of each cloud service provider using the weighted summation method;

[0025] Take the cloud service provider with the highest comprehensive score as the target scheduling cloud manufacturer.

[0026] Optionally, after collecting multi-cloud data from the multi-cloud environment, it further includes:

[0027] Clean the multi-cloud data, remove the error data caused by network fluctuations and collection device failures, and fill in the missing values according to the data type and business logic to obtain the first processed data;

[0028] Convert the format of the first processed data according to the conversion rules to obtain the second processed data in a unified format;

[0029] Normalize the second processed data to obtain the data to be analyzed;

[0030] Correspondingly, based on the multi-cloud data, use deep learning algorithms to analyze resources, security, performance, and cost respectively, and obtain the analysis results for each item, including:

[0031] Based on the data to be analyzed, use deep learning algorithms to analyze resources, security, performance, and cost respectively, and obtain the analysis results for each item.

[0032] Optionally, after analyzing resources, security, performance, and cost respectively using deep learning algorithms based on the multi-cloud data and obtaining the analysis results for each item, it further includes:

[0033] Monitor the above analysis situations based on preset thresholds and rules. When an anomaly occurs, determine the alarm level according to the severity and anomaly type;

[0034] Take corresponding alarm methods for alarm according to the alarm level, and send corresponding alarm information.

[0035] The present invention also provides a multi-cloud management system, including:

[0036] A data acquisition module, configured to acquire multi-cloud data from a multi-cloud environment, where the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data;

[0037] An intelligent analysis module, configured to analyze resources, security, performance, and cost respectively based on the multi-cloud data by using deep learning algorithms to obtain various analysis situations;

[0038] An intelligent decision-making module, configured to formulate a resource management strategy and a security management strategy based on the above analysis situations by using reinforcement learning algorithms and random forests;

[0039] A resource management module, configured to dynamically manage computing resources, network resources, and storage resources respectively based on the resource management strategy and the security management strategy by using a multi-objective optimization function and a support vector machine.

[0040] The present invention also provides a multi-cloud management device, including:

[0041] A memory, configured to store a computer program;

[0042] A processor, configured to implement the above multi-cloud management method when executing the computer program.

[0043] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the above multi-cloud management method is implemented.

[0044] It can be seen that the present invention collects multi-cloud data from a multi-cloud environment, where the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data; based on the multi-cloud data, the deep learning algorithm is used to analyze resources, security, performance, and cost respectively to obtain the analysis results of each item; based on the analysis results of each item, the reinforcement learning algorithm and the random forest are used to formulate resource management strategies and security management strategies; based on the resource management strategies and security management strategies, the multi-objective optimization function and the support vector machine are used to dynamically manage computing resources, network resources, and storage resources respectively. The AI (Artificial Intelligence) intelligent analysis based on the multi-cloud data of the present invention can accurately analyze resources, costs, performance, and security in a multi-cloud environment; use AI intelligent drive to formulate management strategies, and use AI intelligent dynamic adjustment of resource allocation and application deployment to ensure the efficient operation of business applications, improve the user experience, effectively improve resource utilization rate, avoid resource idleness and waste, and greatly reduce the enterprise's cost expenditure on cloud resources; through automated resource management, the workload of management personnel is greatly reduced, manual operation errors are reduced, and the overall efficiency and quality of multi-cloud management are improved.

[0045] In addition, the present invention also provides a multi-cloud management system, device, and computer-readable storage medium, which also have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a multi-cloud management method provided by an embodiment of the present invention;

[0048] Figure 2 It is a flow example diagram of a multi-cloud management method provided by an embodiment of the present invention;

[0049] Figure 3 It is a schematic architecture diagram of a multi-cloud management system provided by an embodiment of the present invention;

[0050] Figure 4 It is a schematic structural diagram of a multi-cloud management system provided by an embodiment of the present invention;

[0051] Figure 5 It is a schematic structural diagram of a multi-cloud management device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0053] With the rapid development of information technology, the process of enterprise digital transformation has been accelerating continuously, and more and more enterprises have started to adopt a multi-cloud strategy. The multi-cloud environment brings many advantages to enterprises, such as improving resource flexibility, avoiding vendor lock-in, and optimizing costs. However, the multi-cloud environment also brings a series of management challenges. The interfaces, architectures, resource management methods, security policies, etc. of different cloud service providers are all different. Enterprises need to switch between multiple different consoles to configure and manage resources, resulting in low work efficiency. Moreover, due to the lack of a unified view and coordination mechanism, problems such as uneven resource allocation and inconsistent security policies frequently occur, greatly limiting the effective utilization of the multi-cloud environment by enterprises.

[0054] Currently, there are also some methods and tools for multi-cloud management, but most of them have limitations. For example, traditional management methods mainly rely on manual configuration and rule-based automated scripts. For complex multi-cloud environments, manual configuration is prone to errors and cannot respond to dynamic resource requirements and fault situations in real time. Although rule-based automated scripts can improve efficiency to a certain extent, they lack learning and adaptive capabilities and are difficult to handle the changing conditions and complex scenarios in the multi-cloud environment. Therefore, there is an urgent need for an innovative method to solve these problems.

[0055] To solve the above problems, the present invention provides an AI-based multi-cloud management method. By leveraging the data analysis, learning, and prediction capabilities of AI, it realizes unified management of a multi-cloud environment composed of different cloud service providers (public cloud, private cloud, hybrid cloud, etc.), including centralized monitoring, configuration, and scheduling of resources, solves the problem of decentralized management caused by multi-cloud architecture differences; dynamically optimizes resource allocation, predicts change trends based on business requirements, resource usage history, and real-time data, avoids resource waste and shortages, improves resource utilization rate, and saves costs; automatically detects and responds to security threats, analyzes the security characteristics of each cloud to establish adaptive strategies, and ensures information security in the multi-cloud environment; quickly locates faults, determines the fault source and issues an alarm based on continuous monitoring of the system operation status and machine learning algorithms, and ensures business continuity; at the same time, provides a better user experience for managers and users, simplifies management operations, provides an intuitive and convenient management interface, and enables managers to manage resources more easily. The present invention can enable enterprises to effectively utilize the multi-cloud environment, accelerate the process of enterprise digital transformation, and improve the competitive advantage of enterprises in the market. For details, please refer to Figure 1 , Figure 1 FIG. Figure 1 is a flowchart of a multi-cloud management method provided by an embodiment of the present invention, which specifically provides a method for unifying management, monitoring, scheduling, and optimization of cloud services of multiple cloud service providers in combination with AI. The method may include:

[0056] S101: Collect multi-cloud data from the multi-cloud environment, where the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data.

[0057] The execution entity of this embodiment is a terminal. This embodiment does not limit the type of the terminal, as long as it can complete the operations of the multi-cloud management method. Specifically, the purpose of step S101 is to collect various types of data from the environments of different cloud service providers (including public clouds, private clouds, and hybrid clouds) at regular intervals through a data collection agent program connected to the APIs of each cloud service provider. Such data includes resource data (such as CPU (Central Processing Unit) utilization rate, memory utilization rate, storage capacity usage, network bandwidth utilization rate, etc., and computing, storage, and network resources in each cloud platform), security data (access logs, security vulnerability scan results, authentication and authorization information, etc.), configuration and running status data (configuration information of cloud services and running status data of business applications), and cost data (such as billing models of different cloud services, details of resource usage fees, usage of pay-as-you-go and reserved instances, etc.). By interacting with the APIs of different clouds through multiple interfaces and protocols, the comprehensiveness and accuracy of data acquisition are ensured, providing raw data for subsequent processing and analysis. Among them, the resource data can be further refined into computing resource data (such as CPU utilization rate, memory utilization rate, etc.), storage resource data (information such as capacity, type, read-write frequency, etc.), and network resource data (such as network bandwidth utilization rate, network latency, packet loss rate, etc.).

[0058] Further, to improve accuracy, after collecting multi-cloud data from the multi-cloud environment as described above, the following steps may also be included:

[0059] Step 11: Clean the multi-cloud data, remove the error data caused by network fluctuations and collection device failures, and fill in the missing values according to the data type and business logic to obtain the first processed data;

[0060] Step 12: Convert the format of the first processed data according to the conversion rules to obtain the second processed data in a unified format;

[0061] Step 13: Normalize the second processed data to obtain the data to be analyzed.

[0062] Steps 11 to 13 are the processes of cleaning, transforming, and normalizing the collected original multi-cloud data. Specifically, it can include: cleaning the noise, errors, and missing values in the data, such as removing the incorrect data records caused by network failures or abnormal collection devices, and processing the missing parts in the data through reasonable filling or valuation methods. Moreover, transforming the data with different formats and semantics, and unifying the heterogeneous data from various cloud services into a standard format that can be understood and processed. For example, standardizing the conversion of data with different billing units and pricing methods to facilitate subsequent analysis. At the same time, performing normalization operations on the data, including normalizing the data, so that data with different magnitudes and ranges are comparable in subsequent analysis, thereby improving the accuracy and efficiency of the entire system's analysis and decision-making.

[0063] S102: Based on the multi-cloud data, use deep learning algorithms to analyze resources, security, performance, and cost respectively to obtain the analysis situations of each item.

[0064] Specifically, in this embodiment, for the multi-cloud data, multiple different deep learning algorithms are used to analyze from four aspects of resources, security, performance, and cost respectively to obtain the analysis situations of each aspect.

[0065] Furthermore, the above-mentioned analysis of resources, security, performance, and cost respectively based on the multi-cloud data using deep learning algorithms to obtain the analysis situations of each item can specifically include the following steps:

[0066] Step 21: Based on the resource data, use a long short-term memory network model to perform resource analysis to obtain the resource usage analysis situation; the architecture of the long short-term memory network model includes an input layer, a hidden layer, and an output layer; the input features at each time step in the input layer correspond to the resource usage metrics at a specific moment, and the hidden layer includes multiple input gates, forget gates, and output gates; the early stopping method is introduced in the training process of the long short-term memory network model, and the Adam optimizer is used to adjust the learning rate.

[0067] Step 21 is used to analyze resource data, specifically using the Long Short-Term Memory (LSTM) network model. The reason is that considering that resource data usually has time series characteristics and there are long-term dependency relationships. For example, the resource usage situation (resource data) at the current moment may be related to the usage situations in the past few hours or even days. The LSTM (Long Short-Term Memory) model can selectively remember or forget previous information, thus effectively capturing such long-term dependency relationships. Moreover, the LSTM model can continuously adjust its weight parameters through training to adapt to different data sets and tasks. In resource usage analysis, the resource usage patterns of different business applications and different time periods may vary greatly. The LSTM model can learn these differences, automatically extract the features and patterns in the data, and thus accurately identify the peak and trough periods of resource usage, the resource requirements of different business applications, etc. Compared with some other traditional time series prediction methods, when dealing with time series data with long-term dependencies and complex patterns, the LSTM model can continuously optimize the model, reduce the error between the predicted value and the true value, and provide a more accurate basis for resource management and scheduling.

[0068] The construction process of the LSTM model: input layer, hidden layer, and output layer. The input layer receives the preprocessed time series data, and the input features at each time step correspond to the resource usage metrics at specific moments. The hidden layer consists of multiple LSTM units, and these units have a unique memory gate (input gate, forget gate, output gate) structure, which can selectively remember or forget previous information, thus capturing the long-term dependency relationships in the resource usage data. For example, when predicting the CPU usage rate of cloud services in the next few hours, the model can accurately predict based on the usage patterns at the same time period in the past few days, combined with the recent real-time change trends. The output layer is designed according to the specific prediction task. If predicting the future value of a single resource metric, a fully connected layer is used to output a single value; if predicting the usage trends of multiple resources simultaneously, multiple output nodes can be designed.

[0069] The training process of the LSTM model: The historical resource data is divided into a training set, a validation set, and a test set in chronological order, with a ratio of 70%, 15%, and 15% respectively. The training set is used to train the LSTM model, and the weight parameters of the LSTM model are continuously adjusted through the backpropagation algorithm to minimize the mean squared error (MSE) or other appropriate loss functions between the predicted value and the true value. During the training process, the early stopping method is introduced, that is, when the loss on the validation set no longer decreases, the training is stopped to prevent overfitting. At the same time, the Adam (Adaptive Moment Estimation) optimizer dynamically adjusts the learning rate as the training progresses. The adaptive learning rate adjustment strategy can be combined to improve the convergence speed and accuracy of the model.

[0070] Step 22: Based on the security data, use an autoencoder to obtain reconstructed data, input the security data and the reconstructed data into a discriminator for security analysis, and obtain the security analysis situation.

[0071] Step 22 is used to analyze the security data, specifically using an autoencoder and a discriminator. The reason is that the structure of the autoencoder is relatively simple and the computational efficiency is high. It can encode, decode, and calculate the reconstruction error for newly input data in a short time, and is suitable for real-time or near-real-time security analysis scenarios. It can process a large amount of data generated by the cloud platform in a timely manner, quickly detect anomalies and trigger security alerts, helping to take timely measures to address security threats and reduce potential losses. The security data (system logs, network traffic data, file access records, etc.) generated by the cloud platform is usually high-dimensional data containing a large amount of redundant information. The encoder part of the autoencoder can compress the high-dimensional data into a low-dimensional hidden representation, extract the key features of the data in this process, remove redundancy, making subsequent analysis and processing more efficient, and also helping to discover the internal patterns in the data that are not easily directly observable. The autoencoder is an unsupervised learning algorithm and does not require a large amount of manually labeled normal and abnormal data for training. In the cloud security analysis scenario, it is often difficult and costly to obtain a large amount of accurately labeled abnormal data, while the autoencoder can directly use a large amount of unlabeled normal data for training, automatically learn the distribution and features of the normal data, which greatly reduces the difficulty and cost of data preparation, and improves the practicality and scalability of the model. The autoencoder can be easily combined with other technologies such as the discriminator in GAN (Generative Adversarial Network) to form a more powerful anomaly detection model. Through the adversarial training mechanism, the autoencoder can learn more robust normal data features, further improving the accuracy and reliability of anomaly detection. This flexibility enables the autoencoder to adapt to different security analysis requirements and scenarios, complement other security technologies, and enhance the overall security protection ability.

[0072] Autoencoder part: Collect various types of security data under the normal operating state of the cloud platform as training data. The autoencoder consists of two parts, an encoder and a decoder. The encoder compresses the input high-dimensional data (such as the packet feature vector of network traffic) into a low-dimensional hidden representation, which is achieved through a multi-layer neural network. Each layer gradually reduces the data dimension to extract the key features of the data. The decoder then reconstructs data as similar as possible to the original input based on the low-dimensional representation output by the encoder, also through a multi-layer neural network, gradually restoring the original dimension of the data. During the training process of the autoencoder, the reconstruction error (such as mean squared error) is used as the loss function, and the weights of the encoder and decoder are continuously optimized through the backpropagation algorithm to minimize the reconstruction error, ensuring that the autoencoder can accurately learn the internal patterns of normal data.

[0073] Discriminator part: Introduce the discriminator in GAN, and the discriminator is trained simultaneously with the autoencoder. The input of the discriminator includes the collected security data and the data reconstructed by the autoencoder, and its goal is to distinguish the two as accurately as possible. The discriminator adopts a binary classification neural network architecture, and outputs a probability value representing the authenticity of the data through feature analysis of the input data. During the training process of the autoencoder, it not only needs to minimize the reconstruction error, but also tries to deceive the discriminator, that is, make the discriminator determine the reconstructed data as real data. Through this adversarial training mechanism, the autoencoder can learn more robust normal data features and improve the sensitivity of anomaly detection. In the detection stage, when new data is input into the autoencoder, if the reconstruction error exceeds the threshold based on statistical analysis set in the training process in advance, and the discriminator determines that the reconstructed data is significantly different from the real data, it can be determined as an anomaly, triggering a security alarm to indicate that there may be potential malware activities or other security threats.

[0074] Step 23: Organize the configuration and operation status data into transactions, scan all transactions to count the frequency of each item's occurrence, form frequent item sets for items that meet the minimum support threshold according to the frequency, obtain the correlation between cloud services based on the frequent item sets, and obtain the performance analysis situation based on the correlation.

[0075] Step 23 is used for the analysis of configuration and operation status data, specifically using the Apriori algorithm. The specific reason is that there are a large number of different types of cloud services in the cloud environment. The Apriori algorithm can discover which cloud services are often used together through the analysis of transaction data, reveal the association relationships and usage patterns among cloud services, and help understand the actual requirements of business processes for cloud services. The Apriori algorithm filters frequent item sets based on the support threshold, and this method can effectively reduce the amount of data processing. In the cloud environment, the data volume is usually very large. By setting an appropriate support threshold, those cloud service combinations that do not appear frequently can be quickly filtered out, and only those frequent item sets with practical significance and value are concerned, improving the efficiency and accuracy of data mining. The Apriori algorithm adopts a layer-by-layer search strategy, starting from frequent 1-item sets, and gradually generating and filtering higher-order frequent item sets. This method has high efficiency in processing large-scale data, can obtain relatively accurate results within a reasonable time, and is applicable to the analysis scenarios of a large number of business applications and cloud service usage data in the cloud environment.

[0076] Specifically, step 23 is used to evaluate the performance bottlenecks and optimization spaces of different cloud service components, and find out the key factors affecting the overall performance through association analysis. The configuration and operation status data collected from the multi-cloud management platform (such as detailed records of various cloud services used by different business applications, including information such as the start time, stop time, usage frequency, and resource consumption of the services) are organized into a transaction format. Each transaction represents the usage of cloud services by a business application within a specific time period. For example, a transaction can be [cloud server instance A, object storage service B, network load balancer C], indicating that a certain business uses these three cloud services simultaneously during one operation. Set the minimum support threshold, which is used to filter frequent item sets. Scan all transactions, count the frequency of each item (single cloud service) appearing, and find out the frequent 1-item sets that meet the minimum support. Based on the frequent 1-item sets, candidate 2-item sets, candidate 3-item sets, etc. are continuously generated through combination, and the transactions are scanned again to judge their support. Repeat this process until no new candidate item sets can be generated. The final obtained frequent item sets reflect the frequent co-usage situations among cloud services. For example, if the frequent item set contains [cloud server instance A, database service D], it indicates that these two cloud services are often used together in many business applications, and there is a strong association between them.

[0077] Step 24: Based on the cost data, use the linear regression model for cost analysis to obtain the cost analysis situation; the linear regression model includes a computing cost linear regression model, a storage cost linear regression model, and a network cost linear regression model.

[0078] Step 24 is used to analyze cost data, specifically using a linear regression model. The reason is that the principle of the linear regression model is relatively simple, with well-established statistical inference methods and evaluation metrics. In terms of calculation, the solution algorithm for linear regression is relatively simple and efficient, without the need for complex iterative calculations or a large amount of computing resources. In the scenario of computing costs in cloud services, in many cases, there is an approximately linear relationship between costs and influencing factors. For example, the longer the cloud server runs, the higher the computing cost usually is, and this increase may be approximately linear. The linear regression model can well fit this data with a linear trend. By training on a large amount of historical data, it finds the straight line that best represents the data relationship, thus making a relatively accurate prediction of the computing cost.

[0079] Specifically, the cost data includes detailed cost data collected by the billing system of the cloud service provider, including the usage duration of computing instances, the costs corresponding to the specifications of computing resources (such as the number of CPU cores and memory size), the costs generated by the capacity and type of storage resources (such as block storage, object storage, etc.) and the usage duration, as well as the costs related to the usage volume and peak rate of network bandwidth. The cost data is associated and integrated with the corresponding resource data to ensure that all costs can be mapped to the resource consumption in specific business application scenarios. Regression models are established separately for different components of computing costs, storage costs, and network costs.

[0080] Exemplarily, taking computing costs as an example, the main factors affecting computing costs (such as cloud server running time, CPU load peak, cloud server type, etc.) are used as independent variables in the linear regression model of computing costs, and the computing cost is used as the dependent variable. Through the fitting training of a large amount of historical data, the regression coefficients are determined so that the model can relatively accurately predict the computing cost based on the current resource usage situation. Further, the linear programming algorithm is used to optimize the costs: First, clarify the optimization objective, such as minimizing the total cost under the premise of meeting the business performance requirements; define the constraint conditions, such as the minimum performance requirements of each business application for resources and the total resource limit. Then, take the cost predicted by the linear regression model as part of the objective function, and use a linear programming solver to find the optimal resource allocation plan that meets the constraint conditions to achieve the maximization of cost-effectiveness.

[0081] S103: Based on the various analysis situations, use the reinforcement learning algorithm and random forest to formulate resource management strategies and security management strategies.

[0082] Based on the above analysis, the resources, costs, security, and performance aspects are integrated and analyzed, and resource management strategies and security management strategies are formulated using reinforcement learning algorithms and random forest algorithms. Among them, the resource management strategy is affected by three aspects: resources, costs, and performance. That is, the resource management strategy can affect costs, performance, and resource usage. For example, in cost decision-making, by analyzing price fluctuations and discount strategies of different cloud services, cost optimization suggestions are provided to determine whether resource deployment needs to be adjusted to reduce costs. For example, increase resource reserves during price trough periods; when dealing with performance issues, measures such as adjusting the application deployment method, optimizing cloud service configurations, or reallocating resources can be selected to improve overall performance; in resource configuration decision-making, by increasing or decreasing computing, storage, and network resources in different cloud services, and considering cost factors, the most cost-effective resource allocation plan is selected. For security, whether the security policy needs to be adjusted can be determined based on the security analysis, such as updating access control rules, implementing additional encryption measures, etc., to address potential security risks. By continuously learning and adapting to changes in the multi-cloud environment, this module ensures the scientificity and rationality of decision-making to achieve efficient management of the multi-cloud environment.

[0083] Furthermore, based on the analysis, resource management strategies and security management strategies are formulated using reinforcement learning algorithms and random forest, which can specifically include:

[0084] Step 31: Abstract the multi-cloud environment into a Markov decision process, introduce a neural network, and obtain a resource allocation strategy by using the neural network to approximate the Q-value function; the state space in the Markov decision process includes the resource usage analysis and cost analysis of each cloud service, the action space is resource allocation and scheduling operations, and the reward function is the improvement of resource utilization rate, the reduction of costs, and the satisfaction degree of business requirements.

[0085] In terms of resource management strategies, a reinforcement learning algorithm is specifically adopted. The reason is that the multi-cloud environment is highly dynamic and complex, and factors such as resource usage, business load, and cost are constantly changing. The reinforcement learning algorithm can abstract the multi-cloud environment into a Markov decision process and can well adapt to this dynamically changing environment. By continuously interacting with the environment, it can learn the optimal decision-making strategies in different states to handle various complex situations. Reinforcement learning can design a reward function that includes multi-dimensional factors such as resource utilization rate, cost, and degree of business demand satisfaction to comprehensively consider multiple objectives in resource management. This approach enables decision-making not to be limited to the optimization of a single objective but to balance multiple objectives, achieve efficient utilization of resources, effective control of costs, and stable operation of the business, thereby making decisions that better meet the actual needs. The reinforcement learning algorithm has the ability of autonomous learning and does not require a large amount of labeled data for training. Instead, it continuously adjusts the strategy through trial and error in the environment and gradually finds the optimal resource allocation and scheduling scheme. Over time and with the accumulation of experience, the algorithm can continuously optimize the decision-making to adapt to the changing environment and business needs and improve the efficiency and effectiveness of resource management.

[0086] Specifically, reinforcement learning is used to abstract the multi-cloud environment into a Markov decision process, where the state space includes the resource usage of each cloud service (such as CPU utilization, memory occupancy, remaining storage capacity, etc.), the current business load, and cost metrics (real-time prices of resources, cumulative costs, etc.). The action space is a series of possible resource allocation and scheduling operations, such as migrating a certain task from a high-cost cloud service instance to a low-cost and idle instance, or increasing the resource quota of a cloud service instance. A reward function is designed to comprehensively consider the improvement of resource utilization rate (the reward increases with higher resource utilization rate), the reduction of costs (a positive reward is given for cost reduction), and the satisfaction of business requirements (a negative reward is given if the business experiences lags or failures due to resource adjustment). Since the state space and action space may be relatively complex in a multi-cloud environment, the Q-value table of traditional Q-learning is difficult to handle. Therefore, a neural network is introduced to approximate the Q-value function. The input of the neural network is the state vector, and the output is the Q-value estimation corresponding to each action, effectively handling such complex state and action spaces, being able to learn and process high-dimensional state vectors, and providing a basis for the intelligent agent to make decisions in a complex environment. To solve the instability problem in the reinforcement learning training process, an experience replay mechanism is adopted. The experiences (state, action, reward, next state) during the interaction between the intelligent agent and the environment are stored in the replay buffer. During training, a batch of samples is randomly selected from it for training to avoid the influence of sample correlation on the training effect. At the same time, a target network is set up, and the parameters of the training network are periodically copied to the target network for calculating the target Q-value, making the training process more stable. Through continuous training, DQN can dynamically optimize the resource allocation strategy in a complex multi-cloud environment.

[0087] Step 32: Based on the security analysis situation, use a random forest for security decision-making. When the security decision determines that there is a security risk, then use a rule engine to trigger corresponding rules according to the risk type and relevant data of the security risk to generate a security management strategy; the random forest consists of trained classification and regression decision trees.

[0088] In terms of security management strategies, the random forest algorithm is specifically adopted. The reason is that when constructing each decision tree in the random forest, a subset of features is randomly selected, which enables different decision trees to focus on different combinations of security data features. As a result, information in the data can be mined more comprehensively, and various features in the security-related data can be fully utilized to make security risk judgments, improving the generalization ability of the model. The decision tree structure constructed using CART (Classification and Regression Tree) is intuitive. The path from the root node to the leaf node can clearly show the decision-making process and basis. The splitting attributes and conditions of each node have clear meanings, enabling users to easily understand how the model makes security risk judgments based on different security-related data features, facilitating security managers' analysis and decision-making. During the process of constructing the decision tree, operations such as calculating the Gini index or information gain of different attributes are relatively simple, and the computational complexity is relatively low. A large amount of security data can be processed and a decision tree model can be constructed in a relatively short time, meeting the requirements for rapid processing and analysis of multi-cloud platform security data and promptly discovering security risks.

[0089] Data obtained based on security performance analysis, including but not limited to user login behavior (login time, location, device, whether there are multiple failed logins, etc.), network access logs (source IP (Internet Protocol), destination IP, access port, protocol, access frequency, etc.), and system file operation records (file creation, modification, deletion time, operation subject, etc.). Further clean this data, remove invalid records (such as error logs generated due to network failures), encode the classified data, and discretize the continuous data (such as dividing the login time into different time periods) to facilitate processing by the decision tree model. Select CART. Starting from the root node, calculate the Gini index or information gain of different attributes (such as login location, access port, etc.) based on the training data, select the attribute with the best splitting effect to divide the dataset into sub-datasets, and repeat this process until the stopping condition is met (such as the purity of the sub-dataset reaches a certain standard, the depth of the tree reaches the upper limit, etc.) to construct a decision tree. Random forest trains multiple decision trees through multiple random samplings and combines the prediction results of multiple decision trees to improve the accuracy and robustness of the decision. Through the trained decision tree model, based on the input security-related data, quickly determine whether there is a security risk and the approximate type of the risk. Based on the output of the decision tree, the rule engine further refines the security policy. The rule engine maintains a series of predefined security rules, such as "if a user logs in from multiple different regions within a short period of time and attempts to access sensitive resources multiple times, immediately block the user account and issue an alarm", "if a certain IP address continuously accesses a specific port at a high frequency and the port does not belong to the normal business scope, add the IP to the temporary blacklist", etc. When the decision tree determines that there is a security risk, the rule engine triggers the corresponding rule according to the specific risk type and relevant data to generate specific security management policies, such as updating access control rules, implementing additional encryption measures, increasing the sensitivity of the intrusion detection system, etc., to comprehensively ensure the security of the cloud platform.

[0090] S104: Based on the resource management policy and the security management policy, use the multi-objective optimization function and the support vector machine to dynamically manage the computing resources, network resources, and storage resources respectively.

[0091] It should be noted that in this embodiment, the management of resources depends on the above-mentioned resource management strategy and security management strategy. That is to say, based on the resource management strategy and security management strategy, the dynamic management of resources is ultimately realized to achieve low cost, high performance, high security, and high resource utilization. The essence lies in finding the optimal resource management method when resources, costs, performance, and security all meet the standards. By interacting with the API of the cloud service provider, the dynamic management of computing resources, network resources, and storage resources is realized. Taking the dynamic management of computing resources as an example, this process covers multiple key operations, such as creating new virtual machine instances to provide sufficient computing power support for business expansion; starting the virtual machine instances to make them quickly put into operation to ensure the normal operation of the business system; stopping the temporarily idle or no longer needed virtual machine instances to effectively save energy and resources; and also including flexibly adjusting the configuration of virtual machine instances according to actual business needs. Whether it is increasing the number of CPU cores, expanding the memory capacity, or optimizing other relevant parameters, it can ensure that the computing resources always match the business needs precisely, and ensure that the entire business process operates efficiently and stably under the support of optimal computing resources.

[0092] Furthermore, based on the resource management strategy and security management strategy, the multi-objective optimization function and support vector machine are used to dynamically manage computing resources, network resources, and storage resources respectively. Specifically, it can include:

[0093] Step 41: Construct a multi-objective optimization function based on cost and resource utilization rate, and use the linear programming algorithm to solve the multi-objective optimization function based on the constraint conditions, and configure the virtual machine according to the solution result.

[0094] Specifically, for the management of computing resources, to maximize resource utilization and minimize costs, it is first necessary to accurately define a multi-objective optimization function. Let the resource utilization rate be U and the cost be C. A multi-objective optimization function can be constructed, for example: F = w1 * U - w2 * C, where w1 and w2 are weight coefficients set according to business requirements and priorities. According to the multi-objective optimization function, the management of computing resources is associated with costs and resources, that is, associated with resource management strategies. The resource utilization rate U can be calculated by monitoring the ratio of the usage duration of key resources such as CPU and memory of each virtual machine instance to the total available duration in real time; the cost C is calculated by accumulating according to a certain time period based on the billing model of the cloud service provider, considering factors such as the rental cost of instances and data transmission costs. The linear programming algorithm is used to solve the multi-objective optimization function and determine the constraint conditions, such as the minimum performance requirements of each business application for computing resources (the lower limits of the number of CPU cores, memory size, etc. required to ensure the normal operation of the business), and the upper limit of the total resources provided by the cloud platform (the maximum number of virtual machine instances allowed by the cloud service provider, the upper limit of available physical resources, etc.). These constraint conditions and the multi-objective optimization function are input into the linear programming solver to solve for the optimal virtual machine instance configuration plan, including the timing and parameters for creation, startup, shutdown, and configuration adjustment, to ensure the reasonable allocation of computing resources while paying attention to cost changes and avoiding unnecessary resource waste resulting in increased costs.

[0095] Step 42: Based on the resource management strategy, adjust the storage capacity, set the data backup frequency, and set the storage type according to the storage capacity in different cloud storage systems.

[0096] Specifically, for the management of storage resources, it is possible to allocate and adjust the storage capacity, set storage strategies such as data backup frequency and storage type (hot storage or cold storage), etc. in different cloud storage systems as needed to optimize storage costs. It can be seen that the management of storage resources is associated with resource management strategies.

[0097] Step 43: Based on the resource management strategy and the security management strategy, use the support vector machine regression algorithm and the radial basis function to adjust the network bandwidth allocation and firewall rules.

[0098] Specifically, for the management of network resources, adjust the network bandwidth allocation, set the virtual network topology, set firewall rules, etc. to ensure the stability and efficient operation of the network while controlling the network cost. That is, network resources need to rely on resource management strategies and security management strategies. By selecting the support vector machine (SVM) regression algorithm, the analyzed data is divided into training set, validation set and test set in chronological order, generally with a ratio of 80%, 10%, and 10%. Use the training set to train the model, and find the optimal decision boundary by solving the quadratic programming problem to minimize the mean square error (MSE) between the predicted value and the true value. For nonlinear problems, the radial basis function (RBF) can be used to optimize the model performance by adjusting the kernel parameters (such as gamma value) and penalty coefficient (such as C value). In the prediction process, once the model predicts that a cloud resource may fail within a certain period of time in the future (such as 24 hours), the fault prevention mechanism is immediately started. Send warning information to the resource management module, detailing the basis for fault prediction, the expected time of occurrence, the equipment and resources involved, etc. The resource management module takes corresponding preventive maintenance measures based on the warning information. For cloud servers, if the risk of CPU failure is predicted, the business running on it can be migrated to other idle cloud services in advance; for storage devices, if the disk read and write error rate is predicted to increase, data backup can be performed in advance and spare storage devices can be prepared so that they can be replaced in time when a failure occurs, ensuring the continuous and stable operation of resources and reducing the risk of increased management costs caused by sudden failures. It can be seen that network resource management is related to resource management strategy and security management strategy.

[0099] The support vector machine regression algorithm can build a model with good generalization ability on limited sample data. Even when processing complex historical operation data of cloud environment, it can better capture the rules in the data, accurately predict unknown data, reduce the risk of overfitting, and reliably predict whether cloud resources may fail. In cloud environment, it may be difficult or costly to obtain a large amount of labeled data. The support vector machine regression algorithm can also perform well in the case of small sample data. It can make full use of existing data information and mine potential features in the data, unlike some other algorithms that may require a large amount of data to train an effective model. The support vector machine regression algorithm has a certain robustness to noise and outliers in the data. The operation data in the cloud environment may be disturbed by various factors, generating some noise or outliers. The support vector machine regression algorithm can reduce the impact of these factors on the model to a certain extent and maintain relatively stable prediction performance. The model performance can be optimized by adjusting parameters such as kernel parameters and penalty coefficients. The model can be flexibly adjusted according to different cloud environment data characteristics and prediction requirements to achieve the best prediction effect.

[0100] Further, after analyzing resources, security, performance, and cost respectively using deep learning algorithms based on multi-cloud data and obtaining the analysis results for each item, it may further include:

[0101] Step 51: Score each cloud service provider based on the analysis results for each item;

[0102] Step 52: Based on the goals of maximizing resource utilization, minimizing cost, maximizing performance, and maximizing security level, calculate the comprehensive scores of each cloud service provider using the weighted summation method;

[0103] Step 53: Select the cloud service provider with the highest comprehensive score as the target scheduling cloud vendor.

[0104] Specifically, in order to consider multiple goals such as maximizing resource utilization, minimizing cost, maximizing performance, and maximizing security level simultaneously, the weighted summation method is used to determine the comprehensive evaluation index. The weight of the resource goal is α, the weight of the cost goal is β, the weight of the performance goal is γ, and the weight of the security goal is δ. The score of the resource dimension is represented by "R", the score of the cost dimension is represented by "C", the score of the performance dimension is represented by "P", and the score of the security dimension is represented by "S". Among them, α + β + γ + δ = 1. Calculate the comprehensive scores of each cloud service provider through the formula W = α * R + β * C + γ * P + δ * S, and finally select the one with the highest score as the target scheduling cloud vendor. It should be noted that this comprehensive score changes dynamically within a specified time range, and this dynamic nature can better ensure that the system reaches the optimal state in terms of resource utilization efficiency, cost control level, performance performance, and security protection ability.

[0105] Further, after analyzing resources, security, performance, and cost respectively using deep learning algorithms based on multi-cloud data and obtaining the analysis results for each item, it may further include:

[0106] Step 61: Monitor the analysis results for each item based on preset thresholds and rules, and when an anomaly occurs, determine the alarm level according to the severity and anomaly type;

[0107] Step 62: Take corresponding alarm methods for alarm according to the alarm level and send corresponding alarm information.

[0108] Specifically, the entire multi-cloud environment is monitored in real time. Based on preset thresholds and rules, potential problems and risks are identified by combining various analysis results obtained through intelligent analysis. By continuously tracking key metrics such as resource usage, cost metrics, business application performance, and security status, when these metrics exceed the normal range, for example, when the CPU or memory usage is too high, new security vulnerabilities appear, the business application response time is too long, or the cost exceeds the budget, the alarm mechanism is immediately triggered. The alarm information can be notified to the management personnel through various methods, such as emails, text messages, etc. In addition, this module can classify and prioritize the alarm information and provide suggestions, enabling the management personnel to quickly understand the severity and urgency of the problems, so as to take corresponding measures in a timely manner to ensure the stable operation of the multi-cloud environment and avoid cost increases caused by abnormal situations.

[0109] Applying the multi-cloud management method provided by the embodiments of the present invention, multi-cloud data is collected from the multi-cloud environment, and the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data; based on the multi-cloud data, deep learning algorithms are used to analyze resources, security, performance, and cost respectively to obtain various analysis results; based on the various analysis results, reinforcement learning algorithms and random forests are used to formulate resource management strategies and security management strategies; based on the resource management strategies and security management strategies, multi-objective optimization functions and support vector machines are used to dynamically manage computing resources, network resources, and storage resources respectively. The present invention can achieve accurate analysis of resources, costs, performance, and security in the multi-cloud environment through AI intelligent analysis based on multi-cloud data; use AI intelligent drive to formulate management strategies, and use AI intelligent to dynamically adjust resource allocation and application deployment to ensure the efficient operation of business applications, improve the user experience, effectively improve resource utilization rate, avoid resource idleness and waste, and greatly reduce the enterprise's cost expenditure on cloud resources; through automated resource management, the work burden of management personnel is greatly reduced, and manual operation errors are reduced, improving the overall efficiency and quality of multi-cloud management.

[0110] To make the present invention easier to understand, specifically refer to Figure 2 , Figure 2 which is a flowchart example of a multi-cloud management method provided by the embodiments of the present invention and specifically may include:

[0111] Collect data from a multi-cloud environment, preprocess the data, and store the preprocessed data for intelligent analysis. If the intelligent analysis results find that these metrics exceed the normal range, such as high CPU or memory usage, the emergence of new security vulnerabilities, long business application response times, or cost overruns, etc., the alarm mechanism is immediately triggered; if the intelligent analysis results find that the metrics are normal, intelligent decision-making is performed, and resource scheduling is carried out based on the strategies of intelligent decision-making; if the scheduling is successful, it ends, and if the scheduling fails, an alarm prompt is given.

[0112] The multi-cloud management system provided by the embodiments of the present invention will be introduced below. The multi-cloud management system described below can be mutually corresponding and referred to with the multi-cloud management method described above.

[0113] Specifically, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a multi-cloud management system provided by an embodiment of the present invention, and may include:

[0114] A data collection module 100, configured to collect multi-cloud data from a multi-cloud environment, where the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data;

[0115] An intelligent analysis module 200, configured to analyze resources, security, performance, and cost respectively based on the multi-cloud data by using deep learning algorithms to obtain various analysis situations;

[0116] An intelligent decision-making module 300, configured to formulate resource management strategies and security management strategies based on the various analysis situations by using reinforcement learning algorithms and random forests;

[0117] A resource management module 400, configured to dynamically manage computing resources, network resources, and storage resources respectively based on the resource management strategies and security management strategies by using multi-objective optimization functions and support vector machines.

[0118] Based on the above embodiments, the intelligent analysis module 200 may include:

[0119] A resource analysis unit, configured to analyze resources based on the resource data by using a long short-term memory network model to obtain a resource usage analysis situation; the architecture of the long short-term memory network model includes an input layer, a hidden layer, and an output layer; the input features at each time step in the input layer correspond to resource usage metrics at specific moments, and the hidden layer includes multiple input gates, forget gates, and output gates; the early stopping method is introduced in the training process of the long short-term memory network model, and the learning rate is adjusted by using an Adam optimizer;

[0120] A security analysis unit, which is used to obtain reconstructed data by using an autoencoder based on the security data, input the security data and the reconstructed data into a discriminator for security analysis, and obtain a security analysis situation;

[0121] A performance analysis unit, which is used to organize the configuration and operation status data into transactions, scan all the transactions to count the occurrence frequencies of each item, form frequent item sets for the items that meet the minimum support threshold according to the frequencies, obtain the relevance between cloud services according to the frequent item sets, and obtain a performance analysis situation according to the relevance;

[0122] A cost analysis unit, which is used to perform cost analysis by using a linear regression model based on the cost data, and obtain a cost analysis situation; the linear regression model includes a computing cost linear regression model, a storage cost linear regression model, and a network cost linear regression model.

[0123] Based on the above embodiments, the intelligent decision-making module 300 may include:

[0124] A resource decision-making unit, which is used to abstract the multi-cloud environment into a Markov decision process, introduce a neural network, and obtain a resource allocation strategy by using the neural network to approximate the Q-value function; the state space in the Markov decision process includes the resource usage analysis situation and cost analysis situation of each cloud service, the action space is resource allocation and scheduling operations, and the reward function is the improvement of resource utilization rate, the reduction of cost, and the satisfaction degree of business requirements;

[0125] A security decision-making unit, which is used to perform security decision-making by using a random forest based on the security analysis situation. When the security decision determines that there is a security risk, a rule engine is used to trigger corresponding rules according to the risk type and relevant data of the security risk, and generate a security management strategy; the random forest is composed of trained classification and regression decision trees.

[0126] Based on the above embodiments, the resource management module 400 may include:

[0127] A computing resource management unit, which is used to construct the multi-objective optimization function based on cost and resource utilization rate, and use a linear programming algorithm to solve the multi-objective optimization function based on the constraint conditions, and configure virtual machines according to the solution results;

[0128] A storage resource management unit, which is used to adjust the storage capacity, set the data backup frequency, and set the storage type according to the storage capacity in different cloud storage systems based on the resource management strategy;

[0129] A network resource management unit, which is used to adjust network bandwidth allocation and firewall rules by using a support vector machine regression algorithm and a radial basis function based on the resource management strategy and the security management strategy.

[0130] Based on the above embodiments, the multi-cloud management system may further include:

[0131] A scoring module for scoring each cloud service provider based on the above analysis situations;

[0132] A calculation module for calculating the comprehensive scores of each cloud service provider by using the weighted summation method based on the goals of maximizing resource utilization, minimizing cost, maximizing performance, and maximizing security level;

[0133] A target determination module for taking the cloud service provider with the highest comprehensive score as the target scheduling cloud manufacturer.

[0134] Based on the above embodiments, the multi-cloud management system may further include:

[0135] A data cleaning module for cleaning the multi-cloud data, removing the error data caused by network fluctuations and collection device failures, and filling the missing values according to the data type and business logic to obtain the first processed data;

[0136] A format unification module for converting the format of the first processed data according to the conversion rules to obtain the second processed data with a unified format;

[0137] A normalization module for normalizing the second processed data to obtain the data to be analyzed;

[0138] Correspondingly, the intelligent analysis module 200 is specifically configured to analyze resources, security, performance, and cost respectively based on the data to be analyzed by using deep learning algorithms to obtain various analysis situations.

[0139] Based on the above embodiments, the multi-cloud management system may further include:

[0140] A monitoring module for monitoring the above analysis situations based on preset thresholds and rules, and determining the alarm level according to the severity and abnormal type when an abnormality occurs;

[0141] An alarm module for taking corresponding alarm methods to give an alarm according to the alarm level and sending corresponding alarm information.

[0142] It should be noted that the order of the modules and units in the above multi-cloud management system can be changed before and after without affecting the logic.

[0143] Applying the multi-cloud management system provided by the embodiments of the present invention, through the data collection module 100, which is used to collect multi-cloud data from the multi-cloud environment, and the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data; the intelligent analysis module 200, which is used to analyze resources, security, performance, and cost respectively based on the multi-cloud data by using deep learning algorithms to obtain various analysis situations; the intelligent decision-making module 300, which is used to formulate resource management strategies and security management strategies based on the various analysis situations by using reinforcement learning algorithms and random forests; the resource management module 400, which is used to dynamically manage computing resources, network resources, and storage resources based on the resource management strategies and security management strategies by using multi-objective optimization functions and support vector machines. This system can achieve precise analysis of resources, costs, performance, and security in the multi-cloud environment through AI intelligent analysis based on multi-cloud data; use AI intelligent drive to formulate management strategies, and use AI intelligent to dynamically adjust resource allocation and application deployment to ensure the efficient operation of business applications, improve user experience, effectively improve resource utilization rate, avoid resource idleness and waste, and greatly reduce the cost expenditure of enterprises in cloud resources; through automated resource management, it greatly reduces the workload of management personnel, reduces manual operation errors, and improves the overall efficiency and quality of multi-cloud management.

[0144] For a better understanding of the present invention, please specifically refer to Figure 4 , Figure 4 which is a schematic structural diagram of a multi-cloud management system provided by an embodiment of the present invention, and specifically may include:

[0145] A multi-cloud environment, which includes public clouds, private clouds, and hybrid clouds; a data collection module (the same as the above data collection module 100), a data preprocessing module (composed of the above data cleaning module, format unification module, and normalization module), an intelligent analysis module (the same as the above intelligent analysis module 200), an intelligent decision-making module (the same as the above intelligent decision-making module 300), a resource management module (the same as the above resource management module 400), and a monitoring and warning module (composed of the above monitoring module and warning module).

[0146] Next, the multi-cloud management device provided by the embodiments of the present invention will be introduced, and the multi-cloud management device described below can be mutually referred to with the multi-cloud management method described above.

[0147] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a multi-cloud management system device provided by an embodiment of the present invention, and may include:

[0148] A memory 10, which is used to store computer programs;

[0149] A processor 20 for executing a computer program to implement the above multi-cloud management method.

[0150] The memory 10, the processor 20, and the communication interface 31 all complete their mutual communication through the communication bus 32.

[0151] In an embodiment of the present invention, the memory 10 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions. In an embodiment of the present invention, the memory 10 may store programs for implementing the following functions:

[0152] Collect multi-cloud data from a multi-cloud environment. The multi-cloud data includes resource data, security data, configuration and running status data, and cost data;

[0153] Based on the multi-cloud data, use deep learning algorithms to analyze resources, security, performance, and cost respectively to obtain various analysis situations;

[0154] Based on various analysis situations, use reinforcement learning algorithms and random forests to formulate resource management strategies and security management strategies;

[0155] Based on the resource management strategy and the security management strategy, use multi-objective optimization functions and support vector machines to dynamically manage computing resources, network resources, and storage resources respectively.

[0156] In a possible implementation, the memory 10 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store data created during use.

[0157] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include NVRAM. The memory stores an operating system and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.

[0158] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices. The processor 20 may be a microprocessor or any conventional processor, etc. The processor 20 may call the program stored in the memory 10.

[0159] The communication interface 31 can be the interface of the communication module for connecting to other devices or systems.

[0160] Of course, it should be noted that Figure 5 the structure shown does not constitute a limitation on the multi-cloud management device in the embodiments of the present invention. In practical applications, the multi-cloud management device may include more or fewer components than Figure 5 those shown, or combine some components.

[0161] Next, the computer-readable storage medium provided by the embodiments of the present invention will be introduced. The computer-readable storage medium described below can be correspondingly referred to the multi-cloud management method described above.

[0162] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above multi-cloud management method are implemented.

[0163] The computer-readable storage medium can include various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0164] In the present specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0165] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0166] Finally, it should also be noted that in this text, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.

[0167] The above has introduced in detail a multi-cloud management method, system, device and computer-readable storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A multi-cloud management method, characterized in that: include: Collect multi-cloud data from a multi-cloud environment, wherein the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data; Based on the multi-cloud data, a deep learning algorithm is used to analyze resources, security, performance, and cost respectively to obtain various analysis results; Based on the above analysis, use reinforcement learning algorithm and random forest to formulate resource management strategy and security management strategy; Based on the resource management strategy and security management strategy, multi-objective optimization function and support vector machine are used to dynamically manage computing resources, network resources and storage resources respectively.

2. The multi-cloud management method according to claim 1, characterized in that: Based on the multi-cloud data, deep learning algorithms are used to analyze resources, security, performance, and costs, respectively, to obtain various analysis results, including: Based on the resource data, a long short-term memory network model is used to perform resource analysis to obtain a resource usage analysis situation; the architecture of the long short-term memory network model includes an input layer, a hidden layer and an output layer; the input feature of each time step in the input layer corresponds to a resource usage indicator at a specific moment, and the hidden layer includes multiple input gates, forget gates and output gates; the early stopping method is introduced into the training process of the long short-term memory network model, and the Adam optimizer is used to adjust the learning rate; Based on the security data, reconstructed data is obtained by using an autoencoder, and the security data and the reconstructed data are input into a discriminator for security analysis to obtain a security analysis situation; Arrange the configuration and operation status data into transactions, count the frequency of occurrence of each item by scanning all the transactions, form frequent item sets based on the frequency, obtain the correlation between cloud services based on the frequent item sets, and obtain performance analysis based on the correlation; Based on the cost data, a linear regression model is used to perform cost analysis to obtain a cost analysis situation; the linear regression model includes a computational cost linear regression model, a storage cost linear regression model, and a network cost linear regression model.

3. The multi-cloud management method according to claim 1, characterized in that: Based on the above analysis, reinforcement learning algorithms and random forests are used to formulate resource management strategies and security management strategies, including: The multi-cloud environment is abstracted into a Markov decision process, and a neural network is introduced. The resource allocation strategy is obtained by using the neural network to approximate the Q value function; the state space in the Markov decision process includes the resource usage analysis and cost analysis of each cloud service, the action space is the resource allocation and scheduling operation, and the reward function is the improvement of resource utilization, the reduction of cost and the satisfaction of business needs; Based on the security analysis, random forest is used to make security decisions. When the security decision determines that there is a security risk, the rule engine is used to trigger the corresponding rules according to the risk type and related data of the security risk to generate a security management strategy; the random forest is composed of trained classification and regression decision trees.

4. The multi-cloud management method according to claim 1, characterized in that: Based on the resource management strategy and security management strategy, multi-objective optimization function and support vector machine are used to dynamically manage computing resources, network resources and storage resources respectively, including: Constructing the multi-objective optimization function based on cost and resource utilization, solving the multi-objective optimization function using a linear programming algorithm based on constraint conditions, and configuring the virtual machine according to the solution result; Based on the resource management strategy, adjust the storage capacity, set the data backup frequency and set the storage type according to the storage capacity in different cloud storage systems; Based on the resource management strategy and security management strategy, support vector machine regression algorithm and radial basis function are used to adjust network bandwidth allocation and firewall rules.

5. The multi-cloud management method according to claim 1, characterized in that: Based on the multi-cloud data, the deep learning algorithm is used to analyze resources, security, performance and cost respectively, and after obtaining various analysis results, the following is also included: Score each cloud service provider based on the above analysis; Based on the goals of maximizing resource utilization, minimizing costs, maximizing performance, and maximizing security, the weighted sum method is used to calculate the comprehensive score of each cloud service provider; The cloud service provider with the highest comprehensive score is selected as the target cloud vendor for scheduling.

6. The multi-cloud management method according to claim 1, characterized in that: After collecting multi-cloud data from a multi-cloud environment, it also includes: Cleaning the multi-cloud data to remove erroneous data caused by network fluctuations and acquisition device failures, and filling missing values ​​according to data types and business logic to obtain first processed data; Convert the first processed data into a uniform format according to the conversion rule to obtain second processed data in a uniform format; performing normalization processing on the second processed data to obtain data to be analyzed; Accordingly, based on the multi-cloud data, a deep learning algorithm is used to analyze resources, security, performance, and cost respectively to obtain various analysis conditions, including: Based on the data to be analyzed, a deep learning algorithm is used to analyze resources, security, performance and cost respectively to obtain various analysis results.

7. The multi-cloud management method according to claim 1, characterized in that: Based on the multi-cloud data, the deep learning algorithm is used to analyze resources, security, performance and cost respectively, and after obtaining various analysis results, the following is also included: Monitor the above analysis situations based on preset thresholds and rules, and when an abnormality occurs, determine the alarm level according to the severity and type of abnormality; According to the alarm level, a corresponding alarm method is adopted to issue an alarm and send corresponding alarm information.

8. A multi-cloud management system, characterized in that: include: A data collection module, used to collect multi-cloud data from a multi-cloud environment, wherein the multi-cloud data includes resource data, security data, configuration and operation status data, and cost data; An intelligent analysis module, for analyzing resources, security, performance, and cost respectively based on the multi-cloud data using a deep learning algorithm to obtain various analysis results; An intelligent decision-making module, which is used to formulate resource management strategies and security management strategies based on the above-mentioned analysis situations using reinforcement learning algorithms and random forests; The resource management module is used to dynamically manage computing resources, network resources and storage resources respectively based on the resource management strategy and security management strategy using a multi-objective optimization function and a support vector machine.

9. A multi-cloud management device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the multi-cloud management method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the multi-cloud management method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Cost optimization method in multi-cloud environment

    CN121056255A