A reinforcement learning system and method for resource scheduling in a multi-cloud environment
By constructing a dynamic graph structure and reinforcement learning system in a multi-cloud environment, resource scheduling is monitored and optimized in real time, solving the problem of insufficient global optimization capability of resource scheduling in a multi-cloud environment, and realizing efficient, reliable and global optimization of cross-cloud resource scheduling.
Patent Information
- Application Number
- CN202511120316.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-12
AI Technical Summary
In a multi-cloud computing environment, traditional methods are insufficient in dealing with heterogeneity, dynamism, and multi-objective optimization needs, which limits the global optimization capability of multi-cloud resource scheduling.
The environmental monitoring module collects performance monitoring data in the multi-cloud environment in real time, constructs a dynamic graph structure, uses graph neural networks to generate structured state vectors, combines reinforcement learning strategy network analysis to generate cross-cloud resource scheduling instructions, combines autoregressive decoder to process action parameters, generates application programming interface instructions, and optimizes resource scheduling through the strategy adjustment module.
It significantly improves the global optimization capability of resource scheduling in multi-cloud environments, ensures the reliable execution of scheduling instructions and cross-platform compatibility, realizes global multi-objective collaborative optimization, and avoids resource waste and performance imbalance.
Smart Images

Figure CN120631595B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-cloud resource scheduling, in particular to a reinforcement learning system and method for resource scheduling in a multi-cloud environment. BACKGROUND
[0002] In a multi-cloud computing environment, how to effectively control programs, achieve cross-cloud task scheduling and resource allocation, and achieve globally optimal cost-effectiveness, performance, and resource utilization is an important technical problem faced by the current cloud computing field.
[0003] Traditional methods have obvious shortcomings in dealing with the heterogeneity, dynamics, complex dependency relationships, and multi-objective optimization requirements of multi-cloud environments, resulting in limited global optimization capabilities for multi-cloud resource scheduling.
[0004] Therefore, a reinforcement learning system and method for resource scheduling in a multi-cloud environment are proposed. SUMMARY
[0005] The present application aims to provide a reinforcement learning system and method for resource scheduling in a multi-cloud environment, which collects real-time first state performance monitoring data of multiple heterogeneous cloud platforms in a multi-cloud environment through an environment monitoring module, and constructs a dynamic graph structure; based on graph neural networks, the dynamic graph is subjected to multiple rounds of message passing and node aggregation to generate a structured state vector that integrates topological features; based on a policy network, the state vector is analyzed to output action parameters including resource selection probabilities and configuration parameters, the action parameters are analyzed by a self-recurrent decoder to generate a cross-cloud resource scheduling instruction sequence, and the sequence is converted into an application program interface instruction for execution. A strategy adjustment module constructs a multi-cloud environment resource scheduling score to drive dynamic optimization of the policy network. The present application can effectively improve the global optimization capability of resource scheduling in a multi-cloud environment.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0007] A reinforcement learning system for resource scheduling in a multi-cloud environment, comprising:
[0008] An environment monitoring module for collecting real-time first state performance monitoring data of multiple heterogeneous cloud platforms in a multi-cloud environment;
[0009] A state representation module for constructing a dynamic graph structure based on resource entities and pending workloads based on the monitoring data, and using a graph neural network to analyze the graph structure and generate a structured state vector;
[0010] A reinforcement learning agent module for analyzing the structured state vector through a reinforcement learning policy network to obtain action parameters based on resource scheduling;
[0011] an action decoding module configured to process the action parameters through a self-recursive decoding process to obtain a sequence of cross-cloud resource scheduling instructions;
[0012] an action execution module configured to convert the sequence of cross-cloud resource scheduling instructions into application program interface instructions and submit the application program interface instructions to a plurality of cloud platforms for execution;
[0013] a policy adjustment module configured to obtain second state performance monitoring data after execution of the plurality of heterogeneous cloud platforms, construct a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjust the reinforcement learning policy network based on the score.
[0014] Preferably, the first state performance monitoring data comprises virtual machine instance monitoring data, container instance monitoring data, database service instance monitoring data, storage volume monitoring data, and pending workload monitoring data.
[0015] Preferably, the nodes of the graph structure are used to represent the entity representations of virtual machine instances, container instances, database service instances, storage volumes, and pending workloads; the edges of the graph structure are used to represent the relationships between the entities; and the types of the relationships comprise network connection relationships between resource entities, deployment subordination relationships between workloads and resources, running subordination relationships between workloads and resources, data access dependency relationships between workloads, dependency relationships between workloads, and invocation dependency relationships between services.
[0016] Preferably, the process of obtaining the structured state vector comprises: inputting the dynamic graph structure and the attributes of the nodes and edges of the dynamic graph structure into the graph neural network; performing multiple rounds of message passing and node state aggregation operations through the graph neural network to generate node embedding representations that fuse the attributes of each node and the attributes of adjacent nodes; and aggregating the node embedding representations to obtain the structured state vector.
[0017] Preferably, the specific process of obtaining the action parameters based on resource scheduling comprises: feeding the structured state vector as input into the reinforcement learning policy network; performing forward calculation through the policy network to output the action parameters representing resource scheduling decision intentions, wherein the action parameters comprise probability distributions of selecting specific resources and numerical values for defining resource configurations.
[0018] Preferably, the obtaining process of the cross-cloud resource scheduling instruction sequence comprises: initializing the autoregressive decoding process and taking the action parameter as input; in each decoding step, predicting and determining the next action component part according to the current decoding state and the previously generated sequence part by using the autoregressive model; appending the determined component part to the current action sequence; and repeating the decoding step until the cross-cloud resource scheduling instruction sequence containing complete scheduling instructions is output.
[0019] Preferably, the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighted combination of a plurality of basic indexes of a plurality of heterogeneous cloud platforms; the plurality of basic indexes include resource cost index, application performance index, SLA compliance index and resource utilization efficiency index.
[0020] A reinforcement learning method for resource scheduling in a multi-cloud environment, comprising:
[0021] S1. Real-time collection of first state performance monitoring data of a plurality of heterogeneous cloud platforms in a multi-cloud environment;
[0022] S2. Construction of a dynamic graph structure based on resource entities and to-be-processed workloads based on monitoring data, and generation of a structured state vector by using a graph neural network to analyze the graph structure;
[0023] S3. Analysis of the structured state vector by a reinforcement learning strategy network to obtain an action parameter based on resource scheduling;
[0024] S4. Processing of the action parameter by an autoregressive decoding process to obtain a cross-cloud resource scheduling instruction sequence;
[0025] S5. Conversion of the cross-cloud resource scheduling instruction sequence into an application program interface instruction and submission to a plurality of cloud platforms for execution;
[0026] S6. Obtaining of second state performance monitoring data after execution of the plurality of heterogeneous cloud platforms, construction of a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjustment of the reinforcement learning strategy network based on the score.
[0027] Compared with the prior art, the present application has the following beneficial effects:
[0028] 1、The application constructs a dynamic graph structure based on the state performance monitoring data of multiple heterogeneous cloud platforms after execution, models the complex dependency relationship (such as network connection, deployment dependency, data access) between resource entities such as virtual machines, containers and databases and workloads in real time, combines the multi-round message passing mechanism of graph neural network (GNN), accurately captures the topology dynamics of the multi-cloud environment, significantly improves the accuracy of resource state representation, makes the scheduling decision more adaptive to the dynamic change scene, and thus is beneficial to improving the global optimization ability of resource scheduling in the multi-cloud environment.
[0029] 2、The application processes the action parameters through an autoregressive decoding process to obtain a cross-cloud resource scheduling instruction sequence; converts the cross-cloud resource scheduling instruction sequence into an application program interface instruction and submits it to multiple cloud platforms for execution; ensures reliable execution of the cross-cloud scheduling instruction and cross-platform compatibility, thereby facilitating improvement of the global optimization ability of resource scheduling in the multi-cloud environment.
[0030] 3、The application designs a multi-cloud environment resource scheduling score that integrates resource cost indicators, application performance indicators, SLA compliance indicators and resource utilization efficiency indicators, drives the reinforcement learning strategy network to iterate in the direction of comprehensive optimization, avoids resource waste or performance imbalance caused by single indicator optimization, and realizes global multi-objective collaborative optimization of resource scheduling in the multi-cloud environment. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 A structural schematic diagram of a reinforcement learning system for resource scheduling in a multi-cloud environment is provided for the embodiments of the application.
[0032] Figure 2 A flowchart of a reinforcement learning method for resource scheduling in a multi-cloud environment is provided for the embodiments of the application. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0034] Embodiment one: In order to improve the global optimization ability of resource scheduling in a multi-cloud environment, a reinforcement learning system for resource scheduling in a multi-cloud environment is applied, Figure 1 A structural schematic diagram of a reinforcement learning system for resource scheduling in a multi-cloud environment is provided for the embodiments of the application, which comprises:
[0035] An environment monitoring module is used to collect first state performance monitoring data of multiple heterogeneous cloud platforms in a multi-cloud environment in real time.
[0036] Further, the multi-cloud environment includes an AWS platform and an Azure platform; the first state performance monitoring data includes virtual machine instance monitoring data, container instance monitoring data, database service instance monitoring data, storage volume monitoring data, and pending workload monitoring data;
[0037] The virtual machine instance monitoring data includes virtual machine instance configuration data (instance ID, etc.), virtual machine instance performance data (including CPU utilization, memory utilization, disk read-write IOPS, disk read-write throughput, cross-cloud network delay, bandwidth utilization, etc.), virtual machine instance cost data (consumption cost per unit hour), and virtual machine instance state data (including running state and stopping state);
[0038] The container instance monitoring data includes container instance configuration data (including POD ID and host node ID); the host node indicates which virtual machine instance it is in; container instance running state data (running success state, running failure state, and abort state); resource data (CPU core number and memory size); and container instance performance data (network throughput and cross-container communication delay);
[0039] The database service instance monitoring data includes database service instance configuration data (instance name, database engine type and platform to which it belongs, and storage configuration), database service instance performance data (CPU utilization, memory utilization, storage usage, and average query delay), database service instance cost data (consumption cost per unit hour), and database service instance state data (preparation state and update state, etc.);
[0040] The storage volume monitoring data includes storage volume configuration data (Volume ID and platform to which it belongs, etc.), storage volume performance data (read IOPS, write IOPS, average read delay, and average write delay, etc.), storage volume state data (in-use state and available state, etc.), and storage volume cost data (storage cost);
[0041] The pending workload monitoring data includes pending workload attribute data (task ID, priority, and submission time), SLA requirement (workload deadline), queue state data (queue order, queue length, and queue waiting time), and workload demand resource amount (CPU core number and memory size);
[0042] The state representation module is configured to construct a dynamic graph structure based on the resource entities and the pending workloads based on the monitoring data, and analyze the graph structure by using a graph neural network to generate a structured state vector.
[0043] Further, the nodes of the graph structure are used to embody virtual machine instances, container instances, database service instances, storage volumes, and entity representations of to-be-processed workloads; the edges of the graph structure are used to represent relationships between the entities; types of the relationships include: network connection relationships between resource entities, deployment dependency relationships between workloads and resources, running dependency relationships between workloads and resources, data access dependency relationships between workloads, dependency relationships between workloads, and invocation dependency relationships between services;
[0044] Further, the dynamic graph structure is ; representing nodes; representing edge relationships; representing node attributes; representing edge attributes;
[0045] Each node includes a virtual machine instance, a container instance, a database service instance, a storage volume, and a to-be-processed workload;
[0046] The node attribute is represented as a feature vector, which is integrated by data of the node;
[0047] The virtual machine instance attribute includes a vector composed of an instance ID, CPU utilization, memory utilization, disk read-write IOPS, disk read-write throughput, cross-cloud network delay, bandwidth utilization, unit-hour consumption cost, and a state;
[0048] The container instance attribute is a vector composed of a POD ID, a host node ID, container instance running state data, a number of CPU cores, a memory size, network throughput, and cross-container communication delay;
[0049] The database service instance attribute includes a vector composed of an instance name, a database engine type, a platform to which the database service instance belongs, a storage configuration, CPU utilization, memory utilization, storage usage, average query delay, unit-hour consumption cost, and database service instance state data;
[0050] The storage volume attribute includes a vector composed of a Volume ID, a platform to which the storage volume belongs, read IOPS, write IOPS, average read delay, average write delay, storage volume state data, and storage cost;
[0051] The to-be-processed workload attribute includes a vector composed of a task ID, a priority, a submission time, a workload deadline, a queue order, a queue length, a queue waiting time, a number of CPU cores required by the workload, and a memory size of the workload;
[0052] The network connection relationship between resource entities represents the logical network connection state between different resource entities (such as virtual machines, containers, databases), reflecting the communication capability and performance across clouds and between resources in the same cloud;
[0053] The edge attribute corresponding to the network connection relationship between resource entities includes a vector composed of cross-cloud communication delay, intra-cloud communication delay, bandwidth usage, and cross-cloud transmission cost;
[0054] The deployment dependency relationship between workloads and resources represents the binding relationship of workloads (such as containers, tasks) being deployed to resource entities (such as virtual machines, storage volumes);
[0055] The edge attribute corresponding to the deployment dependency relationship between workloads and resources includes a vector composed of deployment timestamp, expected replica number, actual running replica number, and scaling strategy;
[0056] The running dependency relationship between workloads and resources represents the actual resource quota (such as allocated CPU cores, memory size) occupied by the workload during runtime and its real-time utilization rate (such as CPU usage rate 80%), and also contains the upper limit constraint of the resource quota (such as maximum memory limit of 8GB).
[0057] The edge attribute corresponding to the running dependency relationship between workloads and resources includes a vector composed of allocated CPU cores, allocated memory size, resource quota upper limit, CPU utilization rate, and memory utilization rate;
[0058] The data access dependency relationship between workloads represents the read-write dependency relationship of workloads on data storage resources (such as databases, storage volumes, message queues), including access mode (read / write / read-write), data size (such as 500MB / second), and criticality level (such as high / medium / low);
[0059] The edge attribute corresponding to the data access dependency relationship between workloads includes a vector composed of access mode and criticality level;
[0060] The dependency relationship between workloads represents the execution order or data transfer dependency between tasks, for example, task A is completed before task B can be started, or task C depends on the output data of task D;
[0061] The edge attribute corresponding to the dependency relationship between workloads includes a vector composed of trigger condition, transmission data volume, and dependency timeout time;
[0062] The calling dependency relationship between services represents the calling chain relationship between containerized microservices through API or event triggering;
[0063] The edge attribute corresponding to the calling dependency relationship between services includes a vector constructed by average response time, error rate, communication protocol, and request frequency;
[0064] Further, the structured state vector acquisition process includes: inputting the dynamic graph structure and the attributes of its nodes and edges into the graph neural network; performing multiple rounds of message passing and node state aggregation operations by the graph neural network; the node state aggregation operation generates a node representation that integrates local and global features by aggregating the node's own attributes and the attributes of its neighbor nodes;
[0065] In this embodiment, 3 rounds of message passing are performed; a node embedding representation that integrates its own attributes and the attributes of adjacent nodes is generated for each node in the graph; and the node embedding representations are aggregated to obtain the structured state vector;
[0066] Further, message passing includes: each node sends a message to its adjacent nodes, and the message content is composed of node attributes and edge attributes; for example, a virtual machine instance node transmits network delay and bandwidth utilization to a connected container instance node;
[0067] All node embeddings are aggregated by averaging to generate a structured state vector;
[0068] The reinforcement learning agent module is used to analyze the structured state vector through a reinforcement learning policy network to obtain resource scheduling-based action parameters;
[0069] This embodiment constructs a dynamic graph structure based on the state performance monitoring data after the execution of multiple heterogeneous cloud platforms, and models the complex dependency relationships (such as network connections, deployment dependencies, and data access) between resource entities such as virtual machines, containers, and databases and workloads in real time. Combined with the multi-round message passing mechanism of the graph neural network (GNN), the topological dynamics of the multi-cloud environment are accurately captured, the accuracy of resource state representation is significantly improved, the scheduling decision is more adaptive to dynamic changes, and the global optimization capability of resource scheduling in the multi-cloud environment is improved. Limited.
[0070] Further, the specific acquisition process of the resource scheduling-based action parameters includes: feeding the structured state vector as input into the reinforcement learning policy network; performing forward calculation by the policy network to output the action parameters representing the resource scheduling decision intent, the action parameters including a probability distribution of selecting a specific resource and a numerical value for defining resource configuration.
[0071] The probability distribution represents the probability of selecting each cloud platform resource (such as AWS: 0.6, Azure: 0.3); under the current state, the probability of selecting AWS resources is 60%, and the probability of selecting Azure resources is 30%;
[0072] The resource configuration numerical value includes CPU core number, memory size, etc. (such as 4 cores, 8 GB);
[0073] Further, the reinforcement learning policy network comprises an input layer, a hidden layer and an output layer; the input layer is configured to input the structured state vector into the reinforcement learning policy network; the hidden layer comprises 256 neurons; and the output layer is configured to output an action parameter based on resource scheduling.
[0074] An action decoding module is configured to process the action parameter through a self-recurrent decoding process to obtain a cross-cloud resource scheduling instruction sequence.
[0075] Further, the process of obtaining the cross-cloud resource scheduling instruction sequence comprises: initializing the self-recurrent decoding process and inputting the action parameter; in each decoding step, using the self-recurrent model (Transformer model) to predict and determine a next action component part according to a current decoding state and a previously generated sequence part; appending the determined component part to the current action sequence; and repeating the decoding step until the cross-cloud resource scheduling instruction sequence containing a complete scheduling instruction is output.
[0076] An action execution module is configured to convert the cross-cloud resource scheduling instruction sequence into an application program interface instruction and submit the application program interface instruction to multiple cloud platforms for execution.
[0077] The embodiment processes the action parameter through a self-recurrent decoding process to obtain a cross-cloud resource scheduling instruction sequence, converts the cross-cloud resource scheduling instruction sequence into an application program interface instruction and submits the application program interface instruction to multiple cloud platforms for execution, thereby ensuring reliable execution of the cross-cloud scheduling instruction and cross-platform compatibility, and thus facilitating improvement of the global optimization capability of resource scheduling in a multi-cloud environment.
[0078] A policy adjustment module is configured to obtain second state performance monitoring data after execution of the multiple heterogeneous cloud platforms, construct a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjust the reinforcement learning policy network based on the score.
[0079] Further, the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighted combination of multiple basic indexes of the multiple heterogeneous cloud platforms; the multiple basic indexes comprise a resource cost index, an application performance index, an SLA compliance index and a resource utilization efficiency index.
[0080] The resource cost index obtaining process is to calculate a product of a unit-hour consumption fee and a usage hour number of each node, and obtain a sum of the product and a cross-cloud transmission cost.
[0081] The process of obtaining application performance metrics involves analyzing virtual machine instance performance data, container instance performance data, database service instance performance data, and storage volume performance data. Specifically, the following steps are taken: weighting each data point from the virtual machine instance performance data to obtain first performance data; weighting each data point from the container instance performance data to obtain second performance data; weighting each data point from the database service instance performance data to obtain third performance data; weighting each data point from the storage volume performance data to obtain fourth performance data; and then weighting these first, second, third, and fourth performance data to obtain the application performance metrics. The weighting method is based on balanced weights, with each data point having a weight of 1 / the number of data points.
[0082] SLA compliance is obtained by quoting the SLA latency requirements of all tasks in the workload with the actual latency time.
[0083] The resource utilization efficiency metric is obtained by weighting the CPU utilization and memory utilization of all nodes; specifically, it includes the following:
[0084] ;
[0085] in, Indicators representing resource utilization efficiency; Indicates the number of nodes; Indicates the first CPU utilization of each node; Indicates the first Memory utilization of each node; Indicates CPU utilization weight; Indicates the weight of memory utilization; and Take values of 0.5 and 0.5 respectively;
[0086] Furthermore, the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighted combination of multiple basic indicators from multiple heterogeneous cloud platforms; specifically including:
[0087] ;
[0088] in, This indicates the resource scheduling score for a multi-cloud environment; Indicates platform The improvement value of resource utilization efficiency indicators; Indicates platform The savings value of resource cost indicators; Indicates platform The improvement in application performance indicators; , and Two comprehensive scores are obtained respectively based on the second state performance monitoring data and the first state performance monitoring data of each platform respectively, and a difference value is obtained; the comprehensive score obtaining process is: first, the weighted values of multiple basic indexes of each platform are summed to obtain the score value of each platform, and the score values of each platform are summed to obtain the comprehensive score; the adjusting the reinforcement learning strategy network based on the score comprises: if the multi-cloud environment resource scheduling score is less than a preset threshold, adjusting the action parameter based on resource scheduling. The action parameter includes a probability distribution of selecting a specific resource and a numerical value for defining resource configuration; until the multi-cloud environment resource scheduling score is greater than the preset threshold.
[0089] The embodiment designs a multi-cloud environment resource scheduling score integrating resource cost indicators, application performance indicators, SLA compliance indicators and resource utilization efficiency indicators, drives the reinforcement learning strategy network to iterate in the direction of comprehensive optimization, avoids resource waste or performance imbalance caused by single indicator optimization, and realizes global multi-objective collaborative optimization of resource scheduling in a multi-cloud environment.
[0090] Embodiment two: in order to improve the global optimization capability of resource scheduling in a multi-cloud environment, a reinforcement learning method for resource scheduling in a multi-cloud environment is applied, Figure 2 A flowchart of a reinforcement learning method for resource scheduling in a multi-cloud environment provided by the embodiment of the present application, comprising:
[0091] S1. Real-time collection of first state performance monitoring data of multiple heterogeneous cloud platforms in a multi-cloud environment;
[0092] S2. Constructing a dynamic graph structure based on resource entities and to-be-processed workloads based on monitoring data, and analyzing the graph structure by using a graph neural network to generate a structured state vector;
[0093] S3. Analyzing the structured state vector by a reinforcement learning strategy network to obtain an action parameter based on resource scheduling;
[0094] S4. Processing the action parameter through a self-recurrent decoding process to obtain a cross-cloud resource scheduling instruction sequence;
[0095] S5. Converting the cross-cloud resource scheduling instruction sequence into an application program interface instruction and submitting it to multiple cloud platforms for execution;
[0096] S6. Obtaining second state performance monitoring data after execution of the multiple heterogeneous cloud platforms, constructing a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjusting the reinforcement learning strategy network based on the score.
[0097] Further, the first state performance monitoring data comprises virtual machine instance monitoring data, container instance monitoring data, database service instance monitoring data, storage volume monitoring data and pending workload monitoring data.
[0098] Further, the nodes of the graph structure are used to embody entity representations of virtual machine instances, container instances, database service instances, storage volumes and pending workloads; the edges of the graph structure are used to represent relationships between the entities; and the types of the relationships comprise network connection relationships between resource entities, deployment dependency relationships between workloads and resources, running dependency relationships between workloads and resources, data access dependency relationships between workloads, dependency relationships between workloads and invocation dependency relationships between services.
[0099] Further, the obtaining process of the structured state vector comprises: inputting the dynamic graph structure and attributes of nodes and edges thereof into the graph neural network; performing multiple rounds of message passing and node state aggregation operations by the graph neural network to generate node embedding representations of each node in the graph, which fuse attributes of the node itself and neighboring nodes; and aggregating the node embedding representations to obtain the structured state vector.
[0100] Further, the specific obtaining process of the action parameter based on resource scheduling comprises: feeding the structured state vector as input into the reinforcement learning policy network; and performing forward calculation by the policy network to output the action parameter representing a resource scheduling decision intention, the action parameter comprising a probability distribution of selecting a specific resource and a numerical value used to define resource configuration.
[0101] Further, the obtaining process of the cross-cloud resource scheduling instruction sequence comprises: initializing the autoregressive decoding process and feeding the action parameter as input; in each decoding step, predicting and determining a next action component part by the autoregressive model according to a current decoding state and a previously generated sequence part; appending the determined component part to the current action sequence; and repeating the decoding step until the cross-cloud resource scheduling instruction sequence containing complete scheduling instructions is output.
[0102] Further, the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighted combination of multiple basic indexes of multiple heterogeneous cloud platforms; and the multiple basic indexes comprise resource cost indexes, application performance indexes, SLA compliance indexes and resource utilization efficiency indexes.
[0103] Although the embodiments of the present application have been shown and described, it is to be understood that various changes, modifications, substitutions and alterations can be made to the embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A reinforcement learning system for resource scheduling in a multi-cloud environment, characterized in that, The method comprises the following steps: An environment monitoring module is used to collect first state performance monitoring data of a plurality of heterogeneous cloud platforms in a multi-cloud environment in real time; A state representation module is used to construct a dynamic graph structure based on resource entities and to-be-processed workloads based on the monitoring data, and to generate a structured state vector by analyzing the graph structure using a graph neural network; the nodes of the graph structure are used to represent the entity representations of virtual machine instances, container instances, database service instances, storage volumes, and to-be-processed workloads; the edges of the graph structure are used to represent the relationships between the entities; the types of the relationships include network connection relationships between resource entities, deployment dependency relationships between workloads and resources, running dependency relationships between workloads and resources, data access dependency relationships between workloads, dependency relationships between workloads, and invocation dependency relationships between services; A reinforcement learning agent module is used to analyze the structured state vector through a reinforcement learning policy network to obtain resource scheduling-based action parameters; the specific acquisition process of the resource scheduling-based action parameters includes feeding the structured state vector as input into the reinforcement learning policy network; performing forward calculation through the policy network to output the action parameters representing the resource scheduling decision intention, the action parameters including a probability distribution of selecting a specific resource and a numerical value used to define resource configuration; An action decoding module is used to obtain a cross-cloud resource scheduling instruction sequence by processing the action parameters through a self-recurrent decoding process; the acquisition process of the cross-cloud resource scheduling instruction sequence includes initializing the self-recurrent decoding process and inputting the action parameters; in each decoding step, the self-recurrent model is used to predict and determine the next action component based on the current decoding state and the previously generated sequence part; the determined component is appended to the current action sequence; and the decoding step is repeated until the cross-cloud resource scheduling instruction sequence containing complete scheduling instructions is outputted; An action execution module is used to convert the cross-cloud resource scheduling instruction sequence into application program interface instructions and submit them to a plurality of cloud platforms for execution; A policy adjustment module is used to obtain second state performance monitoring data after the execution of the plurality of heterogeneous cloud platforms, construct a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjust the reinforcement learning policy network based on the score; the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighting and combining a plurality of basic indicators of the plurality of heterogeneous cloud platforms; the plurality of basic indicators include resource cost indicators, application performance indicators, SLA compliance indicators, and resource utilization efficiency indicators. 2.The resource scheduling reinforcement learning system in a multi-cloud environment of claim 1, wherein: The first state performance monitoring data includes virtual machine instance monitoring data, container instance monitoring data, database service instance monitoring data, storage volume monitoring data, and to-be-processed workload monitoring data. 3.The system of claim 1, wherein: The obtaining process of the structured state vector comprises: inputting the dynamic graph structure and the attributes of nodes and edges thereof into the graph neural network; performing multiple rounds of message passing and node state aggregation operations by the graph neural network to generate node embedding representations of nodes in the graph which fuse their own attributes and the attributes of adjacent nodes; and aggregating the node embedding representations to obtain the structured state vector.
4. A reinforcement learning method for resource scheduling in a multi-cloud environment, characterized in that, Comprise: S1. Collecting first state performance monitoring data of multiple heterogeneous cloud platforms in a multi-cloud environment in real time; S2. Constructing a dynamic graph structure based on resource entities and workloads to be processed based on the monitoring data, and generating a structured state vector by analyzing the graph structure using a graph neural network; the nodes of the graph structure are used to represent virtual machine instances, container instances, database service instances, storage volumes, and entity representations of workloads to be processed; the edges of the graph structure are used to represent the relationships between the entities; the types of the relationships include network connection relationships between resource entities, deployment dependency relationships between workloads and resources, running dependency relationships between workloads and resources, data access dependency relationships between workloads, dependency relationships between workloads, and invocation dependency relationships between services; S3. Analyzing the structured state vector by a reinforcement learning policy network to obtain resource scheduling-based action parameters; the specific obtaining process of the resource scheduling-based action parameters comprises: feeding the structured state vector as input into the reinforcement learning policy network; performing forward calculation by the policy network to output the action parameters representing the intention of resource scheduling decisions, the action parameters including probability distribution of selecting specific resources and numerical values for defining resource configurations; S4. Processing the action parameters by a self-recurrent decoding process to obtain a cross-cloud resource scheduling instruction sequence; the obtaining process of the cross-cloud resource scheduling instruction sequence comprises: initializing the self-recurrent decoding process and inputting the action parameters; in each decoding step, using the self-recurrent model to predict and determine the next action component based on the current decoding state and the previously generated sequence part; appending the determined component to the current action sequence; and repeating the decoding step until the cross-cloud resource scheduling instruction sequence containing complete scheduling instructions is outputted; S5. Converting the cross-cloud resource scheduling instruction sequence into application program interface instructions and submitting them to multiple cloud platforms for execution; S6. Obtaining second state performance monitoring data after execution of the multiple heterogeneous cloud platforms, constructing a multi-cloud environment resource scheduling score based on the first state performance monitoring data and the second state performance monitoring data, and adjusting the reinforcement learning policy network based on the score; the multi-cloud environment resource scheduling score is a comprehensive evaluation value formed by weighted combination of multiple basic indicators of the multiple heterogeneous cloud platforms; the multiple basic indicators include resource cost indicators, application performance indicators, SLA compliance indicators, and resource utilization efficiency indicators.
Citation Information
Patent Citations
Intelligent well building safety management and control system
CN120013258A
Multi-cloud environment task-oriented graph enhancement two-level architecture A3C scheduling method and system
CN120179340A