Cloud resource scheduling method and device, computer readable storage medium and product

By constructing a fully connected permissionless graph of heterogeneous hardware resources, combined with the graph attention network and Graph Transformer model, combined with the reinforcement learning generation and scheduling strategy, the problem of single traditional cloud resource scheduling mode is solved, and more accurate and efficient heterogeneous resource scheduling is achieved.

CN120075902AActive Publication Date: 2025-05-30ZGC INSTITUTE OF UBIQUITOUS-X INNOVATION & APPLICATIONS

Patent Information

Application Number
CN202510257962.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-30
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The traditional cloud resource scheduling model is single, making it difficult to effectively manage complex heterogeneous resource topology relationships and resource dependencies, affecting the overall scheduling performance.

Method used

By constructing a fully connected, authorized undirected graph of heterogeneous hardware resources, and using the graph attention network GAT and Graph Transformer models to extract global embedding features, combined with reinforcement learning model to generate scheduling strategies.

Benefits of technology

It improves the scheduling accuracy and topological perception of multiple heterogeneous resources, provides more accurate and efficient heterogeneous resource scheduling strategies, and improves overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075902A_ABST
    Figure CN120075902A_ABST
Patent Text Reader

Abstract

The invention provides a cloud resource scheduling method and device, a computer readable storage medium and a product, and the method comprises the steps: constructing a full-connection weighted undirected graph of heterogeneous hardware resources according to the heterogeneous hardware resources in a cloud edge server; inputting the full-connection weighted undirected graph into a target network model, and obtaining global embedding features output by the target network model; the target network model is formed by cascading a GAT model with a Graph Transform model, and the GAT model is a multi-layer GAT model using a multi-head attention mechanism; inputting the global embedded features and the user scheduling intention into a reinforcement learning model, and obtaining a scheduling strategy output by the reinforcement learning model; according to the method, a more accurate and efficient heterogeneous resource scheduling strategy can be provided, so that the scheduling performance and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a cloud resource scheduling method, device, computer-readable storage medium, and product. Background Art

[0002] In the cloud computing scenario for wireless networks, Kubernetes (abbreviated as K8s, which is formed by replacing the 8 characters "ubernete" in the middle of the name with 8) has become the preferred framework in the wireless edge cloud scenario due to its flexible expansion and container orchestration and management capabilities.

[0003] K8s scheduling is the core link in the entire container orchestration system. It is responsible for reasonably allocating containers to different nodes in the cluster according to multi-dimensional factors such as resource requirements, affinity rules, and quality of service requirements. The rationality of scheduling directly determines the resource utilization rate of the system, the response speed of applications, and the overall stability. A reasonable scheduling strategy can maximize the use of computing resources, avoid resource waste, and at the same time ensure the priority operation of critical applications, guaranteeing the continuity and efficiency of the business.

[0004] In traditional cloud platforms, the default scheduling strategy of K8s mainly considers the utilization rates of two elements, namely the Central Processing Unit (CPU) and memory, and ignores other important heterogeneous resources, such as the Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), etc. The scheduling dimension is relatively single. That is to say, currently, K8s is relatively mature in scheduling conventional resources such as CPU and memory, but when dealing with heterogeneous resources such as GPU, FPGA, and Data Processing Unit (DPU), its scheduling ability is limited, and it is difficult to effectively manage complex topological relationships and resource dependencies. Even if the management of these resources can be achieved through extended components, for the scheduling of heterogeneous resources, the scheduling mode is single, and there is a lack of attention to the topological connection relationships and connection performance indicators between heterogeneous resources. It can be seen that traditional schedulers still lack the optimization of multi-objective and multi-dimensional scheduling, lack the precise scheduling ability at a finer-grained resource management level for multiple heterogeneous resource clusters, and the lack of perception of topological connection relationships also makes this type of method difficult to effectively adjust scheduling weights, thus affecting the overall performance.

[0005] Summary of the Invention

[0006] The purpose of the embodiments of this application is to provide a cloud resource scheduling method, device, computer-readable storage medium, and product to solve the problem that the traditional cloud resource scheduling mode is single, thus affecting the overall scheduling performance.

[0007] To solve the above problems, an embodiment of the present application provides a cloud resource scheduling method, including:

[0008] Construct a fully connected weighted undirected graph of the heterogeneous hardware resources according to the heterogeneous hardware resources in the cloud edge server;

[0009] Input the fully connected weighted undirected graph into a target network model to obtain global embedding features output by the target network model; the target network model is formed by cascading a Graph Attention Network (GAT) model and a Graph Transformer model, where the GAT model is a multi-layer GAT model using a multi-head attention mechanism;

[0010] Input the global embedding features and the user scheduling intention into a reinforcement learning model to obtain a scheduling policy output by the reinforcement learning model.

[0011] Among them, constructing the fully connected weighted undirected graph of the heterogeneous hardware resources according to the heterogeneous hardware resources in the cloud edge server includes:

[0012] Abstract each heterogeneous hardware resource in the cloud edge server as a node in the fully connected weighted undirected graph;

[0013] Abstract the connection information between the heterogeneous hardware resources as an edge in the fully connected weighted undirected graph;

[0014] Construct the fully connected weighted undirected graph according to the feature representations of each node and the feature representations of each edge.

[0015] Among them, constructing the fully connected weighted undirected graph according to the feature representations of each node and the feature representations of each edge includes:

[0016] Construct feature vectors of each node and feature vectors of each edge;

[0017] Construct the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge.

[0018] Among them, constructing the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge includes:

[0019] Dynamically standardize the feature vectors of each node to obtain a node feature matrix of the fully connected weighted undirected graph;

[0020] Dynamically standardize the feature vectors of each edge to obtain an edge feature matrix of the fully connected weighted undirected graph;

[0021] Construct the fully-connected weighted undirected graph according to the node feature matrix and the edge feature matrix.

[0022] Among them, the feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resource corresponding to the node is included in scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resource corresponding to the node.

[0023] And / or

[0024] The feature vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in scheduling, and the fourth element is used to indicate the topological connection type of the edge.

[0025] Among them, the feature vector of the node further includes: a fifth element, and the fifth element is used to indicate other quantifiable features of the heterogeneous hardware resource corresponding to the node.

[0026] And / or

[0027] The feature vector of the edge further includes: a sixth element, and the sixth element is used to indicate other quantifiable features of the edge.

[0028] Among them, inputting the fully-connected weighted undirected graph into the target network model to obtain the global embedding features output by the target network model includes:

[0029] Using a multi-layer GAT model, dynamically calculate the attention weights of each edge of the fully-connected weighted undirected graph through the multi-head attention mechanism, and layer by layer update the features of each node of the fully-connected weighted undirected graph to obtain the multi-dimensional embedding features of each node.

[0030] Input the multi-dimensional embedding features of each node into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

[0031] Among them, the using a multi-layer GAT model, dynamically calculate the attention weights of each edge of the fully-connected weighted undirected graph through the multi-head attention mechanism, and layer by layer update the features of each node of the fully-connected weighted undirected graph to obtain the multi-dimensional embedding features of each node includes:

[0032] For each node in the fully-connected weighted undirected graph, respectively perform a first operation to obtain the embedding feature of the node; wherein, the first operation includes:

[0033] Use the first attention weight to perform a weighted sum of the features of the neighbor nodes of the node to obtain the first feature of the node; the first attention weight corresponds to the first attention head of the multi-head attention mechanism.

[0034] Concatenate the result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism with the first feature of the corresponding node to obtain the second feature of the node;

[0035] Apply a non-linear activation function after merging and averaging the second feature to obtain the multi-dimensional embedding feature of the node.

[0036] Among them, inputting the multi-dimensional embedding feature of each node into the Graph Transformer model to obtain the global embedding feature output by the Graph Transformer model includes:

[0037] According to the attention scores of the Graph Transformer model, perform global information aggregation on the multi-dimensional embedding features of each node to obtain the target embedding feature of each node;

[0038] Aggregate the target embedding features of the nodes through global average pooling to generate the global embedding feature.

[0039] Among them, the method further includes:

[0040] In the reinforcement learning model, calculate the reward information according to the influence of the action selected by the policy network on the environment update and perform backpropagation to update the weights of the policy network;

[0041] Among them, the environment includes at least one of the following: environment state, connection state, task arrival and completion state;

[0042] The reward information includes at least one of the following: resource utilization rate, task completion time, and energy consumption.

[0043] An embodiment of the present application further provides a scheduling device, including a processor and a transceiver. The transceiver receives and sends data under the control of the processor, and the processor is used to perform the following operations:

[0044] Construct a fully connected weighted undirected graph of the heterogeneous hardware resources according to the heterogeneous hardware resources in the cloud-edge server;

[0045] Input the fully connected weighted undirected graph into the target network model to obtain the global embedding feature output by the target network model; the target network model is formed by cascading a graph attention network (GAT) model and a graph transformation (Graph Transformer) model, where the GAT model is a multi-layer GAT model using the multi-head attention mechanism;

[0046] Input the global embedding feature and the user scheduling intention into the reinforcement learning model to obtain the scheduling policy output by the reinforcement learning model.

[0047] Wherein, the processor is further configured to perform the following operations:

[0048] Abstract each heterogeneous hardware resource in the cloud edge server as a node in the fully connected weighted undirected graph;

[0049] Abstract the connection information between the heterogeneous hardware resources as an edge in the fully connected weighted undirected graph;

[0050] Construct the fully connected weighted undirected graph according to the feature representations of each node and the feature representations of each edge.

[0051] Wherein, the processor is further configured to perform the following operations:

[0052] Construct a feature vector for each node and a feature vector for each edge;

[0053] Construct the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge.

[0054] Wherein, the processor is further configured to perform the following operations:

[0055] Dynamically normalize the feature vectors of each node to obtain the node feature matrix of the fully connected weighted undirected graph;

[0056] Dynamically normalize the feature vectors of each edge to obtain the edge feature matrix of the fully connected weighted undirected graph;

[0057] Construct the fully connected weighted undirected graph according to the node feature matrix and the edge feature matrix.

[0058] Wherein, the feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resource corresponding to the node is included in scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resource corresponding to the node;

[0059] And / or,

[0060] The feature vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in scheduling, and the fourth element is used to indicate the topological connection type of the edge.

[0061] Wherein, the feature vector of the node further includes: a fifth element, and the fifth element is used to indicate other quantifiable features of the heterogeneous hardware resource corresponding to the node;

[0062] And / or,

[0063] The eigenvector of the edge further includes: a sixth element for indicating other quantifiable features of the edge.

[0064] Wherein, the processor is further configured to perform the following operations:

[0065] Using a multi-layer GAT model, dynamically calculate the attention weights of each edge of the fully-connected weighted undirected graph through the multi-head attention mechanism, and layer by layer update the features of each node of the fully-connected weighted undirected graph to obtain the multi-dimensional embedding features of each node;

[0066] Input the multi-dimensional embedding features of each node into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

[0067] Wherein, the processor is further configured to perform the following operations:

[0068] For each node in the fully-connected weighted undirected graph, perform a first operation respectively to obtain the embedding features of the node; wherein, the first operation includes:

[0069] Use the first attention weight to perform a weighted sum of the features of the neighbor nodes of the node to obtain the first feature of the node; the first attention weight corresponds to the first attention head of the multi-head attention mechanism;

[0070] Concatenate the result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism with the first feature of the corresponding node to obtain the second feature of the node;

[0071] Apply a non-linear activation function after merging and averaging the second feature to obtain the multi-dimensional embedding features of the node.

[0072] Wherein, the processor is further configured to perform the following operations:

[0073] According to the attention scores of the Graph Transformer model, perform global information aggregation on the multi-dimensional embedding features of each node to obtain the target embedding features of each node;

[0074] Aggregate the target embedding features of the node through global average pooling to generate the global embedding features.

[0075] Wherein, the processor is further configured to perform the following operations:

[0076] In the reinforcement learning model, calculate the reward information according to the influence of the actions selected by the policy network on the update of the environment and perform backpropagation to update the weights of the policy network;

[0077] Among them, the environment includes at least one of the following: environmental status, connection status, task arrival and completion status;

[0078] The reward information includes at least one of the following: resource utilization rate, task completion time, and energy consumption.

[0079] An embodiment of the present application further provides a scheduling device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the cloud resource scheduling method described above is implemented.

[0080] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the cloud resource scheduling method described above are implemented.

[0081] An embodiment of the present application further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps of the cloud resource scheduling method described above are implemented.

[0082] The above technical solution of the present application has at least the following beneficial effects:

[0083] In the cloud resource scheduling method, device, computer-readable storage medium, and product of the embodiment of the present application, a fully connected weighted undirected graph is constructed according to the heterogeneous hardware resources in the cloud edge server to improve the scheduling accuracy of various heterogeneous resources and the perception ability of the topological structure; and the global embedding features of the fully connected weighted undirected graph are extracted by using a multi-head attention mechanism-based multi-layer GAT model cascaded with the GraphTransformer model, which can not only more finely embed local and global graph structure information, but also more accurately model heterogeneous resource nodes and their topological relationships by dynamically adjusting the weights of nodes and edges; finally, the global embedding features and the user scheduling intention are input into the reinforcement learning model to obtain the scheduling strategy, so as to provide a more accurate and efficient heterogeneous resource scheduling strategy. Description of the Drawings

[0084] Figure 1 It represents the flowchart of the steps of the cloud resource scheduling method provided by the embodiment of the present application;

[0085] Figure 2 It represents the overall framework to which the cloud resource scheduling method provided by the embodiment of the present application is applied;

[0086] Figure 3 It represents an example of the fully connected weighted undirected graph provided by the embodiment of the present application;

[0087] Figure 4 It represents an example of the multi-layer GAT model using the multi-head attention mechanism in the cloud resource scheduling method provided by the embodiment of the present application;

[0088] Figure 5 An example showing the CPU utilization of each node in a K8s (Kubernetes) cluster;

[0089] Figure 6 An example showing the memory utilization of each node in a K8s (Kubernetes) cluster;

[0090] Figure 7 An example showing the disk utilization of each node in a K8s (Kubernetes) cluster;

[0091] Figure 8 Another example showing the CPU utilization of each node in a K8s (Kubernetes) cluster;

[0092] Figure 9 An example showing the Pod deployment time of each scenario in a K8s (Kubernetes) cluster;

[0093] Figure 10 An example showing the maximum number of Pod deployments for each scenario in a K8s (Kubernetes) cluster;

[0094] Figure 11 A schematic diagram showing the structure of the scheduling device provided in the embodiments of the present application. Detailed implementation manners

[0095] To make the technical problems, technical solutions, and advantages to be solved by the present application clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0096] As Figure 1 shown, the embodiments of the present application provide a cloud resource scheduling method, including:

[0097] Step 101, construct a fully connected weighted undirected graph of the heterogeneous hardware resources according to the heterogeneous hardware resources in the cloud edge server.

[0098] Optionally, the heterogeneous hardware resources in the cloud edge server can also be referred to as heterogeneous resource nodes. In this step, a fully connected weighted undirected graph is constructed for the heterogeneous resource nodes and their topological connection relationships, and more fine-grained multiple performance metrics are used to encode and represent different types of heterogeneous resource nodes in order to enhance the precise scheduling ability, so as to improve the scheduling accuracy of multiple heterogeneous resources and the perception ability of the topological structure. Compared with the traditional scheduling method, this solution can better support resource diversity and dynamic changes, and effectively improve the scheduling efficiency, resource utilization rate, and the overall performance of complex applications.

[0099] Step 102: Input the fully connected weighted undirected graph into the target network model to obtain the global embedding features output by the target network model. The target network model is formed by cascading a Graph Attention Network (GAT) model and a Graph Transformer model. The GAT model is a multi-layer GAT model using the multi-head attention mechanism.

[0100] Optionally, the target model proposed in the embodiment of this application, which is a multi-layer GAT using multi-head attention cascaded with a Graph Transformer, can not only more meticulously embed local and global graph structure information, but also more precisely model heterogeneous resource nodes and their topological relationships by dynamically adjusting the weights of nodes and edges. The global feature modeling of the Graph Transformer can make up for the limitations of the GAT in the local neighborhood range, thereby generating embedding features (also called embedding representations) that combine global and local information. In the scenario of combining local and global, using a multi-layer GAT cascaded with a Graph Transformer and a fully connected weighted undirected graph can effectively aggregate useful information related to the target node at both the global and local levels, thus providing a precise and efficient feature embedding representation for complex heterogeneous resource scheduling, enabling the scheduling system to effectively handle the multi-level and multi-distribution dependency relationships among different nodes, edges, and the entire graph.

[0101] Step 103: Input the global embedding features and the user's scheduling intention into the reinforcement learning model to obtain the scheduling policy output by the reinforcement learning model.

[0102] Optionally, input the global embedding features output by the target model and the user's scheduling intention into a Reinforcement Learning (RL) model. Based on the resource requirements and priorities input by the user, dynamically adjust the scheduling policy through a reward mechanism, and perform adaptive adjustment learning for multi-objective heterogeneous resource optimization scheduling such as efficient utilization of heterogeneous resources under balanced load, rapid deployment of Pods, and maximization of the number of deployed Pods, to provide a more precise and efficient heterogeneous resource scheduling policy.

[0103] In summary, in view of the resource characteristics of the wireless edge cloud scenario, the embodiments of the present application construct a fully connected weighted undirected graph for heterogeneous resource nodes and the topological connection relationships between heterogeneous resource nodes, and use a splicing network structure of multi-layer GAT with multi-head attention cascaded with Graph Transformer, which enhances the embedding representation ability of the traditional Graph Neural Network (GNN) on multi-dimensional features of heterogeneous resource nodes and makes up for the deficiency of the traditional GNN in topological feature expression. Among them, in order to better obtain the embedding representation of the graph and the ability to better allocate weights in the above-mentioned weighted undirected graph, multi-layer GAT with multi-head attention is used to perform multi-layer feature embedding representation on heterogeneous resource nodes and topological connection relationships. Then, the feature embedding representation is input into the RL framework. Based on the resource requirements and priorities input by the user, the learning weights are adaptively adjusted through a reward mechanism, so as to dynamically adjust the scheduling optimization strategy, enabling the scheduling system to effectively handle the multi-level and multi-distribution dependency relationships between different nodes, edges, and the entire graph, realizing the multi-objective optimization goals of more accurate scheduling optimization and high-efficiency resource utilization of CPU, memory, disk, and GPU under balanced load, reducing the Pod deployment time, and increasing the maximum number of Pod deployments, thereby improving the scheduling performance and accuracy.

[0104] In one implementation manner, the overall framework to which the cloud resource scheduling method provided by the embodiments of the present application is applied is as Figure 2 shown, specifically:

[0105] First, a fully connected weighted undirected graph is constructed for heterogeneous resource nodes and the topological connection relationships between them. More fine-grained multiple performance indicators are used to encode and represent different types of heterogeneous resource nodes in order to enhance the precise scheduling ability. Dynamic standardization is introduced for the representation of different types of heterogeneous resource nodes and topological connections, so as to construct an overall fully connected device node graph by recording the connection relationships and weights between different heterogeneous devices, and better represent various types and levels of heterogeneous resources and the connections between them. Specifically, all heterogeneous hardware (such as CPU, GPU, FPGA, DPU, memory, etc.) in the server has a topological relationship with this node. We abstract the topological relationship among them as the above-mentioned edge information, describe the relationship between devices as the relationship between points in a weighted undirected graph, and assign different topological relationship weights to different relationships. Then it is converted into a fully connected graph and input into the subsequent algorithm.

[0106] Secondly, the architecture of cascading Graph Transformer with multiple layers of GAT can not only embed local and global graph structure information more meticulously, but also achieve a more accurate modeling of heterogeneous resource nodes and their topological relationships by dynamically adjusting the weights of node and edge dependencies. Specifically, GAT is optimized on the local graph structure through the multi-head self-attention mechanism, and the attention weights of each neighbor node are adjusted through backpropagation to achieve dynamic weighted aggregation of neighborhood node information. The embedding process of multiple layers of GAT gradually extracts feature representations at different levels, enabling it to flexibly capture the complex relationships between resource nodes and connections. Then, the output of multiple layers of GAT is fed into Graph Transformer, and its global attention mechanism is used to further aggregate information from the entire graph, enhancing the representation ability of the global context. The global feature modeling of Graph Transformer can make up for the limitations of GAT in the local neighborhood range, thus generating an embedding representation that combines global and local information. The network structure of cascading Graph Transformer with multiple layers of GAT using multi-head attention not only improves the fine-grained performance of the model in local node interactions, but also expands the model's understanding ability of all graph resource nodes and topological connections.

[0107] Finally, although the embedding representation after cascading Graph Transformer with multiple layers of GAT using multi-head attention can efficiently capture the dependency relationships and interaction patterns between local and global resource nodes and thus dynamically adjust the weights of their complex connection relationships, in the user scheduling intention, there may be a higher emphasis or priority on specific resources, and the specific priority requirements and weight updates of multiple optimization objectives in the scheduling intention are mainly carried out in the RL framework part using the policy network. Combining the policy network with the reinforcement learning (RL) framework, the global embedding features are input into the RL framework of the multi-policy network. Based on the emphasis on resource type and priority requirements input by the user, multiple optimization objectives (such as resource utilization rate, load balancing, task completion time, etc.) are weighed and dynamically adjusted through the reward mechanism, providing a more accurate and efficient heterogeneous resource scheduling scheme. This scheme can better meet the multi-level and multi-type heterogeneous resource scheduling requirements and achieve efficient allocation and optimization of resources under complex topological structures.

[0108] In some embodiments, the advantage of using a multi-layer GAT cascaded Graph Transformer network with multi-head attention to embed and represent resource and topology features in both local and global scopes and then using an RL framework based on the policy network PPO to learn and generate an optimized resource scheduling strategy is that it can automatically learn the important relationships among nodes, edges, and the global in local and global graph structure data through an adaptive multi-head attention mechanism. In the absence of explicit priority information, through the reward feedback of reinforcement learning and the weight update of the attention mechanism, the model gradually learns to optimize the resource priorities and importance of the scheduling strategy, and the model has higher flexibility and generalization ability. It can dynamically adjust the attention weights in various different scheduling intents and resource environments, and then generate a better resource scheduling strategy.

[0109] In at least one embodiment of the present application, step 101 includes:

[0110] Abstract each heterogeneous hardware resource in the cloud-edge server as a node in the fully connected weighted undirected graph;

[0111] Abstract the connection information between the heterogeneous hardware resources as an edge in the fully connected weighted undirected graph;

[0112] Construct the fully connected weighted undirected graph according to the feature representations of each node and the feature representations of each edge.

[0113] In one implementation, abstract the heterogeneous hardware resources in a server as a node set V = {v 1 , v 2 , …, v n} in the fully connected weighted undirected graph. All heterogeneous hardware resources (such as CPUs, GPUs, FPGAs, memories, network devices, etc.) have topological relationships with the remaining nodes. Since the relationships between heterogeneous hardware resources are multi-path, the optimization strategy will convert the multi-path into a single-path weight to achieve the generation of the weight path between points. As Figure 3 shown.

[0114] In another implementation, abstract the connection information between the heterogeneous hardware resources as an edge set E = {e ij |(v i , v j) ∈ V}, where the performance metrics include information such as connection loss, latency, throughput, etc. In addition, the connection topology between various hardware resources needs to be considered. The heterogeneous hardware within the server is described as a weighted undirected graph data structure. The server itself is abstracted as a node in the graph, and all heterogeneous hardware within the server (such as CPUs, GPUs, FPGAs, memory, network devices, etc.) has a topological relationship with this node. Based on the heterogeneous hardware resources within the server and their topological relationships, and according to the established policies of the cloud platform, the weights of these topological relationships are determined, and finally this information is stored in the form of an adjacency list.

[0115] The overall input data is a list of nodes and the required resources. First, all heterogeneous hardware resources within the server are scanned, and then the topological relationships (such as the topological connections mentioned above), transmission paths, etc. between the hardware are analyzed and determined. Then, through the policies of the cloud platform, corresponding initial weight values are assigned to the heterogeneous resources. The device resources are regarded as nodes of a weighted undirected graph, and the connections between heterogeneous resources are regarded as weighted edges. Based on this, a fully connected weighted undirected graph of heterogeneous hardware resources is constructed.

[0116] In at least one embodiment of the present application, constructing the fully connected weighted undirected graph according to the feature representations of each node and each edge includes:

[0117] Constructing the feature vectors of each node and the feature vectors of each edge;

[0118] Constructing the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge.

[0119] Optionally, constructing the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge includes:

[0120] Performing dynamic normalization on the feature vectors of each node to obtain the node feature matrix of the fully connected weighted undirected graph;

[0121] Performing dynamic normalization on the feature vectors of each edge to obtain the edge feature matrix of the fully connected weighted undirected graph;

[0122] Constructing the fully connected weighted undirected graph according to the node feature matrix and the edge feature matrix.

[0123] Different heterogeneous hardware resources can be uniformly represented through feature engineering. A set of common features can be defined for each type of heterogeneous hardware resource, and the same feature vector representation is used for all hardware types. The node features should cover the core information and current state of each resource. After the respective feature vectors of various types of heterogeneous resources are dynamically normalized, they are unified into a common feature space and combined together to form a node feature matrix.

[0124] In one implementation, the feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resources corresponding to the node are included in scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resources corresponding to the node.

[0125] Optionally, the feature vector of the node further includes: a fifth element, and the fifth element is used to indicate other quantifiable features of the heterogeneous hardware resources corresponding to the node. For example, the other quantifiable features include at least one of the following: performance score, importance score, etc.

[0126] For example, each node has a feature vector v i =[e i , u i , p i · , including multiple performance metrics of heterogeneous hardware resources, presence or absence, resource occupancy rate, etc.

[0127] Among them, e i ∈{0, 1}: indicates whether the resource vi is included in scheduling (0 means not included in scheduling, 1 means included in scheduling).

[0128] u i ∈[0, 1]: indicates the occupancy rate of the resource v i (0 means 100% occupancy rate).

[0129] p i ∈[0.1, 1]: indicates a certain performance score of the resource v i (dynamically standardized and extensible).

[0130] In another implementation, the feature vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in scheduling, and the fourth element is used to indicate the topological connection type of the edge.

[0131] Optionally, the feature vector of the edge further includes: a sixth element, and the sixth element is used to indicate other quantifiable features of the edge. For example, the other quantifiable features include at least one of the following: performance score, importance score, etc. Each edge e ij has a feature vector e ij , defined as: e ij =[c ij , t ij , s ij · .

[0132] Among them, c ij ∈{0, 1}: indicates the edge e​​ij Whether to include in scheduling (0 means not included in scheduling, 1 means included in scheduling).

[0133] t ij ∈ {[1, 0, 0], [0, 1, 0], [0, 0, 1]}: One-hot encoding represents edge e ij 's topological connection type. Suppose there are three connection methods. For example, [1, 0, 0], [0, 1, 0], and [0, 0, 1] respectively represent: PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) and NVLink (a bus and its communication protocol developed and launched by NVIDIA), NUMA (NonUniform Memory Access, a non-uniform memory access architecture).

[0134] s ij ∈ [0.1, 1]: Represents a certain performance score of edge e ij (dynamically standardized and extensible).

[0135] In one implementation, the formula for dynamic standardization is as follows:

[0136]

[0137] Among them, α and β are two adjustment parameters. α is used to adjust the range of the standardized result, and β is used to prevent the result from being 0. For example, setting α = 0.9 and β = 0.1 compresses the standardized result to between [0.1, 1.0] to avoid a score of 0. x is the object of dynamic standardization; for example, any element of the feature vector of the above-mentioned node or the feature vector of the edge.

[0138] In the embodiment of the present application, step 102 includes:

[0139] Using a multi-layer GAT model, dynamically calculate the attention weights of each edge of the fully connected weighted undirected graph through the multi-head attention mechanism, and layer by layer update the features of each node of the fully connected weighted undirected graph to obtain the multi-dimensional embedding features of each node;

[0140] Input the multi-dimensional embedding features of each node into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

[0141] Among them, the GAT self-attention mechanism of each layer can still dynamically adjust the edge weights between nodes in the fully connected weighted undirected graph and dynamically evaluate the relative importance between resource nodes. The multi-layer GAT can enable each layer to focus on different relationship weights or feature interactions, extract and refine the features of each resource node layer by layer, gradually update the node features and weight distributions at each layer, making the node embedding representations more complex layer by layer, capturing the implicit hierarchical relationships between nodes, understanding their potential roles and impacts in the topological structure, and thus effectively reflecting the differences and connections between nodes hierarchically.

[0142] In one implementation, the multi-layer GAT model is used to dynamically calculate the attention weights of each edge of the fully connected weighted undirected graph through the multi-head attention mechanism and update the features of each node of the fully connected weighted undirected graph layer by layer to obtain the multi-dimensional embedding features of each node, including:

[0143] For each node in the fully connected weighted undirected graph, a first operation is respectively performed to obtain the embedding feature of the node; wherein, the first operation includes:

[0144] Using the first attention weight to perform a weighted sum of the features of the neighbor nodes of the node to obtain the first feature of the node; the first attention weight corresponds to the first attention head of the multi-head attention mechanism;

[0145] Concatenating the result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism with the first feature of the corresponding node to obtain the second feature of the node;

[0146] Applying a non-linear activation function after performing a combined average on the second feature to obtain the multi-dimensional embedding feature of the node.

[0147] Using the multi-layer GAT model, dynamically calculate the attention weights of each edge through the multi-head attention mechanism, and update the node features layer by layer. For the GAT operation of each layer l, the embedding representation h of node v i of i (l) Using the first attention weight a ij (l,k) Perform a weighted sum of the neighbor node features to update the feature of node i to obtain the first feature (where the feature of node i in the first layer is itself):

[0148]

[0149] wherein, N(v i ) represents node v iThe set of neighbor nodes, W(l,k) is the weight matrix of the linear transformation node features for the k-th head in the l-th layer, and σ is a non-linear activation function (such as LeakyReLU); a ij (l,k) represents the first attention weight.

[0150] Concatenate the result features of the multi-head attention to obtain the second feature:

[0151]

[0152] Then, use the method of merging and averaging and applying a non-linear activation function to obtain the updated representation of the node, that is, the multi-dimensional embedding features of the node:

[0153]

[0154] where K is the number of attention heads, and || represents the vector concatenation operation.

[0155] To increase the expressive power of the multi-layer GAT model, the multi-head attention mechanism is usually used. For each node, multiple attention heads can focus on different feature relationships, such as the power consumption, load, performance, etc. of the node. For the heterogeneous resource cluster scenario, the multi-head attention can capture multiple feature interactions, emphasize the feature weights of different resource dimensions, and provide more comprehensive and fine-grained feature aggregation for each layer of GAT. This mechanism is more effective in capturing the different characteristics of heterogeneous resource nodes, such as Figure 4 shown.

[0156] The multi-head attention mechanism is used to capture different feature relationships between resource nodes, and also affects the attention weights through the edge feature e ij to capture characteristics such as different connection methods (such as MUNA, PCIe, NVLINK) and latency. Assume that the initial feature of each node is represented by the vector h i (0) ∈R d For each pair of connected nodes (i,j) ∈ E in each layer l, there are a total of K attention heads to calculate their respective attention weights from different dimensions. For the attention weight a ij (l,k) of each head k, the calculation formula is as follows:

[0157]

[0158] where, W (l,k) ∈R d×d is the weight matrix representing the linear transformation of the node features for the k-th head in the l-th layer; is the weight matrix representing the linear transformation of the edge features, which transforms features such as the connection method and latency of the edge to the same dimension as the node features; are the trainable parameters of the attention mechanism; N(v i ) represents the set of neighbor nodes of node v i . The || represents the vector concatenation operation.

[0159] Multi-head attention and multi-layer GAT play different roles in the graph structure. In a fully connected weighted undirected graph, the collaboration of the two can further enhance the detailed expression of resource features.

[0160] The design intention of the multi-head attention mechanism is to focus on the relationships between nodes from different angles or perspectives. Each layer can use the multi-head attention mechanism to conduct a sub-evaluation of node relationships. Different attention heads can focus on different feature dimensions, such as CPU utilization, memory occupancy, etc. This enables the multi-layer GAT to effectively encode specific dimensions at each layer and finally form a more comprehensive global embedding. This embedding can better reflect the feature differences between different resource nodes, thereby providing a more fine-grained weight allocation in the scheduling strategy.

[0161] In each layer of GAT, the outputs of each attention head will be integrated together to form the node representation at the current level, gathering the different perspective information generated by the multi-heads together to generate richer node features. The design of the multi-layer GAT can gradually build the global feature representation of the nodes in a progressive manner, and integrate the features of each layer and apply them progressively to the next layer.

[0162] The multi-head attention mechanism focuses on refining feature capture and multi-perspective fusion within the same layer, while the multi-layer GAT focuses on further aggregating and adjusting these fused node features at deeper levels. This hierarchical integration process helps to gradually strengthen or weaken certain features, constructing a hierarchical-dependent embedding representation for the nodes in the fully connected weighted undirected graph. In each layer of GAT, the output of the attention head will focus on specific feature relationships, and in deeper layers, the GAT can gradually reconstruct these features to ensure a more globalized feature embedding in the overall scheduling requirements.

[0163] In at least one embodiment of the present application, inputting the multi-dimensional embedding features of each node into the GraphTransformer model to obtain the global embedding features output by the Graph Transformer model includes:

[0164] Performing global information aggregation on the multi-dimensional embedding features of each node according to the attention scores of the Graph Transformer model to obtain the target embedding features of each node;

[0165] Aggregating the target embedding features of the nodes through global average pooling to generate the global embedding features.

[0166] In the embodiments of the present application, the output embedding of the multi-layer GAT using multi-head attention is used as the input of the Graph Transformer. The Graph Transformer further models the global graph structure through more detailed local feature aggregation processed by the GAT and its global self-attention mechanism, pays attention to the relative importance of nodes in the whole graph, generates a feature representation reflecting the overall graph structure for each node, helps to identify the nodes and resource combinations crucial in large-scale scheduling, enables the network structure of the multi-layer GAT cascaded with the Graph Transformer to more effectively focus on important node relationships, improves the global optimization ability, and helps the policy network effectively learn complex global scheduling patterns.

[0167] Specifically, the multi-dimensional embedded features h of each node after passing through the multi-layer GAT i (L) (i.e., h' in the above embodiments i ′ (l) ) are used as the input of the Graph Transformer. In the original Graph Transformer structure, the linear transformation matrices W Q , W K , W V are used. To enable the Graph Transformer to have stronger non-linear feature expression ability, so as to better model the global graph feature relationships, more fully capture the complex global dependencies between nodes, and better process heterogeneous graphs or multi-modal data, a multi-layer perceptron MLP is used to enhance the expression ability. Specifically:

[0168] The query vector The key vector The value vector

[0169] are independent multi-layer perceptrons, and the specific formula is:

[0170]

[0171] where, W Q1 , W Q2 , W K1 , W K2 , W V1 , W V2 are weight matrices;

[0172] b Q1 , b Q2 , b K1 , b K2 , b V1 , b V2 are biases;

[0173] σ is an activation function (such as ReLU).

[0174] Attention score β in Graph Transformer ij is calculated as follows:

[0175]

[0176] where MLP Q and MLP K are the query and key representations output by the multi-layer perceptron module, and d k is a scaling factor used to prevent numerical values from being too large.

[0177] After obtaining the attention score β ij Graph Transformer performs global information aggregation on the entire graph, and the final embedding of the node is updated as:

[0178]

[0179] where MLP V is the value representation output by the multi-layer perceptron module. The final embedding includes the global information aggregation of each node, which helps to capture the relationships of each node in the entire graph.

[0180] The global embedding can provide the overall state of the entire system, which helps the policy network to understand the global relationships and overall performance among resources; the policy network can combine the global embedding with the local node embeddings to make more comprehensive and optimized scheduling decisions; through the global embedding, the policy network can consider the synergy effects among resources and achieve an improvement in overall performance. In complex resource scheduling scenarios, such as when the scheduling tasks involve global resource balance and multi-dimensional performance optimization, the global embedding can significantly improve the decision-making quality of the policy network.

[0181] To provide a global perspective for the policy network, all node embeddings h i (L) are aggregated through global average pooling to generate the global embedding feature h G . The global average pooling is as follows:

[0182]

[0183] where L represents the total number of GAT layers.

[0184] In at least one embodiment of the present application, the global embedding feature h GCombined with the user's scheduling intention, it serves as the state input of the reinforcement learning model. In the reinforcement learning model (which can also be called the reinforcement learning framework), the environment E defines the external system with which the agent interacts. In the scenario of the optimal scheduling strategy for multi-heterogeneous resources, the environment includes resource status, connection status, task arrival and completion, etc. Under the reinforcement learning framework, the policy network optimizes the parameter θ through interaction with the environment to maximize the cumulative reward.

[0185] The embodiments of this application use the global embedding feature representations of nodes, edges, and global graphs output by the target model as the state input of the reinforcement learning model, combined with the user's scheduling intention. Define the action space to select which nodes meet the resource requirements; design the reward mechanism so that the greater the reward under the policy that meets the given priority; optimize the policy, and update the weights of the policy network through backpropagation and gradient descent methods. Finally, generate a preliminary optimal resource scheduling plan according to the policy output of the reinforcement learning.

[0186] Optionally, the embodiments of this application calculate the reward signal according to the scheduling result and perform backpropagation to update the weights of the policy network. Repeat this process to gradually optimize the scheduling policy. Through multiple iterations of optimization and training, the model can continuously improve the scheduling policy and gradually approach the optimal solution.

[0187] In one implementation, the method further includes:

[0188] In the reinforcement learning model, calculate the reward information according to the impact of the action selected by the policy network on the update of the environment, and perform backpropagation to update the weights of the policy network;

[0189] Wherein, the environment includes at least one of the following: environment status, connection status, task arrival and completion status;

[0190] The reward information includes at least one of the following: resource utilization rate, task completion time, and energy consumption.

[0191] The embodiments of this application calculate the reward signal according to the impact of the action selected by the policy network on the update of the environment status, and update the policy by feeding back the reward from the environment status, that is, the scheduling result. Calculate the reward signal and perform backpropagation to update the weights of the policy network. Repeat this process to gradually optimize the scheduling policy. Through multiple iterations and training, the model continuously improves and optimizes the scheduling policy and gradually approaches the optimal solution.

[0192] In one implementation, the elements of reinforcement learning are as follows:

[0193] State space: The current resource status and user requirements, denoted as S t =(h G , u).

[0194] Action space: Select a resource scheduling strategy that meets resource requirements, denoted as A t = a.

[0195] Reward mechanism: The feedback signal after executing an action, reflecting the scheduling effect, denoted as R t .

[0196] Policy network: Defined as πθ(A t |S t ).

[0197] In one implementation, the policy optimization method is as follows:

[0198] Use Proximal Policy Optimization (PPO) for policy optimization. The goal of PPO is to maximize the following objective function:

[0199]

[0200] where, represents the policy probability ratio; represents the advantage function; ∈ represents a hyperparameter, usually set to 0.2.

[0201] In one implementation, the definition of the advantage function is as follows:

[0202] The advantage function measures the superiority of the current action relative to the average action:

[0203]

[0204] where, V φ (S t ) is the estimated value of the value network (Critic) with parameters φ.

[0205] In one implementation, the multi-task learning and multi-objective optimization strategy is as follows:

[0206] Reward function R t (S t ,A t ) reflects multiple objectives of resource scheduling, and the specific design is as follows:

[0207] R t (S t ,A t ) = w 1 R 1 + w 2 R 2 - w 3 R 3 ; where:

[0208] Resource utilization rate R1 : Calculate the proportion of the allocated resources that are actually used to measure the efficiency of resource allocation.

[0209] Task completion time R 2 : The time required for a task to be completed from allocation, used to measure the scheduling speed.

[0210] Energy consumption R 3 : Calculate the total power consumption of all resources for executing the current task. Optimizing energy consumption can help extend the hardware lifespan and reduce the operating cost.

[0211] w 1 , w 2 , w 3 : Weight coefficient, the weight that is automatically learned and adjusted in the policy network according to the user's resource requirements and priorities, ensuring that the policy network can automatically adjust the target weight according to the specific requirements of different scheduling tasks.

[0212] In one implementation, the training steps of PPO are as follows:

[0213] Collect samples: Interact with the environment E according to the current policy π θ to collect data such as states, actions, and rewards.

[0214] Calculate the advantage function: Calculate the advantage function using the value network

[0215] Update the policy network: Update the policy network parameters θ by maximizing L CLIP (θ).

[0216] Update the value network: Minimize the following loss function to update the value network parameters φ:

[0217] L(φ) = E t [(R t - V φ (S t )) 2

[0218] Iterative training: In each iteration, the policy network and the value network are updated according to the reward feedback of the global embedding, and the scheduling scheme is gradually optimized by increasing the cumulative reward.

[0219] Iteration termination: The termination conditions can be set for the training iteration, such as: if the cumulative reward reaches the set threshold, the system optimization goal is achieved and the iteration can be terminated in advance; set the maximum number of iterations to avoid the risk of overfitting caused by overtraining.

[0220] In one implementation, the simulation results and evaluations are as follows;

[0221] ​For the GAT model provided in the embodiments of this application, after cascading the Graph Transformer model to extract features and generate global embedding features, a resource optimization scheduling policy learned by a reinforcement learning model based on a policy network is used for scheduling performance testing. The experiment customizes a stress test image according to resource pressure software, and for different resource requirements, container images with different resource requirements are customized to simulate container tasks in actual production. Optionally, this example mainly designs three groups of experiments from three aspects: resource utilization rate, Pod deployment time, and number of Pod deployments, and conducts comparative analysis on the data after the experiment.

[0222] To verify the effectiveness and feasibility of the proposed solution, a K8s cluster is built, which includes one control node (Master1) and five worker nodes (Node1 - 5). The total resource information of each node is shown in Table 1.

[0223] Table 1 K8s Resource Node Information List

[0224]

[0225]

[0226] Build T Pod applications, whose resource requirements simulate various different resource-intensive applications in the cloud-edge network, and their resource specifications are shown in Table 2.

[0227] Table 2 Container Resource Requirement Information

[0228]

[0229] In the embodiments of this application, the network structure of the multi-head attention mechanism-based multi-layer GAT model cascading the Graph Transformer model can further learn resource node features and topological connection relationships on the basis of the Graph Neural Network (GNN), and then can more flexibly and accurately adjust multi-dimensional resource weights, maintain the stability between various resources, and achieve load balancing and high resource utilization rate of the K8s cluster. As Figure 5 shown is the CPU utilization rate of each node in the K8s (Kubernetes) cluster, as Figure 6 shown is the memory utilization rate of each node in the K8s (Kubernetes) cluster, as Figure 7 shown is the disk utilization rate of each node in the K8s (Kubernetes) cluster.

[0230] Different from resource-exclusive applications where only one application can be deployed on a single graphics card, resource-sharing Pods are deployed on each Node. The GPU utilization rate of each Node using kubeshare and the scheduling policy based on GNN is tested in the embodiments of this application (where KCSS does not have GPU scheduling). As Figure 8 shown.

[0231] The network structure of cascading Graph Transformer with multi-head attention in multiple layers of GAT can more effectively focus on important node topological relationships, learn multi-dimensional features and topological connection structures of nodes, enhance the global optimization ability, help the policy network effectively learn complex global scheduling patterns, and thus better select nodes that achieve the best balance among multiple criteria related to the cloud infrastructure status and user requirements for Pod deployment. After using the network structure of cascading Graph Transformer with multi-head attention in multiple layers of GAT, the scheduling optimization policy learned by reinforcement learning can optimize the performance in terms of Pod deployment time. As Figure 9 shown.

[0232] Since the default scheduler of K8s scheduler uses a fixed weight to sum the scores of nodes, it cannot judge the bias of Pod applications according to actual needs, and the default deployment will cause resource overload in a certain dimension of the node, making it impossible to deploy more Pod applications. The network structure of cascading Graph Transformer with multi-head attention in multiple layers of GAT learns multi-dimensional features and topological connection structures of nodes and then uses the scheduling optimization policy learned by reinforcement learning.

[0233] When facing the deployment pressure of a large number of application containers, the multi-objective reinforcement learning scheduling optimization algorithm assisted by cascading Graph Transformer with multi-head attention in multiple layers in this paper, under the load balancing policy, considers the actual needs of Pod applications themselves, the multi-dimensional resource usage of nodes themselves, and the topological connection relationships between nodes, and adaptively adjusts multi-dimensional weights with the help of the multi-head self-attention mechanism, and then can make full use of existing cluster resources to deploy more containers. Therefore, using the solution in this paper can further expand the deployment scale of cluster Pods compared with the other three methods, and the deployment of Pod applications can also be more dispersed and balanced, and more Pod applications can be deployed in the edge collaborative computing scenario. As Figure 10 shown.

[0234] In summary, after the multi-layer GAT model using the multi-head attention mechanism cascades with the Graph Transformer model to embed and represent resource and topological features in the local and global scopes, the RL framework based on the policy network PPO is used to learn and generate a multi-objective resource scheduling optimization strategy. Compared with the previous traditional methods and the scheduling optimization strategy based on GNN, the multi-objective reinforcement learning scheduling optimization algorithm assisted by the cascaded multi-layer GAT model and Graph Transformer model provided by the embodiments of the present application can effectively learn the multi-dimensional features of heterogeneous resource nodes and the topological connection relationships between heterogeneous resource nodes, and then better assist the reinforcement learning framework to learn the multi-objective resource optimization scheduling strategy, thereby effectively improving the utilization rate of multi-dimensional heterogeneous resources and meeting the load balancing requirements, enabling the model to adjust the scheduling plan locally and globally to achieve the efficient utilization of heterogeneous resources under balanced load, and achieving the purpose of multi-objective optimization scheduling such as quickly deploying Pods and maximizing the number of deployed Pods.

[0235] As Figure 11 shown, the embodiments of the present application also provide a scheduling device, including a processor 1100 and a transceiver 1110. The transceiver 1110 receives and sends data under the control of the processor 1100, and the processor 1100 is used to perform the following operations:

[0236] Construct a fully-connected weighted undirected graph of the heterogeneous hardware resources according to the heterogeneous hardware resources in the cloud-edge server;

[0237] Input the fully-connected weighted undirected graph into the target network model to obtain the global embedding features output by the target network model; the target network model is formed by cascading a graph attention network GAT model and a graph transformation Graph Transformer model, where the GAT model is a multi-layer GAT model using the multi-head attention mechanism;

[0238] Input the global embedding features and the user scheduling intention into the reinforcement learning model to obtain the scheduling strategy output by the reinforcement learning model.

[0239] In some embodiments of the present application, the processor is further used to perform the following operations:

[0240] Abstract each heterogeneous hardware resource in the cloud-edge server as a node in the fully-connected weighted undirected graph;

[0241] Abstract the connection information between the heterogeneous hardware resources as an edge in the fully-connected weighted undirected graph;

[0242] Construct the fully-connected weighted undirected graph according to the feature representations of each node and the feature representations of each edge.

[0243] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0244] Construct feature vectors of each node and feature vectors of each edge;

[0245] Construct the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge.

[0246] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0247] Dynamically normalize the feature vectors of each node to obtain the node feature matrix of the fully connected weighted undirected graph;

[0248] Dynamically normalize the feature vectors of each edge to obtain the edge feature matrix of the fully connected weighted undirected graph;

[0249] Construct the fully connected weighted undirected graph according to the node feature matrix and the edge feature matrix.

[0250] In some embodiments of the present application, the feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resource corresponding to the node is included in scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resource corresponding to the node;

[0251] And / or

[0252] The feature vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in scheduling, and the fourth element is used to indicate the topological connection type of the edge.

[0253] In some embodiments of the present application, the feature vector of the node further includes: a fifth element, and the fifth element is used to indicate other quantifiable features of the heterogeneous hardware resource corresponding to the node;

[0254] And / or

[0255] The feature vector of the edge further includes: a sixth element, and the sixth element is used to indicate other quantifiable features of the edge.

[0256] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0257] Use a multi-layer GAT model to dynamically calculate the attention weights of each edge of the fully connected weighted undirected graph through the multi-head attention mechanism, and layer by layer update the features of each node of the fully connected weighted undirected graph to obtain the multi-dimensional embedding features of each node;

[0258] Input the multi-dimensional embedding features of each node into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

[0259] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0260] For each node in the fully connected weighted undirected graph, perform a first operation respectively to obtain the embedding features of the node; wherein, the first operation includes:

[0261] Use the first attention weight to perform weighted summation on the features of the neighbor nodes of the node to obtain the first feature of the node; the first attention weight corresponds to the first attention head of the multi-head attention mechanism;

[0262] Concatenate the result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism with the first feature of the corresponding node to obtain the second feature of the node;

[0263] Apply a non-linear activation function after merging and averaging the second feature to obtain the multi-dimensional embedding features of the node.

[0264] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0265] According to the attention scores of the Graph Transformer model, perform global information aggregation on the multi-dimensional embedding features of each node to obtain the target embedding features of each node;

[0266] Aggregate the target embedding features of the node through global average pooling to generate the global embedding features.

[0267] In some embodiments of the present application, the processor is further configured to perform the following operations:

[0268] In the reinforcement learning model, calculate the reward information according to the influence of the action selected by the policy network on the environment update and perform backpropagation to update the weights of the policy network;

[0269] Wherein, the environment includes at least one of the following: environmental state, connection state, task arrival and completion state;

[0270] The reward information includes at least one of the following: resource utilization rate, task completion time, and energy consumption.

[0271] In the embodiments of the present application, on the one hand, a fully-connected weighted undirected graph is constructed for heterogeneous resource nodes, and dynamic standardized multi-performance metrics are introduced for the representation of different types of heterogeneous resource nodes and connections. On the other hand, the multi-head attention mechanism of multiple layers of GAT is used to refine the feature capture and multi-perspective fusion of features in the fully-connected weighted undirected graph within the same layer. Subsequently, the network structure of the Graph Transformer is cascaded, enabling the Graph Transformer to further utilize the global self-attention mechanism to model the graph structure based on the final embedding output of multiple layers of GAT, effectively aggregating key information related to the target node at both the global and local levels, thereby providing accurate and efficient feature embedding representations in complex heterogeneous resource scheduling.

[0272] In the selection of scheduling strategies using artificial intelligence methods, the artificial intelligence method that uses a GNN network to learn and model heterogeneous resource clusters has deficiencies in learning and representing the multi-dimensional features of heterogeneous resource nodes and the topological connections between heterogeneous resource nodes. GAT has more advantages than GNN in resource scheduling, heterogeneous resource management, and complex topological structure modeling. Cascading multiple layers of GAT with multi-head attention and the Graph Transformer can not only more finely embed local and global graph structure information but also more accurately model heterogeneous resource nodes and their topological relationships by dynamically adjusting the weights of nodes and edges. The global feature modeling of the Graph Transformer can make up for the limitations of GAT in the local neighborhood range, thereby generating embedding representations that combine global and local information. In the scenario of combining local and global information, cascading multiple layers of GAT with multi-head attention and the Graph Transformer with a fully-connected weighted undirected graph can effectively aggregate useful information related to the target node at both the global and local levels, thereby providing accurate and efficient feature embedding representations for complex heterogeneous resource scheduling. Then, the global embedding features are input into the RL framework based on the policy network. Based on the resource requirements and priorities input by the user, the learning weights are adaptively adjusted through the reward mechanism, thereby dynamically adjusting the scheduling optimization strategy, enabling the scheduling system to effectively handle the multi-level and multi-distribution dependencies between different nodes, edges, and the entire graph, achieving the multi-objective optimization goals of more accurate scheduling optimization and high resource utilization of CPU, memory, disk, and GPU under balanced load, reducing the Pod deployment time, and increasing the maximum number of Pod deployments. This cloud resource scheduling method can be applied to wireless edge cloud scenarios to improve scheduling performance and accuracy.

[0273] The embodiments of the present application further provide a scheduling device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements each process in the embodiments of the above-mentioned cloud resource scheduling method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0274] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements each process in the above-described embodiments of the cloud resource scheduling method and can achieve the same technical effects. To avoid repetition, details are not described herein again. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0275] The embodiments of the present application also provide a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process in the above-described embodiments of the cloud resource scheduling method and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0276] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.

[0277] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more processes and / or one or more blocks.

[0278] These computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable storage medium generate a paper product including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks.

[0279] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.

[0280] The foregoing is a preferred embodiment of the present application. It should be noted that, for those of ordinary skill in the art, several improvements and modifications can be made without departing from the principle described in the present application, and these improvements and modifications should also be regarded as the protection scope of the present application.

Claims

1. A cloud resource scheduling method, characterized in that: include: According to the heterogeneous hardware resources in the cloud edge server, a fully connected weighted undirected graph of the heterogeneous hardware resources is constructed; Inputting the fully connected weighted undirected graph into a target network model to obtain a global embedding feature output by the target network model; The target network model is formed by a graph attention network GAT model cascaded with a graph transformer model, wherein the GAT model is a multi-layer GAT model using a multi-head attention mechanism; The global embedding features and the user scheduling intention are input into the reinforcement learning model to obtain the scheduling strategy output by the reinforcement learning model.

2. The method according to claim 1, characterized in that According to the heterogeneous hardware resources in the cloud edge server, a fully connected weighted undirected graph of the heterogeneous hardware resources is constructed, including: Abstracting each heterogeneous hardware resource in the cloud edge server as a node in the fully connected weighted undirected graph; Abstracting the connection information between the heterogeneous hardware resources into edges in the fully connected weighted undirected graph; The fully connected weighted undirected graph is constructed according to the feature representation of each node and the feature representation of each edge.

3. The method according to claim 2, characterized in that According to the feature representation of each node and the feature representation of each edge, the fully connected weighted undirected graph is constructed, including: Construct the feature vector of each node and the feature vector of each edge; The fully connected weighted undirected graph is constructed according to the feature vectors of each node and the feature vectors of each edge.

4. The method according to claim 2, characterized in that: Constructing the fully connected weighted undirected graph according to the feature vectors of each node and the feature vectors of each edge, including: Dynamically normalizing the feature vectors of each node to obtain a node feature matrix of the fully connected weighted undirected graph; Dynamically normalizing the feature vectors of each edge to obtain an edge feature matrix of the fully connected weighted undirected graph; The fully connected weighted undirected graph is constructed according to the node feature matrix and the edge feature matrix.

5. The method according to claim 3 or 4, characterized in that: The feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resources corresponding to the node are included in the scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resources corresponding to the node; and / or, The characteristic vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in the scheduling, and the fourth element is used to indicate the topological connection type of the edge.

6. The method according to claim 5, characterized in that The feature vector of the node further includes: a fifth element, the fifth element being used to indicate other quantifiable features of the heterogeneous hardware resources corresponding to the node; and / or, The feature vector of the edge further includes: a sixth element, where the sixth element is used to indicate other quantifiable features of the edge.

7. The method according to claim 2, characterized in that Inputting the fully connected weighted undirected graph into the target network model to obtain the global embedding features output by the target network model includes: Using a multi-layer GAT model, the attention weight of each edge of the fully connected weighted undirected graph is dynamically calculated through a multi-head attention mechanism, and the features of each node of the fully connected weighted undirected graph are updated layer by layer to obtain a multi-dimensional embedding feature of each node; The multi-dimensional embedding features of each node are input into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

8. The method according to claim 7, characterized in that The multi-layer GAT model is used to dynamically calculate the attention weight of each edge of the fully connected weighted undirected graph through a multi-head attention mechanism, and the features of each node of the fully connected weighted undirected graph are updated layer by layer to obtain the multi-dimensional embedding features of each node, including: For each node in the fully connected weighted undirected graph, a first operation is performed respectively to obtain an embedding feature of the node; wherein the first operation includes: Using a first attention weight to perform weighted summation on features of neighboring nodes of a node to obtain a first feature of the node; the first attention weight corresponds to a first attention head of the multi-head attention mechanism; The result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism are concatenated with the first features of the corresponding nodes to obtain the second features of the nodes; A nonlinear activation function is applied to the second features after merging and averaging to obtain a multi-dimensional embedding feature of the node.

9. The method according to claim 7, characterized in that: Input the multi-dimensional embedding features of each node into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model, including: According to the attention score of the Graph Transformer model, global information aggregation is performed on the multi-dimensional embedding features of each node to obtain the target embedding features of each node; The target embedding features of the nodes are aggregated through global average pooling to generate the global embedding features.

10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: In the reinforcement learning model, according to the impact of the action selected by the policy network on the environment update, the reward information is calculated and the weight of the policy network is updated through back propagation; Wherein, the environment includes at least one of the following: environment status, connection status, task arrival and completion status; The reward information includes at least one of the following: resource utilization, task completion time, and energy consumption.

11. A scheduling device, comprising a processor and a transceiver, wherein the transceiver receives and sends data under the control of the processor, characterized in that: The processor is configured to perform the following operations: According to the heterogeneous hardware resources in the cloud edge server, a fully connected weighted undirected graph of the heterogeneous hardware resources is constructed; Inputting the fully connected weighted undirected graph into a target network model to obtain a global embedding feature output by the target network model; The target network model is formed by a graph attention network GAT model cascaded with a graph transformer model, wherein the GAT model is a multi-layer GAT model using a multi-head attention mechanism; The global embedding features and the user scheduling intention are input into the reinforcement learning model to obtain the scheduling strategy output by the reinforcement learning model.

12. The dispatching device according to claim 11, characterized in that: The processor is further configured to perform the following operations: Abstracting each heterogeneous hardware resource in the cloud edge server as a node in the fully connected weighted undirected graph; Abstracting the connection information between the heterogeneous hardware resources into edges in the fully connected weighted undirected graph; The fully connected weighted undirected graph is constructed according to the feature representation of each node and the feature representation of each edge.

13. The dispatching device according to claim 12, characterized in that: The processor is further configured to perform the following operations: Construct the feature vector of each node and the feature vector of each edge; The fully connected weighted undirected graph is constructed according to the feature vectors of each node and the feature vectors of each edge.

14. The dispatching device according to claim 12, characterized in that: The processor is further configured to perform the following operations: Dynamically normalizing the feature vectors of each node to obtain a node feature matrix of the fully connected weighted undirected graph; Dynamically normalizing the feature vectors of each edge to obtain an edge feature matrix of the fully connected weighted undirected graph; The fully connected weighted undirected graph is constructed according to the node feature matrix and the edge feature matrix.

15. The dispatching device according to claim 13 or 14, characterized in that: The feature vector of the node includes: a first element and a second element; the first element is used to indicate whether the heterogeneous hardware resources corresponding to the node are included in the scheduling, and the second element is used to indicate the occupancy rate of the heterogeneous hardware resources corresponding to the node; and / or, The characteristic vector of the edge includes: a third element and a fourth element; the third element is used to indicate whether the edge is included in the scheduling, and the fourth element is used to indicate the topological connection type of the edge.

16. The dispatching device according to claim 15, characterized in that: The feature vector of the node further includes: a fifth element, the fifth element being used to indicate other quantifiable features of the heterogeneous hardware resources corresponding to the node; and / or, The feature vector of the edge further includes: a sixth element, where the sixth element is used to indicate other quantifiable features of the edge.

17. The dispatching device according to claim 12, characterized in that: The processor is further configured to perform the following operations: Using a multi-layer GAT model, the attention weight of each edge of the fully connected weighted undirected graph is dynamically calculated through a multi-head attention mechanism, and the features of each node of the fully connected weighted undirected graph are updated layer by layer to obtain a multi-dimensional embedding feature of each node; The multi-dimensional embedding features of each node are input into the Graph Transformer model to obtain the global embedding features output by the Graph Transformer model.

18. The dispatching device according to claim 17, characterized in that: The processor is further configured to perform the following operations: For each node in the fully connected weighted undirected graph, a first operation is performed respectively to obtain an embedding feature of the node; wherein the first operation includes: Using a first attention weight to perform weighted summation on features of neighboring nodes of a node to obtain a first feature of the node; the first attention weight corresponds to a first attention head of the multi-head attention mechanism; The result features obtained layer by layer by multiple attention heads of the multi-head attention mechanism are concatenated with the first features of the corresponding nodes to obtain the second features of the nodes; A nonlinear activation function is applied to the second features after merging and averaging to obtain a multi-dimensional embedding feature of the node.

19. The dispatching device according to claim 17, characterized in that: The processor is further configured to perform the following operations: According to the attention score of the Graph Transformer model, global information aggregation is performed on the multi-dimensional embedding features of each node to obtain the target embedding features of each node; The target embedding features of the nodes are aggregated through global average pooling to generate the global embedding features.

20. The dispatching device according to any one of claims 11 to 19, characterized in that: The processor is further configured to perform the following operations: In the reinforcement learning model, according to the impact of the action selected by the policy network on the environment update, the reward information is calculated and the weight of the policy network is updated through back propagation; Wherein, the environment includes at least one of the following: environment status, connection status, task arrival and completion status; The reward information includes at least one of the following: resource utilization, task completion time, and energy consumption.

21. A scheduling device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that: When the processor executes the program, the cloud resource scheduling method according to any one of claims 1 to 9 is implemented.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the cloud resource scheduling method as described in any one of claims 1 to 9 are implemented.

23. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the cloud resource scheduling method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Heterogeneous hardware computing power scheduling method and device, equipment and medium

    CN117492986A

  • Cloud edge-end resource scheduling optimization method based on double-layer graph neural network

    CN118134029A

  • Multi-cloud resource scheduling method and device, storage medium and program product

    CN118377617A

  • Cloud edge computing task scheduling method based on reinforcement learning

    CN118740835A

  • Server and a resource scheduling method for use in a server

    US20240192987A1

Cited By

  • Computing power resource optimization method and system based on deep learning

    CN120560859A