Heterogeneous computing power resource allocation method and system for mobile edge computing server cluster
By applying the Confrontation Q network reinforcement learning algorithm and Kubernetes vertical scaling mechanism in the mobile edge computing server cluster, combined with GPU virtualization technology, the joint allocation of heterogeneous computing power resources (CPU and GPU) is achieved, and the problems of GPU resource waste and service level protocol default in the existing technology are solved, and resource utilization and user experience are improved.
Patent Information
- Application Number
- CN202411790492.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-09
AI Technical Summary
When handling heterogeneous computing resources (CPU and GPU), existing mobile edge computing server clusters lack effective automatic adjustment mechanisms, resulting in waste of GPU resources, violation of service level agreements, and reducing user service experience.
A heterogeneous computing power resource allocation method for mobile edge computing server clusters is proposed. The heterogeneous computing power resource allocation in the Pod is dynamically adjusted through the duel Q network reinforcement learning algorithm (Dueling DQN), and combined with Kubernetes' vertical scaling mechanism and GPU virtualization technology, the joint allocation of CPU and GPU resources is realized.
It effectively reduces the economic losses caused by idle computing resources in the cluster, ensures that the user's service quality is within the scope stipulated in the service level agreement, and flexibly adapts to the diversity of task types and the differences in task importance.
Smart Images

Figure CN119960957A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mobile communications, and in particular, to a method and system for allocating heterogeneous computing resources of a mobile edge computing server cluster. Background Art
[0002] Mobile Edge Computing (MEC) is a technical architecture that sinks cloud computing capabilities to the edge of mobile networks. Mobile edge computing provides information technology service environment and cloud computing capabilities at the edge of the wireless access network close to mobile users, sinking some or all functions of traditional centralized cloud computing to the edge of the network, reducing data transmission latency and bandwidth pressure on the core network. A MEC server cluster is a cluster composed of multiple mobile edge computing (MEC) servers. These servers are usually deployed at the edge of the mobile network, close to user devices, to provide low-latency, high-bandwidth computing and storage services. MEC server clusters can work together to process requests from different users and applications, improving the overall performance and reliability of the system. At present, MEC server clusters use container virtualization technology, which has the advantages of efficient deployment and strong scalability, and can simplify the deployment and operation and maintenance process of applications. Among them, the commonly used container virtualization orchestration technology is Kubernetes. In the MEC server cluster built with Kubernetes, each instance that uses resources is called a Pod, and each node that provides resources is called a Node. Common computing resources used in MEC clusters are central processing unit (CPU) and graphics processing unit (GPU). Whenever a Pod is created in a Node, a quota of computing resources will be allocated to it to meet the resource requirements of the tasks running on the Pod. In order to save system resources and reduce the probability of system OOM, Kubernetes uses the vertical Pod Autoscaler (VPA) algorithm to automatically adjust the CPU and memory resource requests and limits of containers according to the actual resource usage of the application. Therefore, the VPA algorithm has also become a key research direction in Kubernetes resource management. At present, in order to meet different application requirements, the traditional single CPU computing resource system has begun to transform into a heterogeneous computing resource system. As a general-purpose processor, the performance of the CPU is inferior to that of the GPU in some scenarios. However, the current VPA algorithm does not involve the automatic adjustment of GPU computing resources, which will lead to the waste of GPU computing resources, violation of the service level agreement (SLA), and reduced user service experience. There are currently few solutions to this problem. Most studies have designed different resource management solutions for CPU and GPU computing resources respectively. However, they have ignored the price and function differences between CPUs and GPUs, as well as the priority allocation of CPU and GPU resources for different types of tasks.
[0003] In the MEC server cluster, the research on vertical auto-scaling involves the resource utilization change in the Pod and the SLA default rate. At present, vertical auto-scaling in the MEC cluster faces the following difficulties: 1) Most of the current technical solutions focus on the automatic scaling of CPU and memory resource requests, but do not involve the automatic scaling of GPU resource requests. In the current AI boom, GPU resources have become a strategic resource. Therefore, it is more necessary to reasonably allocate GPU resources in the cluster to improve its resource utilization. 2) In the current technical field, some technical solutions design different resource management solutions for CPU and GPU resources. However, these design solutions still have shortcomings and do not take into account the significant differences in the requirements of different tasks for these two resources. Traditional tasks often rely more on CPU computing resources to complete, such as daily data processing and office software operation. However, intelligent computing tasks involving AI such as machine learning require the powerful parallel computing capabilities of GPU resources, such as the training and reasoning of deep learning models. In addition, the cost prices of CPU and GPU resources are not the same. Treating the two separately may result in the inability to minimize the economic losses caused by insufficient resource utilization in the cluster. If these two resources cannot be reasonably allocated in a coordinated manner, some tasks may over-occupy one resource while the other resource is idle, which will not only reduce the resource utilization of the entire cluster, but also increase unnecessary cost expenditures. Therefore, it is necessary to jointly allocate heterogeneous computing resources.
[0004] Therefore, how to provide a method for jointly allocating heterogeneous computing resources has become an urgent problem to be solved in this field. Summary of the invention
[0005] The present application proposes a method for allocating heterogeneous computing resources of a mobile edge computing server cluster, comprising the following steps: performing virtualization settings; after completing the virtualization settings, initializing information; verifying whether the conditions for the scaling cooling time are currently met; if the conditions for the scaling cooling time are not met, obtaining scaling unit information; after obtaining the scaling unit information, obtaining monitoring indicators of the Pod; judging whether to trigger scaling according to the obtained monitoring indicators; if scaling is triggered, inputting the resource utilization rate and service quality information of the Pod into an algorithm model to obtain a recommended value of computing resources; scheduling the Pod according to the recommended value of computing resources and executing the scheduling.
[0006] In the heterogeneous computing resource allocation method for the mobile edge computing server cluster as described above, the information initialization includes initializing the scaling cooling time, the scaling cycle T, and the monitoring cycle t.
[0007] In the heterogeneous computing resource allocation method for the mobile edge computing server cluster as described above, obtaining the scaling unit information includes obtaining the CPU and GPU resource utilization of the Pod.
[0008] In the heterogeneous computing resource allocation method of the mobile edge computing server cluster as described above, obtaining the monitoring indicators of the Pod includes: obtaining the computing resource utilization rate of the Pod; and obtaining the service quality information of the Pod.
[0009] The heterogeneous computing resource allocation method of the mobile edge computing server cluster as described above, wherein obtaining the computing resource utilization rate of the Pod includes determining the average resource idle loss C of the Pod in the T period T , expressed as:
[0010]
[0011] Among them, p is the price of each CPU resource, q is the price of each GPU resource, is the average idle CPU computing power resource in period T, is the average idle GPU computing power resources in period T. g represents whether the tasks running on the current Pod need to use GPU computing power resources. 0 means no need, and 1 means yes.
[0012] A heterogeneous computing resource allocation system for a mobile edge computing server cluster specifically includes: a virtualization unit, an initialization unit, a verification unit, a scaling unit information acquisition unit, a monitoring index acquisition unit, a judgment unit, a computing resource recommended value acquisition unit and a scheduling unit; the virtualization unit is used to perform virtualization settings; the initialization unit is used to initialize information; the verification unit is used to verify whether the scaling cooling time condition is currently met; if the scaling cooling time condition is not met, the scaling unit information acquisition unit acquires the scaling unit information; the monitoring index acquisition unit is used to acquire the monitoring index of the Pod; the judgment unit is used to determine whether to trigger scaling according to the acquired monitoring index; if scaling is triggered, the computing resource recommended value acquisition unit inputs the resource utilization rate and service quality information of the Pod into an algorithm model to acquire the computing resource recommended value; the scheduling unit is used to schedule the Pod according to the computing resource recommended value and execute the scheduling.
[0013] In the heterogeneous computing resource allocation system of the mobile edge computing server cluster as described above, the initialization unit performs information initialization including initializing the scaling cooling time, the scaling cycle T, and the monitoring cycle t.
[0014] In the heterogeneous computing resource allocation system for the mobile edge computing server cluster as described above, the scaling unit information acquisition unit acquires the scaling unit information including acquiring the CPU and GPU resource utilization of the Pod.
[0015] In the heterogeneous computing resource allocation system for the mobile edge computing server cluster as described above, the monitoring indicator acquisition unit acquires the monitoring indicator of the Pod, including: acquiring the computing resource utilization rate of the Pod; and acquiring the service quality information of the Pod.
[0016] The heterogeneous computing resource allocation system of the mobile edge computing server cluster as described above, wherein the monitoring indicator acquisition unit obtains the computing resource utilization rate of the Pod including determining the average resource idle loss C of the Pod in the T period T , expressed as:
[0017]
[0018] Among them, p is the price of each CPU resource, q is the price of each GPU resource, is the average idle CPU computing power resource in period T, is the average idle GPU computing power resources in period T, and g represents whether the tasks running on the current Pod need to use GPU computing power resources. 0 means no need, and 1 means yes.
[0019] This application has the following beneficial effects:
[0020] (1) This application is based on deploying the Kubernetes container orchestration tool in the cluster, expanding the native vertical scaling mechanism and scheduling mechanism of Kubernetes, and combining it with GPU virtualization technology. The proposed computing resource dynamic allocation design has higher granularity and stronger adaptability.
[0021] (2) The present application can effectively reduce the economic losses caused by idle computing resources within the cluster, while ensuring that the user's service quality is within the range specified by the service level agreement (SLA), and the present application can flexibly adapt to the diversity of task types and differences in task importance. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is an automatic scaling deployment diagram of a mobile edge computing server cluster provided according to an embodiment of the present application;
[0024] Figure 2It is a schematic diagram of the internal structure of a heterogeneous computing power resource allocation system for a mobile edge computing server cluster provided in an embodiment of the present application;
[0025] Figure 3 It is a flowchart of a method for allocating heterogeneous computing power resources of a mobile edge computing server cluster provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0027] This application proposes a method for allocating heterogeneous computing resources in a mobile edge computing server cluster. This method comprehensively considers the utilization rate of heterogeneous computing resources in the Pod, the SLA of tasks running in the Pod, and the remaining computing resources in the Node, and uses the duel Q network reinforcement learning algorithm to dynamically adjust the amount of heterogeneous computing resources allocated in the Pod, aiming to improve the utilization rate of heterogeneous computing resources in the server cluster, reduce the economic losses caused by idle resources in the cluster, and ensure service quality at the same time.
[0028] In order to solve the problem that the existing vertical scaler does not support scaling GPU resources, resulting in a large amount of economic losses due to idle computing resources, a responsive automatic scaling method based on Dueling Deep Q-Network (Dueling DQN) is designed. According to the utilization rate of heterogeneous computing resources in the Pod and the task type, the optimal allocation of heterogeneous computing resources is obtained to improve the utilization rate of heterogeneous computing resources and ensure the quality of service at the same time. This application intends to extend the native vertical scaler and scheduler of Kubernetes to implement dynamic resource management strategies to meet the research objectives.
[0029] Embodiment 1
[0030] like Figure 1 As shown, it is an automatic scaling deployment diagram of the mobile edge computing server cluster. This embodiment develops a unified cluster status control function in the cluster center node. In the edge node of the cluster, the Pod can choose whether to turn on the automatic allocation mode of computing resources. If the Pod turns on automatic allocation, there will be a heterogeneous computing resource recommendation value calculator to continuously monitor the resource utilization and service quality of the Pod. According to this embodiment, the optimal heterogeneous computing resource allocation value of the Pod can be calculated.
[0031] Specifically, the server cluster includes a Master node and Pods on multiple edge nodes, and computing tasks are deployed on edge nodes in the form of Pods.
[0032] This embodiment is to meet the user service quality under different loads, such as Figure 2 As shown, a heterogeneous computing resource allocation system is set in the cluster cloud node, and the system specifically includes: a virtualization unit 210, an initialization unit 220, a verification unit 230, a scaling unit information acquisition unit 240, a monitoring indicator acquisition unit 250, a judgment unit 260, a computing resource recommended value acquisition unit 270 and a scheduling unit 280.
[0033] The virtualization unit 210 is used to perform virtualization settings, specifically virtualizing CPU and GPU resources, and sharing computing resources in the Pod by virtualizing CPU and GPU resources.
[0034] Specifically, Kubernetes (often referred to as K8s, an open source container orchestration system for automating the deployment, expansion, and management of containerized applications) natively supports the virtualization of CPU resources.
[0035] The smallest unit of resource allocation in k8s is Pod. Pod is defined as an instance that uses resources and can contain one or more containers. The English name of the container is Container.
[0036] For the virtualization of GPU resources, you can use Nvidia's native virtualization solution or the open source GaiaGPU solution, and cooperate with the k8s device-plugin to share a single GPU among Pods.
[0037] The initialization unit 220 is used to initialize information.
[0038] The information initialization includes initializing the expansion and contraction cooling time, the expansion and contraction period T and the monitoring period t.
[0039] A set of Pods in a single node is represented by I = {i1, i2..., i n} indicates that in order to reduce frequent expansion and contraction operations, the unified scaling cycle in the cluster is T, and the unified resource monitoring and Pod status change cycle is t.
[0040] The verification unit 230 is used to verify whether the condition of scaling the cooling time is currently met.
[0041] The basis for checking whether the condition of the current scaling cooling time is met is to subtract the current scaling cooling time from the previous cooling time. If the difference between the two is less than a preset specified threshold, it is considered that the condition of the scaling cooling time is not met, and the scaling unit information acquisition unit 240 is executed.
[0042] If the difference between the two is greater than a preset specified threshold, it is considered that the condition of the expansion cooling time is met, and the process ends.
[0043] The scaling unit information acquiring unit 240 is configured to acquire scaling unit information.
[0044] Obtaining scaling unit information means obtaining the initial information of the Pod that needs to be scaled. The initial information of the Pod includes the CPU and GPU resource utilization of the Pod.
[0045] The monitoring indicator acquisition unit 250 is used to obtain the monitoring indicator of the Pod.
[0046] Obtaining Pod monitoring indicators includes obtaining computing resource utilization and service quality information of each node and Pod in the cluster.
[0047] After obtaining the computing resource utilization and service quality information of each node and Pod, it also includes storing the computing resource utilization and service quality information of each node and Pod in an internal time series cloud database.
[0048] In order to reasonably represent the resource utilization of heterogeneous computing resources, this embodiment defines the idle resource loss C to represent the amount of money wasted due to idle resources. The larger the C, the lower the resource utilization, and vice versa. i (i∈I), its average resource idle loss C in period T T It is expressed as:
[0049]
[0050] Among them, p is the price of each CPU resource, q is the price of each GPU resource, and They are the average idle CPU and GPU computing resources in period T, respectively. g represents whether the tasks running on the current Pod need to use GPU computing resources. 0 means no, and 1 means yes.
[0051] Idle resources It can be obtained from the resource allocation and resource utilization, namely:
[0052]
[0053] in, and The T period is pod i Allocated CPU computing resources and GPU computing resources, and are the average resource utilization of CPU and GPU in period T, expressed as:
[0054]
[0055] in, and is the resource utilization rate at time t, n T The number of monitoring times in each scaling period.
[0056] Since different Pods have different task importances and workloads, they have inconsistent target idle resource losses. Therefore, in this embodiment, target resource losses are defined to represent the target loss resource quantities of different task Pods:
[0057]
[0058] Among them, α is the target resource utilization.
[0059] The service quality information mainly considers the processing delay of user requests and whether it can meet the specified service level agreement (SLA).
[0060] In order to allocate resources reasonably, this embodiment defines tasks into two categories: tasks that only use CPU resources are defined as traditional computing tasks Tk t , define the task that can use both GPU resources and CPU resources as intelligent computing task Tk i .
[0061] Determining the quality of service information therefore includes the following steps:
[0062] Step T1: Determine the violation degree of the task using only CPU resources.
[0063] For tasks that use only CPU resources, the processing time of their requests is T t c It can be calculated as follows:
[0064]
[0065] Among them, I t represents the number of instructions required to process a request, f c represents the processing frequency of the CPU core in the node, and I c Indicates the number of instructions that the CPU core can process in each clock cycle.
[0066] Assume that the number of CPU cores in the node is The number of CPU cores allocated to the Pod is A waiting queue is maintained for each Pod, recording the number of instructions required for the task in the waiting state. When a request arrives at the Pod, the length of the waiting queue is q t Assuming that the instructions required by the task are evenly distributed, the waiting time of the request is T t w It is expressed as:
[0067]
[0068] in
[0069] Ignoring communication delays, the response time of a request is T re as follows:
[0070] T re =T t c +T t w (7)
[0071] The degree of SLA violation involves two aspects: one is the proportion of violation requests, and the other is the proportion of violation time. This embodiment only considers the proportion of violation time. It can be calculated using the following formula:
[0072]
[0073] Where T represents the expansion and contraction period of the elastic expander, Represents task Tk t The maximum request response time defined in the SLA, function f1(x) is defined as follows:
[0074]
[0075] Step T2: Determine the violation degree of tasks that can use both GPU resources and CPU resources.
[0076] Because GPUs perform better when processing large-scale parallel tasks, when there are idle GPU resources in the node, GPU resources should be allocated first to complete reasoning tasks, etc.
[0077] Generally, the execution of reasoning tasks is divided into three steps: data loading, calculation execution, and result feedback. i The reasoning task delay T in i inf Defined as:
[0078] Ti inf =T i load +T i compute +T i feedback (10)
[0079] Among them, T i load With T i feedback They are the time required for data loading and data feedback, which are related to the PCIe bandwidth available to the GPU device and can be expressed as follows:
[0080]
[0081] in, and are the size of the input data and the size of the result data, respectively, i is the batch size, B pcie is the PCIe bandwidth of the GPU device, that is, the speed at which the GPU transmits data to the CPU through the PCIe bus. In this embodiment, the GPU is virtualized to enable a GPU to be shared among multiple Pods. Since its principle is time division multiplexing, the computing time T for executing inference tasks in a single Pod is i compute It can be expressed as follows:
[0082]
[0083] Among them, T s compute The computational latency of running inference tasks on a single GPU, The resource amount of a single GPU.
[0084] The execution time of the inference task is related to the computing power of the GPU and the complexity of the model. Therefore, the computing time T for a single GPU to perform an inference task can be obtained. s compute :
[0085]
[0086] Among them, C m is the complexity of the specified model, f is the peak operating frequency of the GPU, and N cores is the number of CUDA cores of a single GPU, and E is the efficiency factor, which represents the decrease in computing efficiency caused by the decrease in GPU frequency due to resource contention and temperature surge.
[0087] It can be concluded that task Tk i SLA violation rate It is expressed as:
[0088]
[0089] in, For Tk i The maximum request response time defined by the task's SLA.
[0090] Step T3: Determine the total violation rate of all tasks.
[0091] The total violation rate of all tasks in the node A vio (T) is expressed as:
[0092]
[0093] Among them, β is the normalization coefficient, It represents the weight ratio of the violation rate of traditional computing tasks and intelligent computing tasks. For example, β = 0.5 means that the two types of tasks are equally important.
[0094] Therefore, in order to reduce idle resource loss, improve node resource utilization, ensure service quality, and reduce SLA violations within the node, the optimization problem in this embodiment can be defined as follows:
[0095]
[0096]
[0097]
[0098] A vio (T)≤ξ (16)
[0099] Among them, C diff is the difference between idle resource loss and target resource loss during period T, C diff =C target -C T The performance balance coefficients ω1 and ω2 represent the relationship between resource utilization and service quality in the cluster. ω1+ω2=1, and ξ is the maximum SLA default rate. In addition, in order to make the two optimization objectives not affected by the numerical value, the two optimization objectives are normalized respectively.
[0100] The determination unit 260 is configured to determine whether to trigger scaling according to the acquired monitoring indicators.
[0101] If the computing resource utilization rate or the total violation rate of the Pod reaches a preset specified threshold, scaling is considered to be triggered and step S270 is executed, otherwise the process ends.
[0102] The computing power resource recommended value acquisition unit 270 is used to input the resource utilization and service quality information of the Pod into the algorithm model to obtain the computing power resource recommended value.
[0103] The computing resource recommendation value acquisition unit 270 is based on the vertical scaling mechanism and is combined with the Kubernetes vertical autoscaler. It continuously monitors the monitoring indicator acquisition unit 250 and obtains the resource utilization and service quality information of the Pod stored in the time series database and inputs it into the algorithm model.
[0104] The algorithm model is the Dueling DQN model. In this embodiment, the computing resource allocation problem in MEC is expressed as a partially observable Markov decision process (POMDP), which can be expressed as G=<S,A,P,r,Z,O,N,γ> , where S represents the state space of the environment, A includes the actions available to the agent, and the agent selects an action a∈A at each time step or scaling interval T. The state of the environment evolves according to the state transition function P, and the reward function r is used to evaluate the quality of the decisions made. Each agent obtains a specific observation z∈Z through the observation function O(s,a):S×A→Z. The discount factor γ∈[0,1) is used to control the emphasis on future rewards in the decision-making process.
[0105] Each Pod in the MEC edge node is regarded as an intelligent agent, and each intelligent agent is regarded as an agent. Each agent can make independent decisions on computing resource allocation, and the decisions between each agent do not interfere with each other.
[0106] Observation: The observation space o of agent i at time t i.t It is expressed as:
[0107]
[0108] In the partially observable case, and Respectively represent the CPU resource usage and GPU resource usage of the current Pod; and Respectively represent the total CPU and GPU computing resources of the current node.
[0109] The state s at time t t It is expressed as:
[0110]
[0111] Action: Because each action involves the allocation of CPU and GPU computing resources, this embodiment defines a combined action to describe the action when the intelligent computing task is running in the Pod. Action space a i,t It is expressed as follows:
[0112]
[0113] When g = 0, only CPU computing resources are allocated. 0 means no change, which is to prevent performance loss caused by frequent scaling and Pod restarts. cpu and AS gpu Represents the minimum resource unit for CPU and GPU expansion and contraction. ±AS cpu Represents increasing or decreasing AS cpu When g = 1, the action is a combined action, which consists of CPU and GPU resource allocation. In the GPU resource allocation action, ±AS gpu Represents increase or decrease AS gpu MB of GPU memory.
[0114] Reward: The reward should be related to the optimization goal of the system. Since scaling a Pod will cause the Pod to restart, the Pod will not be able to process requests for a period of time. Therefore, when the Pod is scaled, a penalty factor p needs to be added. Therefore, the reward function can be expressed as follows:
[0115]
[0116] Since the model in this embodiment involves actions in two different dimensions, the environmental interaction is more complex. Therefore, this embodiment adopts the Dueling DQN algorithm training model, which is innovative based on the traditional DQN and decomposes the Q value function into a state value function and an advantage function. Through this decomposition, Dueling DQN can better understand the value differences in different states and the advantages of different actions relative to the average value.
[0117] Based on the above, the computing power resource recommended value acquisition unit 270 performs the following sub-steps:
[0118] Step Q1: Train the Dueling DQN model.
[0119] The Dueling DQN model network training includes the following sub-steps:
[0120] Step Q11: Input the obtained Pod computing resource utilization data into the Dueling DQN model network.
[0121] Step Q12: Initialize parameters.
[0122] It initializes the system parameter g, determines the task type, and sets hyperparameters such as learning rate, discount factor, and mini-batch size.
[0123] Initialize the experience replay buffer, which is used to store the experience data of the interaction between the agent and the environment.
[0124] Initialize the main network and the target network, where the main network is used to make decisions and update parameters, and the target network copies parameters from the main network after a certain time interval to stabilize the training process. Both the main network and the target network are composed of a state value function network and an advantage function network. Finally, initialize the state space and action space.
[0125] Step Q13: Initialize the agent state space s t .
[0126] Step Q14: Calculate the state value function and advantage function through the main network, and combine the state value function and advantage function to get the action value function. Finally, select the action according to the action value function to get the current action a. t .
[0127] Specifically, the Pod computing resource utilization And the total computing power resource value in the current Node Input the agent, the agent receives the input data, calculates the state value function and advantage function through the main network, then combines them to get the action value function, and finally selects the action based on the action value function.
[0128] Step Q15: Get the average reward r corresponding to the current action t , and transition to the next state s t+1 .
[0129] The agent performs the selected action, and the environment returns a new state and reward based on the action, and determines whether the terminal state is reached.
[0130] Step Q16: Store the experience data obtained by the interaction between the agent and the environment into the buffer.
[0131] The experience data of the agent's interaction with the environment includes the current state, action, reward, new state, and termination flag.
[0132] Step Q17: Take out a small batch of experience data from the experience replay buffer for training the main network.
[0133] Step Q18: Calculate the target network estimation value, and use the loss function and optimization algorithm to update the main network and the target network.
[0134] For each empirical data sampled, the target value is calculated using the target network, and the loss function is used to measure the difference between the predicted value of the main network and the target value. Finally, the main network and the target network are updated through the back propagation algorithm and the optimization algorithm.
[0135] Step Q19: Determine whether the model has converged.
[0136] If the model converges, the Dueling DQN model is output. Otherwise, the agent state space is initialized again.
[0137] Step Q2: Obtain the recommended computing resource value based on the trained Dueling DQN model.
[0138] The Dueling DQN model derives the recommended scaling value of computing resources based on the input data and verifies its legitimacy.
[0139] And the scheduling unit 280 schedules the Pod according to the recommended value of the computing power resource and executes the scheduling.
[0140] The Pod is scheduled and executed according to the recommended value of computing resources. Specifically, the Pod is rescheduled according to the recommended value of computing resources by finding a node with suitable computing resources, notifying the Kubernetes API Server to execute the scheduling, and allocating the recommended value of computing resources to the Pod.
[0141] Embodiment 2
[0142] like Figure 3 As shown, this embodiment provides a method for allocating heterogeneous computing power resources of a mobile edge computing server cluster, which specifically includes the following steps:
[0143] Step S310: Perform virtualization settings.
[0144] The virtualization setting is to virtualize the CPU and GPU resources. By virtualizing the CPU and GPU resources, the computing power resources can be shared in the Pod.
[0145] Specifically, k8s natively supports virtualization of CPU resources. For virtualization of GPU resources, you can use Nvidia's native virtualization solution or the open source Gaia GPU solution, and use the k8s device-plugin to share a single GPU between Pods.
[0146] Step S320: After completing the virtualization setting, initialize the information.
[0147] The information initialization includes initializing the expansion and contraction cooling time, the expansion and contraction period T and the monitoring period t.
[0148] A set of Pods in a single node is represented by I = {i1, i2..., i n}, in order to reduce frequent expansion and contraction operations, the unified scaling cycle in the cluster is T, and the unified resource monitoring and Pod status change cycle is t.
[0149] Step S330: Check whether the condition of the expansion / contraction cooling time is currently met.
[0150] The basis for checking whether the condition of the scaling cooling time is met is to subtract the current scaling start time from the previous scaling end time. If the difference between the two is less than a preset specified threshold, it is considered that the condition of the scaling cooling time is not met, and step S340 is executed.
[0151] If the difference between the two is greater than a preset specified threshold, it is considered that the condition of the expansion cooling time is met, and the process ends.
[0152] Step S340: Obtaining telescopic unit information.
[0153] Obtaining scaling unit information means obtaining the initial information of the Pod that needs to be scaled. The initial information of the Pod includes the CPU and GPU resource utilization of the Pod.
[0154] Step S350: After obtaining the scaling unit information, obtain the monitoring indicators of the Pod.
[0155] Obtaining Pod monitoring indicators includes obtaining computing resource utilization and service quality information of each node and Pod in the cluster.
[0156] After obtaining the computing resource utilization and service quality information of each node and Pod, it also includes storing the computing resource utilization and service quality information of each node and Pod in an internal time series cloud database.
[0157] In order to reasonably represent the resource utilization of heterogeneous computing resources, this embodiment defines the idle resource loss C to represent the amount of money wasted due to idle resources. The larger the C, the lower the resource utilization, and vice versa. i (i∈I), its average resource idle loss C in period T T It is expressed as:
[0158]
[0159] Among them, p is the price of each CPU resource, q is the price of each GPU resource, and They are the average idle CPU and GPU computing resources in period T, respectively. g represents whether the tasks running on the current Pod need to use GPU computing resources. 0 means no, and 1 means yes.
[0160] Idle resources It can be obtained from the resource allocation and resource utilization, namely:
[0161]
[0162] in, and The T period is pod i Allocated CPU computing resources and GPU computing resources, and are the average resource utilization of CPU and GPU in period T, expressed as:
[0163]
[0164] in, and is the resource utilization rate at time t, n T The number of monitoring times in each scaling period.
[0165] Since different Pods have different task importances and workloads, they have inconsistent target idle resource losses. Therefore, in this embodiment, target resource losses are defined to represent the target loss resource quantities of different task Pods:
[0166]
[0167] Among them, α is the target resource utilization.
[0168] The service quality information mainly considers the processing delay of user requests and whether it can meet the specified service level agreement (SLA).
[0169] In order to allocate resources reasonably, this embodiment defines tasks into two categories: tasks that only use CPU resources are defined as traditional computing tasks Tk t , define the task that can use both GPU resources and CPU resources as intelligent computing task Tk i .
[0170] Determining the quality of service information therefore includes the following steps:
[0171] Step S3501: Determine the violation degree of the task using only CPU resources.
[0172] For tasks that use only CPU resources, the processing time of their requests is Tt c It can be calculated as follows:
[0173]
[0174] Among them, I t represents the number of instructions required to process a request, f c represents the processing frequency of the CPU core in the node, and I c Indicates the number of instructions that the CPU core can process in each clock cycle.
[0175] Assume that the number of CPU cores in the node is The number of CPU cores allocated to the Pod is A waiting queue is maintained for each Pod, recording the number of instructions required for the task in the waiting state. When a request arrives at the Pod, the length of the waiting queue is q t Assuming that the instructions required by the task are evenly distributed, the waiting time of the request is T t w It is expressed as:
[0176]
[0177] in
[0178] Ignoring communication delays, the response time of a request is T re as follows:
[0179] T re =T t c +T t w (7)
[0180] The degree of SLA violation involves two aspects: one is the proportion of violation requests, and the other is the proportion of violation time. This embodiment only considers the proportion of violation time. It can be calculated using the following formula:
[0181]
[0182] Where T represents the expansion and contraction period of the elastic expander, Represents task Tk t The maximum request response time defined in the SLA, function f1(x) is defined as follows:
[0183]
[0184] Step S3502: Determine the violation degree of the task that can use both GPU resources and CPU resources.
[0185] Because GPUs perform better when processing large-scale parallel tasks, when there are idle GPU resources in the node, GPU resources should be allocated first to complete reasoning tasks, etc.
[0186] Generally, the execution of reasoning tasks is divided into three steps: data loading, calculation execution, and result feedback. i The reasoning task delay T in i inf Defined as:
[0187] T i inf =T i load +T i compute +T i feedback (10)
[0188] Among them, T i load With T i feedback They are the time required for data loading and data feedback, which are related to the PCIe bandwidth available to the GPU device and can be expressed as follows:
[0189]
[0190] in, and are the size of the input data and the size of the result data, respectively, i is the batch size, B pcie is the PCIe bandwidth of the GPU device, that is, the speed at which the GPU transmits data to the CPU through the PCIe bus. In this embodiment, the GPU is virtualized to enable a GPU to be shared among multiple Pods. Since its principle is time division multiplexing, the computing time T for executing inference tasks in a single Pod is i compute It can be expressed as follows:
[0191]
[0192] Among them, T s compute The computational latency of running inference tasks on a single GPU, The resource amount of a single GPU.
[0193] The execution time of the inference task is related to the computing power of the GPU and the complexity of the model. Therefore, the computing time T for a single GPU to perform an inference task can be obtained. s compute :
[0194]
[0195] Among them, C m is the complexity of the specified model, f is the peak operating frequency of the GPU, and N cores is the number of CUDA cores of a single GPU, and E is the efficiency factor, which represents the decrease in computing efficiency caused by the decrease in GPU frequency due to resource contention and temperature surge.
[0196] It can be concluded that task Tk i SLA violation rate It is expressed as:
[0197]
[0198] in, For Tk i The maximum request response time defined by the task's SLA.
[0199] Step S3503: Determine the total violation rate of all tasks.
[0200] The total violation rate of all tasks in the node A vio (T) is expressed as:
[0201]
[0202] Among them, β is the normalization coefficient, It represents the weight ratio of the violation rate of traditional computing tasks and intelligent computing tasks. For example, β = 0.5 means that the two types of tasks are equally important.
[0203] Therefore, in order to reduce idle resource loss, improve node resource utilization, ensure service quality, and reduce SLA violations within the node, the optimization problem in this embodiment can be defined as follows:
[0204]
[0205] A vio (T)≤ξ (16)
[0206] Among them, C diff is the difference between idle resource loss and target resource loss during period T, C diff =C target -C T The performance balance coefficients ω1 and ω2 represent the relationship between resource utilization and service quality in the cluster. ω1+ω2=1, and ξ is the maximum SLA default rate. In addition, in order to make the two optimization objectives not affected by the numerical value, the two optimization objectives are normalized respectively.
[0207] Step S360: Determine whether to trigger scaling based on the acquired monitoring indicators.
[0208] If the computing resource utilization rate or the total violation rate of the Pod reaches a preset specified threshold, scaling is considered to be triggered and step S370 is executed, otherwise the process ends.
[0209] Step S370: Input the resource utilization and service quality information of the Pod into the algorithm model to obtain the recommended value of computing power resources.
[0210] The algorithm model is the Dueling DQN model, in which this embodiment expresses the computing resource allocation problem in MEC as a partially observable Markov decision process (POMDP), which can be expressed as G = <S, A, P, r, Z, O, N, γ>, where S represents the state space of the environment, A includes the actions available to the agent, and the agent selects an action a∈A at each time step or telescoping interval T. The state of the environment evolves according to the state transition function P, and the reward function r is used to evaluate the quality of the decision made. Each agent obtains a specific observation value z∈Z through the observation function O(s, a): S×A→Z. The discount factor γ∈[0,1) is used to control the emphasis on future rewards in the decision-making process.
[0211] Each Pod in the MEC edge node is regarded as an intelligent agent, and each intelligent agent is regarded as an agent. Each agent can make independent decisions on computing resource allocation, and the decisions between each agent do not interfere with each other.
[0212] Observation: The observation space o of agent i at time t i.t It is expressed as:
[0213]
[0214] In the partially observable case, and Respectively represent the CPU resource usage and GPU resource usage of the current Pod; and Respectively represent the total CPU and GPU computing resources of the current node.
[0215] The state s at time t t It is expressed as:
[0216]
[0217] Action: Because each action involves the allocation of CPU and GPU computing resources, this embodiment defines a combined action to describe the action when the intelligent computing task is running in the Pod. Action space a i,t It is expressed as follows:
[0218]
[0219] When g = 0, only CPU computing resources are allocated. 0 means no change, which is to prevent performance loss caused by frequent scaling and Pod restarts. cpu and AS gpu Represents the minimum resource unit for CPU and GPU expansion and contraction. ±AS cpu Represents increasing or decreasing AS cpu When g = 1, the action is a combined action, which consists of CPU and GPU resource allocation. In the GPU resource allocation action, ±AS gpu Represents increase or decrease AS gpu MB of GPU memory.
[0220] Reward: The reward should be related to the optimization goal of the system. Since scaling a Pod will cause the Pod to restart, the Pod will not be able to process requests for a period of time. Therefore, when the Pod is scaled, a penalty factor p needs to be added. Therefore, the reward function can be expressed as follows:
[0221]
[0222] Since the model in this embodiment involves actions in two different dimensions, the environmental interaction is more complex. Therefore, this embodiment adopts the Dueling DQN algorithm training model, which is innovative based on the traditional DQN and decomposes the Q value function into a state value function and an advantage function. Through this decomposition, Dueling DQN can better understand the value differences in different states and the advantages of different actions relative to the average value.
[0223] Using the Dueling DQN reinforcement learning algorithm to train the model can effectively train the system model in complex action situations, is suitable for the scenario of compound actions in this embodiment, and can achieve the effect of fast and stable model convergence.
[0224] Based on the above, step S370 includes the following sub-steps:
[0225] Step S3701: Train the Dueling DQN model.
[0226] The Dueling DQN model network training includes the following sub-steps:
[0227] Step S37011: Input the acquired Pod's computing resource utilization data into the Dueling DQN model network.
[0228] Step S37012: Initialize parameters.
[0229] The system parameter g is initialized, the task type is determined, and hyperparameters such as learning rate, discount factor, and mini-batch size are set.
[0230] Initialize the experience replay buffer, which is used to store the experience data of the interaction between the agent and the environment.
[0231] Initialize the main network and the target network, where the main network is used to make decisions and update parameters, and the target network copies parameters from the main network after a certain time interval to stabilize the training process. Both the main network and the target network are composed of a state value function network and an advantage function network. Finally, initialize the state space and action space.
[0232] Step S37013: Initialize the agent state space s t .
[0233] Step S37014: Calculate the state value function and advantage function through the main network, and combine the state value function and advantage function to obtain the action value function. Finally, select the action according to the action value function to obtain the current action a. t .
[0234] Specifically, the Pod computing resource utilization And the total computing power resource value in the current Node Input the agent, the agent receives the input data, calculates the state value function and advantage function through the main network, then combines them to get the action value function, and finally selects the action based on the action value function.
[0235] Step S37015: Get the average reward r corresponding to the current action t , and transition to the next state s t+1 .
[0236] The agent performs the selected action, and the environment returns a new state and reward based on the action, and determines whether the terminal state is reached.
[0237] Step S37016: Store the experience data obtained by the interaction between the agent and the environment into the buffer.
[0238] The experience data of the agent's interaction with the environment includes the current state, action, reward, new state, and termination flag.
[0239] Step S37017: Take out a small batch of experience data from the experience playback buffer for training the main network.
[0240] Step S37018: Calculate the target network estimation value, and use the loss function and optimization algorithm to update the main network and the target network.
[0241] For each empirical data sampled, the target value is calculated using the target network, and the loss function is used to measure the difference between the predicted value of the main network and the target value. Finally, the main network and the target network are updated through the back propagation algorithm and the optimization algorithm.
[0242] Step S37019: Determine whether the model has converged.
[0243] If the model converges, the Dueling DQN model is output. Otherwise, the agent state space is initialized again.
[0244] Step S3702: Obtain the recommended value of computing power resources according to the trained Dueling DQN model.
[0245] The Dueling DQN model derives the recommended scaling value of computing resources based on the input data and verifies its legitimacy.
[0246] Step S380: Schedule the Pod according to the recommended computing resource value and execute the scheduling.
[0247] The Pod is scheduled and executed according to the recommended value of computing resources. Specifically, the Pod is rescheduled according to the recommended value of computing resources by finding a node with suitable computing resources, notifying the Kubernetes API Server to execute the scheduling, and allocating the recommended value of computing resources to the Pod.
[0248] This application has the following beneficial effects:
[0249] (1) This application is based on deploying the Kubernetes container orchestration tool in the cluster, expanding the native vertical scaling mechanism and scheduling mechanism of Kubernetes, and combining it with GPU virtualization technology. The proposed computing resource dynamic allocation design has higher granularity and stronger adaptability.
[0250] (2) The present application can effectively reduce the economic losses caused by idle computing resources within the cluster, while ensuring that the user's service quality is within the range specified by the service level agreement (SLA), and the present application can flexibly adapt to the diversity of task types and differences in task importance.
[0251] Although the present application is described with reference to examples, this is for illustrative purposes only and is not intended to limit the present application, and changes, additions and / or deletions to the embodiments may be made without departing from the scope of the present application.
[0252] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for allocating heterogeneous computing resources of a mobile edge computing server cluster, characterized in that: The following steps are involved: Perform virtualization settings; After completing the virtualization settings, initialize the information; Check whether the current scaling cooldown period conditions are met; If the conditions for the telescopic cooling time are not met, the telescopic unit information is obtained; After obtaining the scaling unit information, obtain the monitoring indicators of the Pod; Determine whether to trigger scaling based on the acquired monitoring indicators; If scaling is triggered, the resource utilization and service quality information of the Pod are input into the algorithm model to obtain the recommended value of computing resources; Schedule Pods based on the recommended computing resource values and execute the schedule.
2. The method for allocating heterogeneous computing resources of a mobile edge computing server cluster according to claim 1, characterized in that: Initializing the information includes initializing the scaling cooling time, the scaling cycle T, and the monitoring cycle t.
3. The method for allocating heterogeneous computing resources of a mobile edge computing server cluster according to claim 1, characterized in that: Obtaining scaling unit information includes obtaining the CPU and GPU resource utilization of the Pod.
4. The method for allocating heterogeneous computing resources of a mobile edge computing server cluster according to claim 1, characterized in that: Obtaining Pod monitoring indicators includes: Get the computing resource utilization of the Pod; Get the quality of service information of the Pod.
5. The method for allocating heterogeneous computing resources of a mobile edge computing server cluster according to claim 4, characterized in that: Obtaining the computing resource utilization of the Pod includes determining the average idle resource loss C of the Pod during the T period. T , expressed as: Among them, p is the price of each CPU resource, q is the price of each GPU resource, is the average idle CPU computing power resource in period T, is the average idle GPU computing power resources in period T. g represents whether the tasks running on the current Pod need to use GPU computing power resources. 0 means no need, and 1 means yes.
6. A heterogeneous computing resource allocation system for a mobile edge computing server cluster, characterized in that: Specifically include: Virtualization unit, initialization unit, verification unit, scaling unit information acquisition unit, monitoring indicator acquisition unit, judgment unit, computing power resource recommended value acquisition unit and scheduling unit; A virtualization unit, used for performing virtualization settings; An initialization unit, used for initializing information; A verification unit, used to verify whether the conditions for the expansion and contraction cooling time are currently met; If the condition of the telescopic cooling time is not met, the telescopic unit information obtaining unit obtains the telescopic unit information; The monitoring indicator acquisition unit is used to obtain the monitoring indicators of the Pod; A judgment unit, used to judge whether to trigger scaling according to the acquired monitoring indicators; If scaling is triggered, the computing resource recommended value acquisition unit inputs the resource utilization and service quality information of the Pod into the algorithm model to obtain the computing resource recommended value; The scheduling unit is used to schedule Pods according to the recommended computing resource values and execute the scheduling.
7. The heterogeneous computing resource allocation system for a mobile edge computing server cluster according to claim 6, characterized in that: The initialization unit performs information initialization including initializing the telescopic cooling time, the telescopic cycle T and the monitoring cycle t.
8. The heterogeneous computing resource allocation system for a mobile edge computing server cluster according to claim 6, characterized in that: Scaling unit information acquisition unit acquisition of scaling unit information includes acquiring the CPU and GPU resource utilization of the Pod.
9. The heterogeneous computing resource allocation system for a mobile edge computing server cluster according to claim 6, characterized in that: The monitoring indicators acquisition unit obtains the monitoring indicators of the Pod, including: Get the computing resource utilization of the Pod; Get the quality of service information of the Pod.
10. The heterogeneous computing resource allocation system for a mobile edge computing server cluster according to claim 9, characterized in that: The monitoring indicator acquisition unit obtains the computing resource utilization of the Pod, including determining the average resource idle loss C of the Pod in the T period T , expressed as: Among them, p is the price of each CPU resource, q is the price of each GPU resource, is the average idle CPU computing power resource in period T, is the average idle GPU computing power resources in period T. g represents whether the tasks running on the current Pod need to use GPU computing power resources. 0 means no need, and 1 means yes.
Citation Information
Cited By
Homogeneous API recommendation method and device based on artificial intelligence, electronic equipment and storage medium
CN120974203A