Task calculation unloading multi-target dynamic decision-making method based on edge intelligence
By adopting a multi-objective dynamic decision-making method based on edge intelligence on intelligent computing nodes, the problem of poor scheduling efficiency and adaptability in high-concurrency deep learning inference task scheduling is solved, and more efficient resource utilization and task processing is achieved.
Patent Information
- Application Number
- CN202510476934.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When the prior art performs scheduling of high-concurrent deep learning inference tasks on intelligent computing nodes, the scheduling efficiency and adaptability are poor, resulting in long-term waiting for task requests and low resource utilization.
The multi-objective dynamic decision-making method based on edge intelligence is adopted. By establishing a delay and energy consumption model for edge intelligence computing and wireless communication, multi-objective optimization function is designed, and it is defined as a multi-objective constraint Markov process, and the deep deterministic strategy gradient algorithm of dual-value network is used for scheduling.
It improves the efficiency and adaptability of task scheduling, and can effectively balance multiple conflicting or related goals in an edge intelligent environment, improves the overall efficiency of the system, reduces task waiting time, and improves resource utilization.
Smart Images

Figure CN120075846A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of communication technologies, and in particular, to a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence. Background Art
[0002] At present, there are three reasons why the heterogeneous and high-concurrency deep learning inference task offloading from a large number of intelligent hardware is challenging for the computing resource scheduling strategy of intelligent nodes: First, the high-concurrency task requests quickly push the intelligent computing nodes to the limit of resource utilization. On this basis, further improving the throughput rate to meet the service level agreement (SLA) puts higher requirements on the robustness and accuracy of the scheduling algorithm. However, even when using reinforcement learning that can adapt to dynamic environments for resource allocation, the fixed reward pattern may cause the decision-making network to focus too much on a single target and ignore relaxing this target to optimize other targets to improve the overall system efficiency. This will cause high-concurrency task requests to have to wait for idle resources to appear for a long time, further exacerbating the resource tension. Second, it is an effective way to improve resource utilization for intelligent computing nodes to use GPU virtualization slicing technology to dynamically allocate computing resources on demand according to task characteristics instead of monopolizing the entire GPU. The scheduling algorithm not only has to consider the heterogeneity of deep learning inference tasks from multiple dimensions but also find the global optimal solution by optimizing continuous variables while balancing multiple conflicting or related targets, which has been widely recognized as an NP-hard problem. Third, due to the different functional positions and activity patterns of the coverage areas of each intelligent computing node, the deep learning task requests show spatial differences. For example, train stations focus on face detection, while shopping malls have frequent access to entertainment software. Under the same intelligent node, deep learning tasks are still affected by factors such as human activity patterns, work and rest habits, and traffic conditions, showing temporal variability. Developing a scheduling strategy with real-time personalized preferences according to the dynamic spatio-temporal distribution of tasks is an urgent problem to be solved.
[0003] It can be seen that there is an urgent need for a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence with high scheduling efficiency and adaptability. Summary of the Invention
[0004] In view of this, the embodiments of the present invention provide a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, which at least partially solves the problem of poor scheduling efficiency and adaptability in the prior art.
[0005] The embodiments of the present invention provide a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, including:
[0006] Step 1, establish a delay and energy consumption model for edge intelligence computing and wireless communication based on orthogonal frequency division multiplexing;
[0007] Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence according to the latency and energy consumption model, and construct a multi-objective optimization problem based on this;
[0008] Step 3: Define the multi-objective optimization problem as a multi-objective Markov process with constraints;
[0009] Step 4: Design a preference deep deterministic policy gradient algorithm scheduling model based on the utility-based double-value network according to the Markov process;
[0010] Step 5: Establish a dynamic adjustment mechanism for the multi-objective utility function and substitute it into the scheduling model to output policy actions.
[0011] According to a specific implementation manner of the embodiment of the present invention, the specific content of the step 1 includes:
[0012] Step 1.1: The GPU virtualization technology divides the physical GPU resources into multiple virtual GPU instances and independently allocates them to the containers storing different deep learning models. The intelligent node divides the GPU into virtual GPUs of any size according to the GPU resource ratio ρ n granularity from the computing power dimension. When the intelligent node submits a deep learning inference task for the nth IoT terminal device, and the GPU resource ratio is ρ n , the resulting computational latency is
[0013]
[0014] where I n =[I 1 , I 2 ,…, I o represents the one-hot encoding vector used to identify the model containers required for the deep learning inference task, represents the vector of time consumption ratios required for single-byte input data inference on different models, T represents the transpose operation of the matrix, P n represents the data sensed for the nth IoT terminal device, and G represents the unified GPU resource specification of the edge AI box;
[0015] Step 1.2: According to the standard time consumption ratio vector and the computational latency, for the computational offloading execution mode with a GPU resource ratio of ρ n %, calculate its computational energy consumption
[0016]
[0017] where δ represents the effective switching capacitance coefficient closely related to the GPU chip structure, which is used to reflect the energy consumption efficiency of the GPU when calculating deep learning inference tasks;
[0018] Step 1.3, the statistical characteristics of the superposition of multipath signals with different phases and amplitudes at the receiving end are modeled as Rayleigh fading. In the Rayleigh fading channel, according to the free space path loss gain h t and a random variable α that follows a Rayleigh distribution t The product calculation for the task time slot t, located at the center frequency f c The non-line-of-sight channel gain experienced by the OFDM subcarriers
[0019]
[0020] Among them, A d is the antenna gain, d 0 is the unit distance length, d i is the distance between the transmitter and the receiver, d e is the path attenuation exponent, α t The probability density function of σ 2 is the parameter of the Rayleigh distribution related to the scattering characteristics of the environment;
[0021] Step 1.4, set the center frequency f c The channel bandwidth of the OFDM subcarrier is expressed as B c , based on the Shannon formula, the maximum signal transmission rate ψ on the channel is calculated c
[0022]
[0023] Among them, P tr is the signal power output by the IoT terminal device antenna, ω 0 Represents the channel background noise power;
[0024] Step 1.5, let c n For deep learning reasoning tasks The number of occupied subcarrier beams, and the calculation of the communication delay using multiple OFDM subcarriers for parallel transmission
[0025]
[0026] Among them, ψ i represents the maximum signal transmission rate on the i-th subcarrier;
[0027] Step 1.6: Calculate communication energy consumption based on communication delay
[0028]
[0029] According to a specific implementation manner of an embodiment of the present invention, step 2 specifically includes:
[0030] Step 2.1, establish a time optimization objective function η according to the computing delay and the communication delay where represents the deadline for the return of the calculation result of the deep learning task; 1
[0031]
[0032] wherein, represents the deadline for the return of the calculation result of the deep learning task;
[0033] Step 2.2, establish an energy consumption optimization objective function η according to the computing energy consumption and the communication energy consumption where represents the deadline for the return of the calculation result of the deep learning task; 2
[0034]
[0035] Step 2.3, establish a service availability optimization objective function η according to the service availability 1-(N failure / N all ) where N 3
[0036]
[0037] wherein, N all represents the total number of task requests, and N failure represents the number of failed tasks;
[0038] Step 2.4, based on the time optimization objective function η 1 , the energy consumption optimization objective function η 2 and the service availability optimization objective function η 3 , construct the computing offloading scheduling problem of the deep learning inference task in edge intelligence into a multi-objective optimization problem
[0039]
[0040] s.t.C 1 :
[0041] C 2 :
[0042] C 3 :
[0043] C 4 :
[0044] Among them, M represents the total number of subcarrier beams in the orthogonal frequency division multiplexing communication system, and N represents the total number of IoT terminal devices that currently request deep learning inference task scheduling from the intelligent node and have not been allocated resources.
[0045] According to a specific implementation manner of an embodiment of the present invention, step 3 specifically includes:
[0046] Define the multi-objective optimization problem as a multi-objective Markov process with constraints and define it as a five-tuple Among them, the state space includes the edge computing environment parameters and the deep learning inference task requests submitted by IoT terminal devices. The state at time slot t is defined as Among them, and respectively represent the remaining GPU computing resources on the intelligent node at the current time slot and the number of unoccupied subcarrier beams in the orthogonal frequency division multiplexing communication system. According to the system model, the action space The action made by the agent at time slot t in s t is defined as a t =(ρ t , c t ), ρ t and c t both fall within the continuous interval of 0 to 1 through the softmax function of the last layer of the network, and then pass through and The proportional scaling operator is scaled to form the final policy action realizes the transformation of the explicit constraints in the multi-objective optimization problem into the implicit form in the action operation. Among them, represents the floor operation, is the probability transition function, describes the probability that the environment transfers to s t when the state is s t and executes the action a t+1 . The discount factor γ ∈ [0, 1] determines the degree of importance of the agent for the current reward and future rewards. The reward function R is a vector representing multi-objective feedback, different from the scalar value of SORL. The vector reward given by the environment according to the agent's action at time slot t is expressed as Use the linear utility function to map the vector reward to a scalar value to provide a quality evaluation of the policy. The linear utility function μ(r t ) = w · η, where w = [w 1 , w 2 , w 3 T and represent the importance weights and optimization function values of each target, respectively, [·] T represents the transpose of a matrix. The weight vector w is smoothly and dynamically adjusted according to the change trend of the vector reward on each optimization objective function. The agent adopts the scalar expected return optimization criterion to learn with the goal of maximizing the utility of the expected return.
[0047] According to a specific implementation manner of an embodiment of the present invention, step 4 specifically includes:[[]]
[0048] Step 4.1, adopt two value networks Q 1 and Q 2 designed with rewards separated from the optimization objective function, and accordingly predict the expected real-time vector reward r t and the expected maximum vector return a policy network θ aiming to maximize the utility of the cumulative vector return;
[0049] Step 4.2, the policy network θ generates a policy action a t based on the scalar expected return optimization criterion according to the state s t at the current time slot t, and its GPU computing resource scheduling action ρ t and the beam scheduling action c t in the communication process and are scaled to form a policy action for the edge intelligent environment to execute and accordingly calculate the optimal policy network θ *
[0050]
[0051] Step 4.3, use the mean square error loss to approximate the true real-time reward r obtained with delay t and accordingly update the parameters of the value network Q 1 Based on the predicted value of the real-time feedback use the one-step temporal difference loss to update the parameters of the value network Q 2 where the expression of the mean square error loss is
[0052]
[0053] The expression of the one-step temporal difference loss is
[0054]
[0055] where κ is the number of Monte Carlo samples in the experience replay process, Denote the scaled action generated by the target policy network θ' of θ, Q' 2 The parameters of the target value network are updated using soft updates and applied according to the rule Q' after every specific number of steps 2 = τQ 2 +(1 - τ)Q' 2 Synchronize with Q 2 where τ is the soft update coefficient;
[0056] Step 4.4, in any state s i the policy network θ selects an action to maximize the utility value of the expected vector return based on under the current fixed weight vector w The policy network θ uses the gradient descent method to minimize the target loss function and updates the parameters of the policy network θ accordingly
[0057]
[0058] where μ is the utility function, Denote the scaled action made by the policy network θ, and during the gradient descent update process of the policy network θ, the value network Q 2 and the value network Q used internally 1 The parameters of and the weight parameters in the utility function μ are regarded as constants, and the target policy network θ' also uses the soft update rule θ' = τθ+(1 - τ)θ' to synchronize with θ.
[0059] According to a specific implementation manner of the embodiment of the present invention, the step 5 specifically includes:
[0060] Step 5.1, calculate the curvatures of the three optimization objective functions respectively based on the Taylor formula to obtain the optimization function value sequence η j The reward value at any point i within the window i ∈ [2, κ - 2] The second derivative of
[0061] Step 5.2, smooth the second derivative ;
[0062] Step 5.3, calculate the importance weights of each objective in the utility function according to the smoothed second derivative and substitute them into the scheduling model to output the policy action.
[0063] The multi-objective dynamic decision-making scheme for task computing offloading based on edge intelligence in the embodiments of the present invention includes: Step 1, establishing a delay and energy consumption model for edge intelligence computing and wireless communication based on orthogonal frequency division multiplexing; Step 2, designing a multi-objective optimization function for task scheduling and resource allocation in edge intelligence according to the delay and energy consumption model and constructing a multi-objective optimization problem based on this; Step 3, defining the multi-objective optimization problem as a multi-objective Markov process with constraints; Step 4, designing a preference deep deterministic policy gradient algorithm scheduling model for a utility-based dual-value network according to the Markov process; Step 5, establishing a dynamic adjustment mechanism for the multi-objective utility function and substituting it into the scheduling model to output policy actions.
[0064] The beneficial effects of the embodiments of the present invention are as follows: Through the solution of the present invention, explicit constraints in the multi-objective optimization problem are incorporated into the action space design to transform the constrained multi-objective optimization problem into a multi-objective Markov process (MOMDP), and a single-policy multi-objective deep deterministic policy gradient (MO-DDPG) model based on the utility function is established. The constrained policy action design adopted by the MO-DDPG model is significantly different from the way of directly embedding constraint conditions into the reward function in the traditional RL paradigm, which can ensure that the decision-making process strictly follows the computing logic of edge intelligence (Edge AI). Secondly, the discrete point numerical differentiation method based on Taylor's formula is applied to the dynamic trend of the multi-objective optimization function in the utility function to insight into the matching degree between the MORL policy preference and the edge intelligence environment to correct the importance of each objective in the utility function. Compared with the common multi-policy methods in MORL, even without calculating the complex Pareto front, the optimal policy can be made synchronously with the environmental changes. Finally, the value network added to the MO-DDPG algorithm for predicting real-time rewards can provide real-time predicted reward values for the value network used to evaluate long-term rewards to solve the problem of sparse rewards, and a multi-task experience replay mechanism is formulated to collaboratively update the parameters of the dual-value network using real rewards to correct the optimal policy learning bias caused by the prediction value error of real-time rewards, improving the scheduling efficiency and adaptability. Description of the Drawings
[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a schematic flowchart of a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence provided by the embodiments of the present invention;
[0067] Figure 2Schematic flowchart of a deep deterministic policy gradient algorithm based on a utility-based dual-value network provided by an embodiment of the present invention;
[0068] Figure 3 Schematic flowchart of a dynamic adjustment mechanism for a multi-objective utility function provided by an embodiment of the present invention. Detailed implementation manners
[0069] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0070] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0071] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present invention, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0072] It also should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically. The drawings only show the components related to the present invention and are not drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0073] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0074] An embodiment of the present invention provides a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, which can be applied to the computing resource scheduling process in the Internet scenario.
[0075] See Figure 1 , which is a schematic flowchart of a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence provided by an embodiment of the present invention. As Figure 1 shown, the method mainly includes the following steps:
[0076] Step 1, establish a delay and energy consumption model for edge intelligence computing and wireless communication based on orthogonal frequency division multiplexing;
[0077] Specifically, in implementation, consider an edge intelligence Edge AI environment formed by N intelligent IoT devices connected to an edge intelligence node covering this range through a wireless communication link experiencing Rayleigh fading based on orthogonal frequency division multiplexing (OFDM) technology. Each intelligent node is equipped with an edge AI box with a unified GPU resource specification of G for calculating the deep learning inference tasks offloaded by IoT terminal devices within the coverage of its wireless access point (AP). The unit of the resource specification is the number of trillion floating-point operations per second (TFLOPS). The single deep learning inference task offloaded by IoT terminal device n to the edge is defined as a triple where I n =[I 1 , I 2 , …, I o is a one-hot encoded vector used to identify the container of the model required for the deep learning inference task, o represents the total number of containers storing different models, and P n , respectively represent the data volume sensed by the nth IoT terminal device and the deadline for the calculation result to be transmitted back. When the deep learning model is requested for the first time, the intelligent node downloads and caches the container containing all components and configurations of the deep learning model from the cloud, and subsequent requests of the same type only need to quickly obtain and start the corresponding container from the cache according to the model one-hot encoding I n . The specific process of establishing the delay and energy consumption model for edge intelligence computing and wireless communication based on orthogonal frequency division multiplexing is as follows:
[0078] Step 1.1 Establish a time and energy consumption model in the intelligent computing process:
[0079] The GPU virtualization technology divides the physical GPU resources into multiple virtual GPU (vGPU) instances and independently allocates them to the containers storing different deep learning models. The intelligent node divides the GPU into virtual GPUs of any size at a percentage ρ n granularity from the computing power dimension. When the intelligent node submits a deep learning inference task for the nth terminal device The allocation ratio is ρ n When allocating GPU computing resources, the edge computing delay generated can be expressed as:
[0080]
[0081] Among them, represents the vector of time consumption ratios required for inferring single-byte input data on different models. For 5 common deep learning models in IoT: convolutional neural network for state recognition, recurrent neural network for time series prediction, autoencoder for anomaly detection, generative adversarial network for data repair or enhancement, and Transformer for human-computer interaction, the specific instances of Google Net, Res Net18, Bi-directional Long Short-Term Memory Network (Bi-directional LSTM), Gated Recurrent Unit (GRU), Variational Autoencoder (VAE), Conditional Generative Adversarial Network (CGAN), and Bidirectional Encoder Representations from Transformers (BERT) perform inferences on the corresponding classic standard data with equivalent small-batch input data on a single-card RTX 4090 GPU hardware platform. The obtained measurement results are normalized to the computational complexity equivalent to single-byte input, and the resulting vector of time consumption ratios Ω is shown in Table 1.
[0082] Table 1
[0083]
[0084]
[0085] Consistent with the widely adopted power consumption model, for the computing offloading execution mode with a GPU resource occupancy of ρ n %, its execution energy consumption is expressed as:
[0086]
[0087] Among them, δ represents the effective switching capacitance coefficient closely related to the GPU chip structure, which is used to reflect the energy consumption efficiency of the GPU when computing deep learning inference tasks.
[0088] Step 1.2 Establish the time and energy consumption models in the wireless communication process:
[0089] The propagation of wireless signals between IoT terminal devices and APs on edge intelligent nodes will encounter the reflection, refraction, and scattering effects of various obstacles (non-line-of-sight propagation). The statistical characteristics of multipath signals with different phases and amplitudes superimposed on each other at the receiving end are usually modeled as Rayleigh fading. In a Rayleigh fading channel, for task time slot t, the non-line-of-sight channel gain experienced by the orthogonal frequency division multiplexing subcarriers located at the center frequency f c is jointly determined by the free space path (line-of-sight propagation) loss gain h and a random variable α t that follows a Rayleigh distribution t .
[0090]
[0091] Among them, A d is the antenna gain used to quantify the amplification ability of the antenna, d 0 is the unit distance length, d 0 usually takes a value of 1, d i is the distance between the transmitter and the receiver, d e is a path attenuation exponent (urban or rural) that depends on the actual environmental characteristics. In an urban environment, d e generally takes values between 3 and 6. The probability density function of α t is σ 2 is a parameter of the Rayleigh distribution related to the environmental scattering characteristics. We set the channel bandwidth of the orthogonal frequency division multiplexing subcarriers located at the center frequency f c to be represented as B c , and calculate the maximum signal transmission rate ψ c on this channel based on the Shannon formula
[0092]
[0093] Among them, P tr is the signal power output by the IoT terminal device antenna, and ω 0 represents the channel background noise power. During the uplink transmission of the deep learning inference task , the communication delay is mainly affected by the sensed data c n . The deadline d n for the calculation result to be sent back and the one-hot encoding I n used to identify the container type can be ignored in the communication delay analysis due to the small amount of data. In the downlink, the transmission of the calculation result from the intelligent node to the IoT terminal is also regarded as a non-primary factor because the data scale of the deep learning inference result is generally not large enough to be transmitted through the confirmation information of the two-way confirmation mechanism in the communication process. Let c nFor deep learning inference tasks The number of subcarrier beams occupied, the communication delay of parallel transmission using multiple orthogonal frequency division multiplexing subcarriers Is calculated by the following formula.
[0094]
[0095] Intelligent nodes have a stable continuous power supply and consume low power for receiving and decoding orthogonal frequency division multiplexing signals. Therefore, the energy consumption in the wireless communication process is mainly concentrated on the IoT (Internet of Things) devices transmitting signals and is calculated using the following formula.
[0096]
[0097] Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence according to the delay and energy consumption model, and construct a multi-objective optimization problem based on this;
[0098] Specifically, in step 2.1, facing the task calculation delay Communication delay Calculation energy consumption Communication energy consumption Service availability 1 - (N failure / N all ), based on the calculation delay Communication delay Establish a time optimization objective function η 1 , based on the calculation energy consumption Communication energy consumption Establish an energy consumption optimization objective function η 2 , based on the service availability, establish a service availability optimization objective function η 3 .
[0099] Since the continuity of services in cloud services not only affects the rights and interests of customers but is also crucial for service providers, especially in high-concurrency scenarios, this paper minimizes the task calculation delay And execution energy consumption On this basis, the goal of maximizing service availability 1 - (N failure / N all ) is taken into consideration, where the total number of task requests is N all , and the number of failed tasks is N failure. There is a relationship of both mutual connection and conflict among the three optimization objectives: pursuing the minimization of task computing latency often comes with an increase in energy consumption, and new task requests may encounter queuing timeouts due to a shortage of computing resources, which in turn has a negative impact on service availability. To more accurately evaluate the optimization scale of the scheduling strategy on the three objectives, three optimization objective functions are redesigned using non-linear functions for evaluation instead of the traditional linear method. For the time optimization objective, blindly reducing the computing time while meeting the requirement of d n will damage other objectives. At the same time, considering the computing result feedback deadline d n in the task request, and the communication latency will reduce the remaining time available for task scheduling and execution, the time optimization objective function η 1 is designed in the form of a logarithmic function to be maximized to balance the impacts of different magnitudes .
[0100]
[0101] For the optimization objective of execution energy consumption, service providers hope that the energy consumption can be reduced without limit. The energy consumption optimization objective function η 2 is designed using an inverse proportional function and maximized for it. Similarly, in an ideal situation, the service availability should approach 1 as much as possible. We utilize the characteristics of the tangent function to design the maximized service availability optimization objective function η 3 as follows.
[0102]
[0103] Step 2.2 Based on the above three clearly maximized optimization objective functions, the computing offloading scheduling problem of deep learning inference tasks in edge intelligence is constructed as a multi-objective optimization problem P.
[0104]
[0105] s.t. C 1 :
[0106] C 2 :
[0107] C 3 :
[0108] C 4 :
[0109] Among them, the constraint C 1 requires that for any task The allocated amount of computing resources must be strictly limited within the total computing resource specifications provided by the intelligent node. Constraint C 2 Limit all tasks being scheduled The total computing resources obtained by all tasks being scheduled shall not exceed the unified GPU resource specification G of the edge AI box on the intelligent node. Constraint C 3 For the subcarrier beam number c actually occupied by tasks occupied by tasks n shall not exceed the total number M of subcarrier beams in the orthogonal frequency division multiplexing communication system. Constraint C 4 Limit all tasks being scheduled The total subcarrier beams obtained by all tasks being scheduled shall not exceed the total number M of subcarrier beams in the orthogonal frequency division multiplexing communication system. Problem P is an NP-hard mixed integer multi-objective non-linear programming problem with ρ n as the continuous decision variable and c n as the discrete decision variable. The optimization variables are of high dimensionality and are coupled with each other. Traditional meta-heuristic algorithms face the dilemmas of huge computational overhead and long solution time when searching for the optimal solution. We will explore using RL-based methods to solve it in the next step.
[0110] Step 3: Define the multi-objective optimization problem as a multi-objective constrained Markov process;
[0111] In specific implementation, the multi-objective optimization problem can be defined as a multi-objective constrained Markov process MOMDP, and MOMDP is defined as a five-tuple <S, A, P, R, γ>. The state space S consists of the edge computing environment parameters and the deep learning inference task requests submitted by IoT terminal devices. The state at time slot t is defined as where and represent the remaining GPU computing resources on the intelligent node at the current time slot and the number of unoccupied subcarrier beams in the orthogonal frequency division multiplexing communication system respectively. According to the system model, the action taken by the agent at time slot t in the action space A is defined as a t made according to s t =(ρ t , c t ), where both ρ t and c t fall within the continuous interval of 0 to 1 through the softmax function of the last layer of the network, and then form the final policy action after being scaled by and scaling operators. This realizes the transformation of the explicit constraints in the multi-objective optimization problem P into the implicit form in the action operation, where Denotes the floor operation. P: S×A×S→[0,1] is the probability transition function, and P t describes that when the environment is in state s t performs action a t and transfers to s t+1 The probability. The discount factor γ∈[0,1] determines the importance that the agent attaches to current rewards and future rewards. The reward function R is a vector representing multi-objective feedback, different from the scalar value of SORL. The vector reward given by the environment according to the agent's action at time slot t is expressed as It is not feasible to directly compare any two policies with vector rewards because it is impossible to provide a complete ordering in the policy space like SORL. We use a linear utility function to map the vector reward to a scalar value to provide a quality assessment of the policy. The linear utility function μ(r t ) = w·η, where w = [w 1 , w 2 , w 3 T and represent the importance weights of each objective and the optimization function values respectively. [·] T represents the transpose of the matrix. The weight vector w is smoothly and dynamically adjusted according to the change trend of the vector reward on each optimization objective function. The agent adopts the scalar expected return (SER) optimization criterion to learn with the goal of maximizing the utility of the expected return.
[0112] Step 4: Design a preference deep deterministic policy gradient algorithm scheduling model for the utility-based double-value network according to the Markov process;
[0113] Specifically, as Figure 2 shown, the process of designing a preference deep deterministic policy gradient algorithm scheduling model for the utility-based double-value network is as follows:
[0114] Step 4.1 Design the value networks (Critic) Q 1 and Q 2 as well as the action policy network (Actor) θ
[0115] To enable the Agent to generate the best policy in real time when facing an edge intelligent system with dynamic preferences. First, we design an MO-DDPG model with a double-value network to simulate the performance feedback of the edge intelligent system executing the current policy in real time without being affected by its feedback delay. The MO-DDPG model with a double-value network includes two value networks Q 1 and Q 2 Used to predict the expected real-time vector reward r for a given state-action pair t and the expected maximum vector return The policy network θ aims to maximize the utility of the cumulative return. The utility function μ converts the multi-dimensional vector reward into a single scalar value to evaluate the quality of the resource scheduling actions of the policy under different edge intelligence environment preferences. The policy network θ generates policy actions a based on the scalar expected return optimization criterion according to the state s at the current time slot t t and its GPU computing resource scheduling action ρ t and the beam scheduling action c during communication t After being scaled by t and and form the policy actions available for execution in the edge intelligence environment The optimal policy network θ * can be solved by the following formula
[0116]
[0117] Due to the delay in the feedback of the real-time reward r t θ * uses the Q 1 value network to predict the real-time reward based on the state s at the current time slot t t and the scaled true action to replace r during the calculation of the cumulative return. The Q t value network and the Q 1 value network actually both predict the current value and future cumulative return of the feedback of 5 performance metrics for the execution of deep learning inference tasks in the edge intelligence environment, and are not affected by the dynamic preferences of the edge intelligence environment caused by changes in task distribution 2
[0118] Step 4.2 Learning method of the Q 1 value network and Q 2
[0119] The Q 1 value network focuses on directly predicting the delayed real reward without considering the impact of the uncertainty of future states. Therefore, the mean squared error loss (MES) is directly used to approximate the real-time reward r obtained by delay t 2 The update of the Q value network parameters is based on the predicted value of the real-time feedback using one-step temporal difference (TD) learning. The TD error is 2 The target value network Q' 2 Establish an objective function with MES as the loss to more stably minimize the TD error.
[0120]
[0121] where κ is the number of Monte Carlo samples in the experience replay process, represents the scaled action generated by the target policy network θ' of θ. Q' 2 The parameters of the target value network are updated softly and applied according to the rule Q' after every specific step size 2 = τQ 2 +(1 - τ)Q' 2 to synchronize with Q 2 , where τ is the soft update coefficient. Q 1 The value network and Q 2 The learning process of the value network is independent of the edge intelligence environment preference and thus is not affected by the utility function μ.
[0122] Step 4.3 Learning method of the action network θ
[0123] The policy action a generated by the policy network θ t will be evaluated by an evaluation mechanism dominated by the dynamic utility function μ. Using the optimization objective function η and its importance weight w, and the predicted high-dimensional performance feedback is jointly reduced to a scalar value. The value of the weight vector w in the utility function μ directly guides the preference direction of the policy actions of the policy network θ. Specifically, at any state s i the policy network θ selects the action to maximize the utility value of the expected vector return based on under the current fixed weight vector w Therefore, the policy network θ uses the gradient descent method to minimize the following loss function.
[0124]
[0125] where represents the scaled action made by the policy network θ. During the gradient descent update process of the policy network θ, Q 2 the value network and the Q used inside it 1 the parameters of the value network and the weight parameters in the utility function μ are regarded as constants. The θ' target policy network also uses the soft update rule θ' = τθ+(1 - τ)θ' to synchronize with θ.
[0126] Step 5, establish a multi-objective utility function dynamic adjustment mechanism and substitute it into the scheduling model to output the policy action.
[0127] In specific implementation, as Figure 3 shown, affected by the spatio-temporal dynamics of task distribution in the edge intelligence environment, the preference among objectives in the utility function should evolve dynamically with the change of task distribution to continuously ensure the generation of the optimal scheduling strategy of the policy network θ in a long time. The degree of adaptation of the policy network θ to the environmental preference is directly reflected in the change trend of the dependent variable of the non-linear optimization function of each objective in the utility function. For the performance metrics and N failure directly feedback by the edge intelligence environment, the intertwined trade-off relationship between them leads to the change trend of their values not only caused by the change of task distribution in the edge intelligence environment but also closely related to the inherent properties of the edge intelligent computing environment. The set of non-linear optimization objective functions η designed with these performance metrics as independent variables can unify the magnitude of policy quality evaluation and effectively reduce the dimension of weight adjustment through comprehensive evaluation by combining similar performance metrics. And the non-linear function is used to more flexibly control the pursuit degree of the policy network θ for each performance metric.
[0128] Step 5.1 Calculate the curvature of the three optimization objective functions respectively based on the Taylor formula
[0129] To judge whether each optimization objective function with performance metrics as input parameters has reached the convergence state according to its change trend, we first use the Taylor formula for discrete point numerical differentiation to cope with the random fluctuation of the optimization objective function value jointly induced by the dynamics of deep learning inference tasks and the time-variability of computing resources, so as to obtain the overall trend of each optimization objective function in the κ Monte Carlo sample number window and avoid the quadratic error caused by first continuousizing the discrete points and then taking the second derivative. We select the values of 5 points for each optimization objective function for finite difference to approximate the 5th-degree polynomial of the Taylor expansion. For the optimization function η of the jth objective j at the i-th reward value in the window i∈[0,κ] The formula for calculating the second derivative through the central difference formula of points is as follows.
[0130]
[0131] Step 5.2 Obtain the sequence of optimization function values η j at any point i in the window i∈[2,κ-2] for the reward value of the second derivative After that, we use moving average to smooth the discrete second derivative to further reduce the random small fluctuations of the Agent in the performance metrics caused by the change of a single task feature in the optimization objective function value.
[0132]
[0133] Among them, n << T is the smoothing window.
[0134] Step 5.3 calculates the importance weights of each objective in the utility function.
[0135] The absolute value of the second derivative of the optimization function value sequence of a certain objective shows a decreasing trend, indicating that the optimization of the objective function has tended to stability. At the same time, the difficulty of further optimizing it also increases. To more effectively balance the optimization requirements among various objectives, it is advisable to reduce the weight allocation of this objective in the overall optimization process so as to reallocate the released weights to other optimization objectives. The weight w of the optimization function of the jth objective j The adjustment formula is as follows.
[0136]
[0137] where β is a hyperparameter used to control the influence degree of the absolute value of the second derivative on the weight. Considering that the Agent's optimization of time, when meeting the task Γ t For the time constraint After that, it is allowed to adjust its preference to other optimization objective functions to allow it to appropriately sacrifice the time optimization objective function η 1 But its value cannot be negative, so as to exchange for the performance indicators in other optimization objective functions and N failure Significantly improved, because when the Agent is too focused on time optimization, it tends to allocate a large amount of GPU computing resources to each task, which will cause subsequent tasks to queue and timeout due to insufficient resources, thus reducing the service availability N failure . For this reason, we set a sacrifice factor α (0 < α < 1), allowing the weight w 1 of the time optimization objective η 1 to decrease, w 1 ← w 1 ·(1 - α). Finally, normalize the weight parameters to ensure that the sum of the weights of all objectives is 1, and then substitute the importance weights into the scheduling model to output the policy action.
[0138] The multi-objective dynamic decision-making method for task computing offloading based on edge intelligence provided in this embodiment transforms the constrained multi-objective optimization problem into a multi-objective Markov process (MOMDP) by incorporating explicit constraints in the multi-objective optimization problem into the action space design, and establishes a single-policy multi-objective deep deterministic policy gradient (MO-DDPG) model based on the utility function. The constrained policy action design adopted by the MO-DDPG model is significantly different from the way of directly embedding the constraint conditions into the reward function in the traditional RL paradigm, which can ensure that the decision-making process strictly follows the computing logic of edge intelligence (Edge AI). Secondly, the discrete point numerical differentiation method based on Taylor's formula is applied to the dynamic trend of the multi-objective optimization function in the utility function to insight into the matching degree between the MORL policy preference and the edge intelligence environment to correct the importance of each objective in the utility function. Compared with the common multi-policy methods in MORL, the optimal policy can be synchronized with the environmental changes without calculating the complex Pareto front. Finally, the value network added to the MO-DDPG algorithm for predicting real-time rewards can provide the reward value predicted in real time for the value network used to evaluate the long-term return to solve the problem of sparse rewards, and a multi-task experience replay mechanism is formulated to collaboratively update the parameters of the dual value network using real rewards to correct the optimal policy learning bias caused by the prediction value error of real-time rewards, improving the scheduling efficiency and adaptability.
[0139] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof.
[0140] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, characterized in that: include: Step 1: Establish a delay and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing; Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the latency and energy consumption model and construct a multi-objective optimization problem accordingly; Step 3, define the multi-objective optimization problem as a multi-objective constrained Markov process; Step 4, design a preference deep deterministic policy gradient algorithm scheduling model based on the utility-based dual value network according to the Markov process; Step 5: Establish a dynamic adjustment mechanism for the multi-objective utility function and substitute it into the scheduling model to output the strategy action.
2. The method according to claim 1, characterized in that The step 1 specifically includes: Step 1.1: GPU virtualization technology divides physical GPU resources into multiple virtual GPU instances and independently allocates them to containers storing different deep learning models. Smart nodes start from the computing power dimension and calculate the GPU resource ratio ρ. n Granularity segmentation GPU is a virtual GPU of any size. When the smart node submits a deep learning inference task Γ for the nth IoT terminal device n The GPU resource ratio is ρ n When , the computational delay is Among them, I n =[I1,I2,…,I o ] represents a one-hot encoded vector used to identify the model container required for deep learning reasoning tasks, represents the time consumption ratio vector required for inference of single-byte input data on different models, T represents the transpose operation of the matrix, P n It represents the data perceived by the nth IoT terminal device, and G represents the unified GPU resource specification of the edge AI box; Step 1.2: Based on the standard time consumption ratio vector and calculation delay, for GPU resource ratio ρ n The computational offloading execution mode is used to calculate its computational energy consumption. Among them, δ represents the effective switching capacitance coefficient which is closely related to the GPU chip structure and is used to reflect the energy efficiency of the GPU when calculating deep learning inference tasks; Step 1.3, the statistical characteristics of the superposition of multipath signals with different phases and amplitudes at the receiving end are modeled as Rayleigh fading. In the Rayleigh fading channel, according to the free space path loss gain h t and a random variable α that follows a Rayleigh distribution t The product calculation for the task time slot t, located at the center frequency f c The non-line-of-sight channel gain experienced by the OFDM subcarriers Among them, A d is the antenna gain, d0 is the unit distance length, d i is the distance between the transmitter and the receiver, d e is the path attenuation exponent, α t The probability density function of σ 2 is the parameter of the Rayleigh distribution related to the scattering characteristics of the environment; Step 1.4, set the center frequency f c The channel bandwidth of the OFDM subcarrier is expressed as B c , based on the Shannon formula, the maximum signal transmission rate ψ on the channel is calculated c Among them, P tr is the signal power output by the IoT terminal device antenna, and ω0 represents the channel background noise power; Step 1.5, let c n For the deep learning reasoning task Γ n The number of occupied subcarrier beams, and the calculation of the communication delay using multiple OFDM subcarriers for parallel transmission Among them, ψ i represents the maximum signal transmission rate on the i-th subcarrier; Step 1.6: Calculate communication energy consumption based on communication delay 3. The method according to claim 2, characterized in that The step 2 specifically includes: Step 2.1, based on the calculated delay and communication delay Establishing time optimization objective function η1 in, Indicates the deadline for the deep learning task calculation results to be returned; Step 2.2, calculate the energy consumption and communication energy consumption Establish energy consumption optimization objective function η2 Step 2.3, based on service availability 1-(N failure / N all ) Establish the service availability optimization objective function η3 Among them, N all Indicates the total number of task requests, N failure Indicates the number of failed tasks; Step 2.4: Based on the time optimization objective function η1, the energy consumption optimization objective function η2, and the service availability optimization objective function η3, the computational offloading scheduling problem of deep learning inference tasks in edge intelligence is constructed as a multi-objective optimization problem. Among them, M represents the total number of subcarrier beams in the orthogonal frequency division multiplexing communication system, and N represents the total number of IoT terminal devices that currently request deep learning inference task scheduling from the intelligent node and have not been allocated resources.
4. The method according to claim 3, characterized in that: The step 3 specifically includes: The multi-objective optimization problem is defined as a multi-objective constrained Markov process and is defined as a five-tuple Among them, the state space Including edge computing environment parameters and deep learning reasoning task requests submitted by IoT terminal devices, the state at time slot t is defined as in, and They represent the remaining GPU computing resources on the smart node in the current time slot and the number of unoccupied subcarrier beams in the OFDM communication system. According to the system model, the action space At time slot t, the agent follows s t The action performed is defined as a t =(ρ t ,c t ), ρ t and c t All pass through the softmax function of the last layer of the network and fall into the continuous interval from 0 to 1, and then pass through and The scaling operator is scaled to form the final strategy action Implemented multi-objective optimization problem The explicit constraint in is transformed into the implicit form in the action operation, where Indicates a round-down operation. is the probability transfer function, Describes the environment in state s t Execute action a t Transfer to t+1 The probability of the discount factor γ∈[0,1] determines the importance of the agent to the current reward and future rewards. The reward function is a vector representing multi-objective feedback. Different from the scalar value of SORL, the vector reward given by the environment according to the action of the agent at time slot t is expressed as Using a linear utility function Mapping the vector reward to a scalar value provides an assessment of the quality of the policy, the linear utility function μ(r t )=w·η,where w=[w1,w2,w3] T and Respectively represent the importance weight and optimization function value of each objective, [·] T Represents the transpose of the matrix. The weight vector w is smoothly and dynamically adjusted according to the changing trend of the vector reward on each optimization objective function. The agent adopts the scalar expected return optimization criterion to learn to maximize the utility of the expected return.
5. The method according to claim 4, characterized in that The step 4 specifically includes: Step 4.1: Take two value networks Q1 and Q2 designed with separate reward and optimization objective functions, and use them to predict the expected real-time vector reward r for a given state-action pair. t and expected maximum vector return The policy network θ aims to maximize the utility of the cumulative vector reward; Step 4.2, the policy network θ is based on the scalar expected return optimization criterion according to the state s at time slot t t Generate policy action a t , and its GPU computing resource scheduling action ρ t and the beam scheduling action c during the communication process t go through and Scaling to form policy actions that can be executed by the edge intelligence environment And calculate the optimal strategy network θ based on this * Step 4.3: Use the mean squared error loss to approximate the delayed real-time reward r t And update the parameters of the value network Q1 accordingly, based on the predicted value of real-time feedback Use the single-step temporal difference loss to update the parameters of the value network Q2, where the expression of the mean square error loss is The expression of the single-step timing difference loss is: Where κ is the number of Monte Carlo samples in the experience replay process, represents the scaled action generated by the target policy network θ' of θ, and the parameters of the target value network Q'2 are synchronized with Q2 using the soft update rule Q'2 = τQ2 + (1-τ)Q'2 after every specific step, where τ is the soft update coefficient; Step 4.4, in any state s i The next policy network θ selects an action Based on maximizing The utility value of the expected vector return under the current fixed weight vector w The policy network θ uses the gradient descent method to minimize the target loss function and update the parameters of the policy network θ accordingly. Among them, μ is the utility function, It represents the scaled action made by the policy network θ. During the update process of the policy network θ using gradient descent, the parameters of the value network Q2 and the value network Q1 used inside it and the weight parameters in the utility function μ are regarded as constants. The target policy network θ' is also synchronized with θ using the soft update rule θ'=τθ+(1-τ)θ'.
6. The method according to claim 5, characterized in that The step 5 specifically includes: Step 5.1, based on Taylor’s formula, calculate the curvatures of the three optimization objective functions respectively and obtain the optimization function value sequence η j The reward value at any point i in the upper i∈[2,κ-2] window The second derivative of Step 5.2, for the second-order derivative Perform smoothing; Step 5.3, according to the second-order derivative after smoothing Calculate the importance weight of each objective in the utility function and substitute it into the scheduling model to output the strategy action.
Citation Information
Patent Citations
Bidirectional cache placement method based on deep reinforcement learning in cellular-free network
CN116321307A
Multi-objective optimization method, system and equipment for RIS auxiliary communication
CN116436512A
Intelligent computing network scheduling method for computing and communication fusion of large model task
CN117667360A
Cited By
Multi-factor environment simulation system for high-speed train passenger comfort research
CN120656354A
Flexible job shop multi-target scheduling method and system based on preference driving
CN121119638A