Multi-objective dynamic decision-making method for task computing offloading based on edge intelligence

By establishing a multi-objective optimization model and dual-value network strategy in edge intelligent computing, the scheduling efficiency and adaptability issues of high-concurrency deep learning tasks are solved, efficient resource allocation and dynamic optimization of tasks are achieved, and the overall performance of the system is improved.

CN120075846BActive Publication Date: 2025-09-23XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510476934.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-09-23
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the existing technology, intelligent computing nodes have poor scheduling efficiency and adaptability when facing high-concurrency deep learning inference tasks, making it difficult to find the global optimal solution in multi-objective optimization. In addition, the spatial and temporal variability of task requests lead to uneven resource utilization.

Method used

A wireless communication model of edge intelligent computing and orthogonal frequency division multiplexing is established, a multi-objective optimization function is designed and defined as a multi-objective constrained Markov process, a preference deep deterministic policy gradient algorithm based on a utility-based dual-value network is used for task scheduling, and resource allocation is optimized through a dynamic adjustment mechanism of a multi-objective utility function.

Benefits of technology

It improves scheduling efficiency and adaptability, can achieve multi-objective optimization in a dynamic environment, ensures that the decision-making process follows the computing logic of edge intelligence, corrects real-time reward prediction errors, and improves resource utilization and service availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075846B_ABST
    Figure CN120075846B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, which belongs to the field of communication technology and specifically includes: step 1, establishing a delay and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing; step 2, designing a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the delay and energy consumption model and constructing a multi-objective optimization problem accordingly; step 3, defining the multi-objective optimization problem as a multi-objective constrained Markov process; step 4, designing a preference deep deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network based on the Markov process; step 5, establishing a dynamic adjustment mechanism for the multi-objective utility function and substituting it into the scheduling model to output the policy action. Through the solution of the present invention, scheduling efficiency and adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of communication technology, and in particular to a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence. Background Art

[0002] Currently, offloading heterogeneous, highly concurrent deep learning inference tasks from massive amounts of intelligent hardware poses a challenge to intelligent node computing resource scheduling strategies for three reasons: First, high concurrency of task requests rapidly pushes intelligent computing nodes to the limits of resource utilization. Further improving throughput to meet service-level agreements (SLAs) places higher demands on the robustness and accuracy of scheduling algorithms. However, even when employing reinforcement learning, which can adapt to dynamic environments, for resource allocation, fixed reward models can cause decision networks to overly focus on a single objective, neglecting to relax that objective and optimize other objectives to improve overall system efficiency. This results in highly concurrent task requests having to wait for long periods for idle resources, further exacerbating resource constraints. Second, intelligent computing nodes can effectively improve resource utilization by utilizing GPU virtualization and partitioning technology to dynamically allocate computing resources based on task characteristics, rather than monopolizing the entire GPU. Scheduling algorithms must not only consider the heterogeneity of deep learning inference tasks from multiple dimensions but also balance multiple conflicting or correlated objectives while optimizing continuous variables to find the global optimal solution. This is widely recognized as an NP-hard problem. Third, deep learning task requests vary spatially across intelligent computing nodes due to their functional positioning and activity patterns in the areas they cover. For example, train stations prioritize facial detection, while shopping malls frequently access entertainment software. Within the same intelligent node, deep learning tasks exhibit temporal variability, influenced by factors such as human activity patterns, rest patterns, and traffic conditions. Developing a scheduling strategy that incorporates real-time personalized preferences based on the dynamic spatiotemporal distribution of tasks is an urgent challenge.

[0003] It can be seen that there is an urgent need for a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence with high scheduling efficiency and adaptability. Summary of the Invention

[0004] In view of this, an embodiment of the present invention provides a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, which at least partially solves the problems of poor scheduling efficiency and adaptability in the existing technology.

[0005] An embodiment of the present invention provides a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, including:

[0006] Step 1: Establish a latency and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing;

[0007] Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the latency and energy consumption model and construct a multi-objective optimization problem based on it.

[0008] Step 3: Define the multi-objective optimization problem as a multi-objective constrained Markov process;

[0009] Step 4: Design a preference-deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network according to the Markov process;

[0010] Step 5: Establish a dynamic adjustment mechanism for the multi-objective utility function and substitute it into the scheduling model to output the strategic action.

[0011] According to a specific implementation of the embodiment of the present invention, step 1 specifically includes:

[0012] Step 1.1: GPU virtualization technology divides physical GPU resources into multiple virtual GPU instances and independently allocates them to containers storing different deep learning models. Smart nodes start from the computing power dimension and calculate the GPU resource ratio ρ. n Granularity splits the GPU into virtual GPUs of any size. When the smart node submits a deep learning inference task for the nth IoT terminal device, The GPU resource ratio is ρ n When , the computational delay is

[0013]

[0014] Among them I n =[I1,I2,…,I o ] represents a one-hot encoded vector used to identify the model container required for deep learning reasoning tasks, represents the time consumption ratio vector required for inference of single-byte input data on different models, T represents the transpose operation of the matrix, P n It represents the data perceived by the nth IoT terminal device, and G represents the unified GPU resource specification of the edge AI box;

[0015] Step 1.2: Based on the standard time consumption ratio vector and calculation delay, for GPU resource ratio ρ n % of the computational offloading execution mode, and calculate its computational energy consumption

[0016]

[0017] Among them, δ represents the effective switching capacitance coefficient that is closely related to the GPU chip structure and is used to reflect the energy efficiency of the GPU when calculating deep learning inference tasks;

[0018] Step 1.3: The statistical characteristics of the superposition of multipath signals with different phases and amplitudes at the receiving end are modeled as Rayleigh fading. In the Rayleigh fading channel, according to the free space path loss gain h t and a random variable α that follows a Rayleigh distribution t The product calculation for the task time slot t, located at the center frequency f c The non-line-of-sight channel gain experienced by the OFDM subcarriers

[0019]

[0020] Among them, A d is the antenna gain, d0 is the unit distance length, d i is the distance between the transmitter and receiver, d e is the path attenuation exponent, α t The probability density function of σ 2 is the parameter of the Rayleigh distribution related to the scattering characteristics of the environment;

[0021] Step 1.4, set the center frequency f c The channel bandwidth of the OFDM subcarrier is expressed as B c , calculate the maximum signal transmission rate ψ on the channel based on Shannon's formula c

[0022]

[0023] Among them, P tr is the signal power output by the IoT terminal device antenna, and ω0 represents the channel background noise power;

[0024] Step 1.5, let c n For deep learning reasoning tasks The number of occupied subcarrier beams, calculate the communication delay of parallel transmission using multiple orthogonal frequency division multiplexing subcarriers

[0025]

[0026] Among them, ψ i represents the maximum signal transmission rate on the i-th subcarrier;

[0027] Step 1.6: Calculate communication energy consumption based on communication delay

[0028]

[0029] According to a specific implementation of the embodiment of the present invention, step 2 specifically includes:

[0030] Step 2.1, based on the calculated delay and communication delay Establish time optimization objective function η1

[0031]

[0032] in, Indicates the deadline for returning the calculation results of the deep learning task;

[0033] Step 2.2, calculate energy consumption based on and communication energy consumption Establish energy consumption optimization objective function η2

[0034]

[0035] Step 2.3, based on service availability 1-(N failure / N all ) Establish service availability optimization objective function η3

[0036]

[0037] Among them, N all Indicates the total number of task requests, N failure Indicates the number of failed tasks;

[0038] Step 2.4: Based on the time optimization objective function η1, the energy consumption optimization objective function η2, and the service availability optimization objective function η3, the computational offloading scheduling problem of deep learning inference tasks in edge intelligence is constructed as a multi-objective optimization problem.

[0039]

[0040] stC1:

[0041] C2:

[0042] C3:

[0043] C4:

[0044] Among them, M represents the total number of subcarrier beams in the orthogonal frequency division multiplexing communication system, and N represents the total number of IoT terminal devices that are currently requesting deep learning inference task scheduling from the intelligent node and have not been allocated resources.

[0045] According to a specific implementation of the embodiment of the present invention, step 3 specifically includes:

[0046] The multi-objective optimization problem is defined as a multi-objective constrained Markov process and is defined as a five-tuple Among them, the state space Including edge computing environment parameters and deep learning inference task requests submitted by IoT terminal devices, the state at time slot t is defined as in, and Represent the remaining GPU computing resources on the smart node in the current time slot and the number of unoccupied subcarrier beams in the orthogonal frequency division multiplexing communication system. According to the system model, the action space At time slot t, the agent t The action performed is defined as a t =(ρ t ,c t ), ρ t and c t All of them fall into the continuous interval of 0 to 1 through the softmax function of the last layer of the network, and then pass through and The scaling operator is used to scale the final strategy action. Implemented multi-objective optimization problem The transformation from explicit constraints in to implicit forms in action operations, where Indicates a round-down operation. is the probability transfer function, Describes the environment in state s t Execute action a t Transfer to s t+1 The probability of the agent's action is , the discount factor γ∈[0,1] determines the importance of the agent to the current reward and future rewards. The reward function R is a vector representing multi-objective feedback. Different from the scalar value of SORL, the vector reward given by the environment according to the agent's action at time slot t is expressed as Using a linear utility function The vector reward is mapped to a scalar value to provide a quality assessment of the policy, and the linear utility function μ(r t ) = w·η, where w = [w1, w2, w3] T and Represent the importance weight and optimization function value of each objective respectively, [·] T Represents the transpose of the matrix, and the weight vector w is smoothly and dynamically adjusted according to the changing trend of the vector reward on each optimization objective function. The intelligent agent adopts the scalar expected return optimization standard to learn to maximize the utility of the expected return.

[0047] According to a specific implementation of the embodiment of the present invention, step 4 specifically includes:

[0048] Step 4.1: Take two value networks Q1 and Q2 designed with separate reward and optimization objective functions, and use them to predict the expected real-time vector reward r for a given state-action pair. t and expected maximum vector return A policy network θ that aims to maximize the utility of the cumulative vector reward;

[0049] Step 4.2, the policy network θ is optimized based on the scalar expected return criterion according to the state s at time slot t t Generate policy action a t , and its GPU computing resource scheduling action ρ t and the beam scheduling action c during the communication process t go through and Scaling to form policy actions that can be executed by the edge intelligence environment And calculate the optimal strategy network θ based on this *

[0050]

[0051] Step 4.3: Use the mean squared error loss to approximate the delayed real-time reward r t And accordingly update the parameters of the value network Q1, based on the predicted value of real-time feedback Use the single-step temporal difference loss to update the parameters of the value network Q2, where the expression of the mean square error loss is

[0052]

[0053] The expression of the single-step timing difference loss is:

[0054]

[0055] Where κ is the number of Monte Carlo samples in the experience replay process, represents the scaled action generated by the target policy network θ' of θ, and the parameters of the target value network Q'2 are synchronized with Q2 using the soft update rule Q'2 = τQ2 + (1-τ)Q'2 after every specific step, where τ is the soft update coefficient;

[0056] Step 4.4, in any state s i The next policy network θ selects an action Based on maximization The utility value of the expected vector return under the current fixed weight vector w The policy network θ uses the gradient descent method to minimize the target loss function and update the parameters of the policy network θ accordingly.

[0057]

[0058] Where μ is the utility function, represents the scaled action made by the policy network θ. During the gradient descent update process of the policy network θ, the parameters of the value network Q2 and the value network Q1 used within it and the weight parameters in the utility function μ are regarded as constants. The target policy network θ' is also synchronized with θ using the soft update rule θ'=τθ+(1-τ)θ'.

[0059] According to a specific implementation of the embodiment of the present invention, step 5 specifically includes:

[0060] Step 5.1: Calculate the curvature of the three optimization objective functions based on Taylor’s formula to obtain the optimization function value sequence η j The reward value at any point i in the upper i∈[2,κ-2] window The second derivative of

[0061] Step 5.2, for the second-order derivative Perform smoothing;

[0062] Step 5.3, according to the second-order derivative after smoothing Calculate the importance weight of each objective in the utility function and substitute it into the scheduling model to output the strategic action.

[0063] The multi-objective dynamic decision-making scheme for task computing offloading based on edge intelligence in an embodiment of the present invention includes: step 1, establishing a delay and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing; step 2, designing a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the delay and energy consumption model and constructing a multi-objective optimization problem based on this; step 3, defining the multi-objective optimization problem as a multi-objective constrained Markov process; step 4, designing a preference deep deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network according to the Markov process; step 5, establishing a dynamic adjustment mechanism for the multi-objective utility function and substituting it into the scheduling model to output the policy action.

[0064] The beneficial effects of the embodiments of the present invention are as follows: through the solution of the present invention, the explicit constraints in the multi-objective optimization problem are integrated into the action space design to transform the constrained multi-objective optimization problem into a multi-objective Markov process (MOMDP), and a single-strategy multi-objective deep deterministic policy gradient (MO-DDPG) model based on the utility function is established. The constrained policy action design adopted by the MO-DDPG model is significantly different from the traditional RL paradigm in which the constraints are directly embedded in the reward function. It can ensure that the decision-making process can strictly follow the computational logic of edge intelligence (Edge AI). Secondly, the discrete point numerical differentiation method based on Taylor's formula is applied to the dynamic trend of the multi-objective optimization function in the utility function to gain insight into the matching degree between the MORL policy preference and the edge intelligence environment to correct the importance of each objective in the utility function. Compared with the common multi-strategy method in MORL, even without calculating the complex Pareto front, the optimal strategy can be made synchronously with environmental changes. Finally, the value network added to the MO-DDPG algorithm for predicting real-time rewards can provide real-time predicted reward values ​​for the value network used to evaluate long-term returns to solve the reward sparsity problem. A multi-task experience replay mechanism is formulated to use real rewards to collaboratively update the parameters of the dual value networks to correct the optimal strategy learning deviation caused by the error in the predicted value of real-time rewards, thereby improving scheduling efficiency and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] Figure 1 A flowchart of a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence provided by an embodiment of the present invention;

[0067] Figure 2 A schematic diagram of a flow chart of a deep deterministic policy gradient algorithm based on a utility-based dual-value network provided by an embodiment of the present invention;

[0068] Figure 3 A flowchart of a dynamic adjustment mechanism for a multi-objective utility function provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0069] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0070] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0071] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present invention, those skilled in the art will appreciate that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.

[0072] It should also be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0073] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.

[0074] An embodiment of the present invention provides a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, which can be applied to the computing resource scheduling process in Internet scenarios.

[0075] See also Figure 1 , is a flow chart of a multi-objective dynamic decision-making method for task computing offloading based on edge intelligence provided by an embodiment of the present invention. Figure 1 As shown, the method mainly includes the following steps:

[0076] Step 1: Establish a latency and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing;

[0077] In the specific implementation, consider the edge intelligence Edge AI environment formed by N smart IoT devices connected to each other through a wireless communication link that experiences Rayleigh fading based on orthogonal frequency division multiplexing (OFDM) technology and an edge smart node covering the range. Each smart node is equipped with an edge AI box with a unified GPU resource specification of G to calculate the deep learning inference tasks offloaded by IoT terminal devices within the coverage of its wireless access point (AP). The unit of resource specification is 1 trillion floating-point operations per second (TFLOPS). A single deep learning inference task offloaded to the edge by the IoT terminal device n is defined as a triplet Among them I n =[I1,I2,…,I o ] is a one-hot encoded vector used to identify the container of the model required for deep learning reasoning tasks, o represents the total number of containers storing different models, P n , Respectively represent the amount of data perceived by the nth IoT terminal device and the deadline for returning the calculation results. When the deep learning model is requested for the first time, the smart node downloads and caches the container containing all components and configurations of the deep learning model from the cloud. Subsequent requests of the same type only need to be encoded according to the model's unique hot encoding I n Quickly retrieve and start the corresponding container from the cache. The specific process of establishing the latency and energy consumption model of edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing is as follows:

[0078] Step 1.1 Establish the time and energy consumption model in the intelligent computing process:

[0079] GPU virtualization technology divides physical GPU resources into multiple virtual GPU (vGPU) instances and independently allocates them to containers storing different deep learning models. Smart nodes are allocated based on the percentage of computing power. n Granularity splits the GPU into virtual GPUs of any size. When the smart node submits a deep learning inference task for the nth terminal device The distribution ratio is ρ n When GPU computing resources are used, the edge computing latency generated can be expressed as:

[0080]

[0081] in, This represents the time consumption ratio vector required for inference on single-byte input data on different models. For five common deep learning models in IoT: convolutional neural networks for state recognition, recurrent neural networks for time series prediction, autoencoders for anomaly detection, generative adversarial networks for data repair or enhancement, and transformers for human-computer interaction, specific instances of Google Net, Res Net18, bidirectional long short-term memory (Bi-directional LSTM), gated recurrent unit (GRU), variational autoencoder (VAE), conditional generative adversarial network (CGAN), and bidirectional Encoder Representations from Transformers (BERT) were used to infer the equivalent mini-batch input data of their respective classic standard data on a single RTX 4090 GPU hardware platform. The measured results were normalized to the computational complexity equivalent to single-byte input, resulting in the time consumption ratio vector Ω, as shown in Table 1.

[0082] Table 1

[0083]

[0084]

[0085] In line with the widely adopted power consumption model, the GPU resource ratio is ρ n % of the computational offload execution mode, its execution energy consumption Expressed as:

[0086]

[0087] Among them, δ represents the effective switching capacitance coefficient that is closely related to the GPU chip structure and is used to reflect the energy efficiency of the GPU when calculating deep learning inference tasks.

[0088] Step 1.2 Establish the time and energy consumption model during wireless communication:

[0089] The propagation of wireless signals between IoT terminal devices and edge intelligent nodes on APs will encounter reflection, refraction, and scattering effects from various obstacles (non-line-of-sight propagation). The statistical characteristics of multipath signals with different phases and amplitudes superimposed on each other at the receiving end are usually modeled as Rayleigh fading. In a Rayleigh fading channel, for a task time slot t, the signal at the center frequency f c The non-line-of-sight channel gain experienced by the OFDM subcarriers is the free space path (line-of-sight propagation) loss gain h t and a random variable α that follows a Rayleigh distribution tThe product of is jointly determined.

[0090]

[0091] Among them, A d Antenna gain is used to quantify the antenna's ability to amplify signals. d0 is the unit distance length, which is usually 1. i is the distance between the transmitter and receiver, d e is a path attenuation index that depends on the actual environment characteristics (urban or rural). In urban environments, d e Generally, the value is between 3 and 6. t The probability density function of σ 2 is the parameter of the Rayleigh distribution related to the scattering characteristics of the environment. We set the center frequency f c The channel bandwidth of the OFDM subcarrier is expressed as B c , calculate the maximum signal transmission rate ψ on the channel based on Shannon's formula c .

[0092]

[0093] Among them, P tr is the signal power output by the IoT terminal device antenna, and ω0 represents the channel background noise power. The communication delay in the uplink transmission process is mainly affected by the perception data c n Impact, the deadline for returning the calculation results d n and a one-hot encoding I for identifying the container type n Due to the small amount of data, it can be ignored in the communication delay analysis. The transmission of the calculation results from the smart node to the IoT terminal in the downlink is also considered a non-major factor. The reason is that the data size of deep learning inference results is generally small, which makes it sufficient to be transmitted as a side effect of the confirmation information of the two-way confirmation mechanism in the communication process. Let c n For deep learning reasoning tasks The number of occupied subcarrier beams, then the communication delay of parallel transmission using multiple orthogonal frequency division multiplexing subcarriers Calculated by the following formula.

[0094]

[0095] Smart nodes have a stable, continuous power supply and consume relatively low power to receive and decode OFDM signals. Therefore, the energy consumption during wireless communication is primarily focused on the signals sent by IoT devices and is calculated using the following formula.

[0096]

[0097] Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the latency and energy consumption model and construct a multi-objective optimization problem based on it.

[0098] When implementing, step 2.1 faces the task calculation delay Communication delay Calculating energy consumption Communication energy consumption Service availability 1-(N failure / N all ), based on the calculation delay Communication delay Establishing time optimization objective function η1, based on calculating energy consumption Communication energy consumption An energy consumption optimization objective function η2 is established, and a service availability optimization objective function η3 is established based on service availability.

[0099] Since business continuity in cloud services not only affects customer interests but is also crucial for service providers, especially in high-concurrency scenarios, this paper minimizes task computing latency. and execution energy consumption The service availability will be maximized based on 1-(N failure / N all ) target into consideration, where the total number of task requests is N all , the number of failed tasks is N failure There are both interrelated and conflicting relationships between the three optimization objectives: the pursuit of minimization of task computing delay is often accompanied by an increase in energy consumption, while new task requests may encounter queue timeouts due to a shortage of computing resources, which in turn has a negative impact on service availability. In order to more accurately evaluate the optimization scale of the scheduling strategy on the three objectives, the three optimization objective functions are redesigned using nonlinear functions for evaluation instead of using traditional linear methods. For the time optimization goal, when d is satisfied, n Blindly reducing the computing time based on the requirements will harm other goals. At the same time, considering the deadline d for returning the computing results in the task request n , and the communication delay It will reduce the actual time margin available for task scheduling and execution. The time optimization objective function η1 is designed in the form of a logarithmic function to maximize the balance of different magnitudes. Influence.

[0100]

[0101] For the optimization goal of energy consumption, the service provider hopes that energy consumption can be reduced without limit. The energy consumption optimization objective function η2 is designed and maximized using an inverse proportional function. Similarly, ideally, the service availability should be as close to 1 as possible. We use the characteristics of the tangent function to design the maximum service availability optimization objective function η3 as follows.

[0102]

[0103] In step 2.2, based on the above three explicit maximization optimization objective functions, the computational offloading scheduling problem of deep learning inference tasks in edge intelligence is constructed as a multi-objective optimization problem P.

[0104]

[0105] stC1:

[0106] C2:

[0107] C3:

[0108] C4:

[0109] Among them, constraint C1 requires that for any task The amount of computing resources allocated must be strictly limited to the total computing resource specifications provided by the smart node. Constraint C2 limits all tasks being scheduled The total amount of computing resources obtained must not exceed the unified GPU resource specification G of the edge AI box on the smart node. Constraint C3 is the actual number of tasks Number of occupied subcarrier beams c n The total number of subcarrier beams in the OFDM communication system should not exceed M. Constraint C4 limits all tasks being scheduled. The sum of the subcarrier beams obtained must not exceed the total number of subcarrier beams M in the OFDM communication system. n is a continuous decision variable, c n This is an NP-hard mixed-integer multi-objective nonlinear programming problem with discrete decision variables. The optimization variables are high-dimensional and mutually coupled. Traditional metaheuristic algorithms face the dilemma of huge computational overhead and lengthy solution time when finding the optimal solution. In the next step, we will explore the use of RL-based methods to solve this problem.

[0110] Step 3: Define the multi-objective optimization problem as a multi-objective constrained Markov process;

[0111] In specific implementation, the multi-objective optimization problem can be defined as a multi-objective constrained Markov process MOMDP. MOMDP is defined as a five-tuple<S,A,P,R,γ> The state space S is composed of edge computing environment parameters and deep learning inference task requests submitted by IoT terminal devices. The state at time slot t is defined as in, and They represent the remaining GPU computing resources on the intelligent node in the current time slot and the number of unoccupied subcarrier beams in the OFDM communication system. According to the system model, the intelligent agent at time slot t in the action space A is based on s t The action performed is defined as a t =(ρ t ,c t ), ρ t and c t All of them fall into the continuous interval of 0 to 1 through the softmax function of the last layer of the network, and then pass through and The scaling operator is used to scale the final strategy action. The transformation from explicit constraints in the multi-objective optimization problem P to implicit forms in the action operation is realized, where Indicates the rounding down operation. P:S×A×S→[0,1] is the probability transfer function, P t Describes the environment in state s t Execute action a t Transfer to s t+1 The discount factor γ∈[0,1] determines the importance the agent places on current rewards and future rewards. The reward function R is a vector representing multi-objective feedback. Different from the scalar value of SORL, the vector reward given by the environment according to the agent's action at time slot t is expressed as Direct comparison of any two policies with vector rewards is not feasible because SORL cannot provide a complete ranking on the policy space. We use a linear utility function The vector reward is mapped to a scalar value to provide a quality assessment of the policy, and the linear utility function μ(r t ) = w·η, where w = [w1, w2, w3] T and Represent the importance weight and optimization function value of each objective respectively, [·] T Represents the transpose of the matrix, and the weight vector w is smoothly and dynamically adjusted according to the changing trend of the vector reward on each optimization objective function. The intelligent agent adopts the scalar expected return (SER) optimization standard to learn to maximize the utility of the expected return.

[0112] Step 4: Design a preference-deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network according to the Markov process;

[0113] When implementing it specifically, Figure 2 As shown in Figure 2, the process of designing a preference-depth deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network is as follows:

[0114] Step 4.1 Design the value network (Critic) Q1 and Q2 and the action policy network (Actor) θ

[0115] To enable the agent to generate the best strategy in real time when facing an edge intelligent system with dynamic preferences, we first design a MO-DDPG model with a dual value network to simulate the performance feedback of the edge intelligent system executing the current strategy in real time without being affected by its feedback delay. The MO-DDPG model with a dual value network includes two value networks Q1 and Q2 designed with separate reward and optimization objective functions, which are used to predict the expected real-time vector reward r for a given state-action pair. t and expected maximum vector return The policy network θ aims to maximize the utility of cumulative returns. The utility function μ converts the multi-dimensional vector reward into a single scalar value to evaluate the quality of the policy's resource scheduling actions under different edge intelligent environment preferences. The policy network θ is based on the scalar expected return optimization criterion according to the state s at the time slot t. t Generate policy action a t , and its GPU computing resource scheduling action ρ t and the beam scheduling action c during the communication process t go through and Scaling to form policy actions that can be executed by the edge intelligence environment Optimal policy network θ * It can be solved by the following formula.

[0116]

[0117] Due to the real-time reward t There is a delay in the feedback, θ * In the process of calculating the cumulative return, the Q1 value network is used according to the state s at the current time slot t. t and scaled real-life action Real-time rewards for predictions To replace r tBoth the Q1 value network and the Q2 value network actually predict the current value and future cumulative return of the five performance indicators of deep learning reasoning tasks in the edge intelligence environment, and are not affected by the dynamic preferences of the edge intelligence environment caused by changes in task distribution.

[0118] Step 4.2 Learning Method of Value Networks Q1 and Q2

[0119] The Q1 value network focuses on directly predicting the delayed real reward without considering the impact of uncertainty in future states. Therefore, the mean squared error loss (MES) is directly used to approximate the delayed real-time reward r t The update of Q2 value network parameters is based on the predicted value of real-time feedback Using single-step temporal difference (TD) learning, the TD error is Then, the target value network Q'2 of Q2 is used to establish the objective function with MES as the loss to minimize the TD error more stably.

[0120]

[0121] Where κ is the number of Monte Carlo samples in the experience replay process, represents the scaled action generated by θ', the target policy network for θ. The parameters of the target value network Q'2 are synchronized with Q2 using soft updates, applying the rule Q'2 = τQ2 + (1-τ)Q'2 at specific time steps, where τ is the soft update coefficient. The learning process of the Q1 and Q2 value networks is independent of the edge intelligence environment's preferences and is therefore unaffected by the utility function μ.

[0122] Step 4.3 Learning method of action network θ

[0123] Policy action a generated by policy network θ t Will accept the evaluation mechanism dominated by the dynamic utility function μ, and use the optimization objective function η and its importance weight w to and The predicted high-dimensional performance feedback is jointly reduced to a scalar value, and the value of the weight vector w in the utility function μ directly guides the preferred direction of the policy action of the policy network θ. Specifically, in any state s i The next policy network θ selects an action Based on maximization The utility value of the expected vector return under the current fixed weight vector w Therefore, the policy network θ minimizes the following loss function using gradient descent.

[0124]

[0125] in, represents the scaled action taken by the policy network θ. During the gradient descent update of the policy network θ, the parameters of the Q2 value network and its internal Q1 value network, as well as the weight parameters in the utility function μ, are treated as constants. The target policy network θ' is also synchronized with θ using the soft update rule θ' = τθ + (1-τ)θ'.

[0126] Step 5: Establish a dynamic adjustment mechanism for the multi-objective utility function and substitute it into the scheduling model to output the strategic action.

[0127] When implementing it specifically, Figure 3 As shown in the figure, affected by the spatiotemporal dynamics of task distribution in the edge intelligent environment, the preferences between the objectives in the utility function should evolve dynamically with the change of task distribution to ensure the optimal scheduling strategy generation of the policy network θ for a long time. The degree of adaptation of the policy network θ to the environmental preference is directly reflected in the dependent variable change trend of the nonlinear optimization function of each objective in the utility function. For the performance indicators directly fed back by the edge intelligent environment, and N failure The intertwined trade-offs between them lead to changing trends in their values, not only driven by shifts in the distribution of tasks in edge intelligent computing environments, but also closely related to the inherent properties of edge intelligent computing environments. Designing a set of nonlinear optimization objective functions η using these performance metrics as independent variables can unify the magnitude of policy quality assessments and effectively reduce the dimensionality of weight adjustment by combining similar performance metrics for comprehensive evaluation. Furthermore, nonlinear functions can more flexibly control the degree to which the policy network θ pursues each performance metric.

[0128] Step 5.1 Calculate the curvature of the three optimization objective functions based on Taylor's formula

[0129] In order to determine whether the convergence state has been reached based on the changing trend of each optimization objective function with performance indicators as input parameters, we first use the Taylor formula to perform discrete point numerical differentiation to deal with the random fluctuations of the optimization objective function values ​​induced by the dynamic nature of deep learning inference tasks and the time-varying nature of computing resources, thereby obtaining the overall trend of each optimization objective function in the κ Monte Carlo sample window and avoiding the quadratic error caused by first making the discrete points continuous and then performing second-order derivatives. We select finite differences for the values ​​of 5 points of each optimization objective function to approximate the fifth-order polynomial of Taylor expansion. The optimization function η of the jth objective j The i-th reward value in the window i∈[0,κ] The formula for calculating the second-order derivative using the point-centered difference formula is shown below.

[0130]

[0131] Step 5.2 obtains the optimized function value sequence η j The reward value at any point i in the upper i∈[2,κ-2] window The second derivative of Finally, we use moving average to smooth the discrete second-order derivatives to further reduce the random small fluctuations in the agent's performance indicators caused by changes in the optimization objective function value due to changes in single task characteristics.

[0132]

[0133] Where n<<T is the smoothing window.

[0134] Step 5.3 calculates the importance weight of each objective in the utility function.

[0135] The decreasing trend of the absolute value of the second-order derivative of the optimization function value sequence of a certain objective indicates that the optimization of the objective function has become stable. At the same time, the difficulty of further optimization has also increased. In order to more effectively balance the optimization requirements among the various objectives, the weight distribution of the objective in the overall optimization process should be reduced so that the released weight can be redistributed to other optimization objectives. The weight w of the optimization function of the jth objective is j The adjustment formula is shown below.

[0136]

[0137] Where β is a hyperparameter used to control the influence of the absolute value of the second-order derivative on the weight. Considering that the agent optimizes the time to meet the task Γ t Time constraints Then it is allowed to adjust its preference to other optimization objective functions to allow it to appropriately sacrifice the time optimization objective function η1, but its value cannot be negative in exchange for performance indicators in other optimization objective functions. and N failure Significant improvement, because when the Agent is too focused on time optimization, it tends to allocate a large amount of GPU computing resources to each task, which will cause subsequent tasks to queue up and time out due to insufficient resources, thereby reducing service availability. failure To this end, we set a sacrifice factor α (0 < α < 1) to allow the weight w1 of the time optimization objective η1 to decrease, w1←w1·(1-α). Finally, the weight parameters are normalized to ensure that the sum of the weights of all objectives is 1. The importance weights are then substituted into the scheduling model to output the policy action.

[0138] The multi-objective dynamic decision-making method for task computation offloading based on edge intelligence provided in this embodiment transforms the constrained multi-objective optimization problem into a multi-objective Markov process (MOMDP) ​​by incorporating explicit constraints in the multi-objective optimization problem into the action space design, and establishes a single-strategy multi-objective deep deterministic policy gradient (MO-DDPG) model based on the utility function. The constrained policy action design adopted by the MO-DDPG model is significantly different from the traditional RL paradigm in which constraints are directly embedded in the reward function, and can ensure that the decision-making process can strictly follow the computational logic of edge intelligence (Edge AI). Secondly, the discrete point numerical differentiation method based on Taylor's formula is applied to the dynamic trend of the multi-objective optimization function in the utility function to gain insight into the matching degree between the MORL policy preference and the edge intelligence environment to correct the importance of each objective in the utility function. Compared with the common multi-strategy method in MORL, even without calculating the complex Pareto front, the optimal strategy can be made synchronously with environmental changes. Finally, the value network added to the MO-DDPG algorithm for predicting real-time rewards can provide real-time predicted reward values ​​for the value network used to evaluate long-term returns to solve the reward sparsity problem. A multi-task experience replay mechanism is formulated to use real rewards to collaboratively update the parameters of the dual value networks to correct the optimal strategy learning deviation caused by the error in the predicted value of real-time rewards, thereby improving scheduling efficiency and adaptability.

[0139] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware or a combination thereof.

[0140] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multi-objective dynamic decision-making method for task computing offloading based on edge intelligence, characterized in that: include: Step 1: Establish a latency and energy consumption model for edge intelligent computing and wireless communication based on orthogonal frequency division multiplexing; The step 1 specifically includes: Step 1.1: GPU virtualization technology divides physical GPU resources into multiple virtual GPU instances and independently allocates them to containers storing different deep learning models. Smart nodes start from the computing power dimension and calculate the GPU resource ratio ρ. n Granularity segmentation GPU is a virtual GPU of any size. When the smart node submits a deep learning inference task Γ for the nth IoT terminal device n The GPU resource ratio is ρ n When , the computational delay is Among them, I n =[I1,I2,…,I o ] represents a one-hot encoded vector used to identify the model container required for deep learning reasoning tasks, represents the time consumption ratio vector required for inference of single-byte input data on different models, T represents the transpose operation of the matrix, P n It represents the data perceived by the nth IoT terminal device, and G represents the unified GPU resource specification of the edge AI box; Step 1.2: Based on the standard time consumption ratio vector and calculation delay, for GPU resource ratio ρ n The computational energy consumption of the computational offloading execution mode is calculated. Among them, δ represents the effective switching capacitance coefficient that is closely related to the GPU chip structure and is used to reflect the energy efficiency of the GPU when calculating deep learning inference tasks; Step 1.3: The statistical characteristics of the superposition of multipath signals with different phases and amplitudes at the receiving end are modeled as Rayleigh fading. In the Rayleigh fading channel, according to the free space path loss gain h t and a random variable α that follows a Rayleigh distribution t The product calculation for the task time slot t, located at the center frequency f c The non-line-of-sight channel gain experienced by the OFDM subcarriers Among them, A d is the antenna gain, d0 is the unit distance length, d i is the distance between the transmitter and receiver, d e is the path attenuation exponent, α t The probability density function of σ 2 is the parameter of the Rayleigh distribution related to the scattering characteristics of the environment; Step 1.4, set the center frequency f c The channel bandwidth of the OFDM subcarrier is expressed as B c , calculate the maximum signal transmission rate ψ on the channel based on Shannon's formula c Among them, P tr is the signal power output by the IoT terminal device antenna, and ω0 represents the channel background noise power; Step 1.5, let c n For the deep learning reasoning task Γ n The number of occupied subcarrier beams, calculate the communication delay of parallel transmission using multiple orthogonal frequency division multiplexing subcarriers Among them, ψ i represents the maximum signal transmission rate on the i-th subcarrier; Step 1.6: Calculate communication energy consumption based on communication delay Step 2: Design a multi-objective optimization function for task scheduling and resource allocation in edge intelligence based on the latency and energy consumption model and construct a multi-objective optimization problem based on it. Step 3: Define the multi-objective optimization problem as a multi-objective constrained Markov process; Step 4: Design a preference-deterministic policy gradient algorithm scheduling model based on a utility-based dual-value network according to the Markov process; Step 5: Establish a dynamic adjustment mechanism for the multi-objective utility function and substitute it into the scheduling model to output the strategic action.

2. The method according to claim 1, characterized in that The step 2 specifically includes: Step 2.1, based on the calculated delay and communication delay Establishment time optimization objective function η1 in, Indicates the deadline for returning the calculation results of the deep learning task; Step 2.2, calculate energy consumption based on and communication energy consumption Establish energy consumption optimization objective function η2 Step 2.3, based on service availability 1-(N failure / N all ) Establish service availability optimization objective function η3 Among them, N all Indicates the total number of task requests, N failure Indicates the number of failed tasks; Step 2.4: Based on the time optimization objective function η1, the energy consumption optimization objective function η2, and the service availability optimization objective function η3, the computational offloading scheduling problem of deep learning inference tasks in edge intelligence is constructed as a multi-objective optimization problem. Among them, M represents the total number of subcarrier beams in the orthogonal frequency division multiplexing communication system, and N represents the total number of IoT terminal devices that are currently requesting deep learning inference task scheduling from the intelligent node and have not been allocated resources.

3. The method according to claim 2, characterized in that The step 3 specifically includes: The multi-objective optimization problem is defined as a multi-objective constrained Markov process and is defined as a five-tuple Among them, the state space Including edge computing environment parameters and deep learning inference task requests submitted by IoT terminal devices, the state at time slot t is defined as in, and Respectively represent the remaining GPU computing resources on the smart node in the current time slot and the number of unoccupied subcarrier beams in the orthogonal frequency division multiplexing communication system. According to the system model, the action space At time slot t, the agent t The action performed is defined as a t =(ρ t ,c t ), ρ t and c t All of them fall into the continuous interval of 0 to 1 through the softmax function of the last layer of the network, and then pass through and The scaling operator is used to scale the final strategy action. Implemented multi-objective optimization problem The transformation from explicit constraints in to implicit forms in action operations, where Indicates a round-down operation. is the probability transfer function, Describes the environment in state s t Execute action a t Transfer to s t+1 The probability of the discount factor γ∈[0,1] determines the importance of the agent to the current reward and future rewards. The reward function Is a vector representing multi-objective feedback. Different from the scalar value of SORL, the vector reward given by the environment according to the action of the agent at time slot t is expressed as Using a linear utility function The vector reward is mapped to a scalar value to provide a quality assessment of the policy, and the linear utility function μ(r t ) = w·η, where w = [w1, w2, w3] T and Represent the importance weight and optimization function value of each objective respectively, [·] T Represents the transpose of the matrix, and the weight vector w is smoothly and dynamically adjusted according to the changing trend of the vector reward on each optimization objective function. The intelligent agent adopts the scalar expected return optimization standard to learn to maximize the utility of the expected return.

4. The method according to claim 3, characterized in that The step 4 specifically includes: Step 4.1: Take two value networks Q1 and Q2 designed with separate reward and optimization objective functions, and use them to predict the expected real-time vector reward r for a given state-action pair. t and expected maximum vector return A policy network θ that aims to maximize the utility of the cumulative vector reward; Step 4.2, the policy network θ is optimized based on the scalar expected return criterion according to the state s at time slot t t Generate policy action a t , and its GPU computing resource scheduling action ρ t and the beam scheduling action c during the communication process t go through and Scaling to form policy actions that can be executed by the edge intelligence environment And calculate the optimal strategy network θ based on this * Step 4.3: Use the mean squared error loss to approximate the delayed real-time reward r t And accordingly update the parameters of the value network Q1, based on the predicted value of real-time feedback Use the single-step temporal difference loss to update the parameters of the value network Q2, where the expression of the mean square error loss is The expression of the single-step timing difference loss is: Where κ is the number of Monte Carlo samples in the experience replay process, represents the scaled action generated by the target policy network θ' of θ, and the parameters of the target value network Q'2 are synchronized with Q2 using the soft update rule Q'2 = τQ2 + (1-τ)Q'2 after every specific step, where τ is the soft update coefficient; Step 4.4, in any state s i The next policy network θ selects an action Based on maximization The utility value of the expected vector return under the current fixed weight vector w The policy network θ uses the gradient descent method to minimize the target loss function and update the parameters of the policy network θ accordingly. Where μ is the utility function, represents the scaled action made by the policy network θ. During the gradient descent update process of the policy network θ, the parameters of the value network Q2 and the value network Q1 used within it and the weight parameters in the utility function μ are regarded as constants. The target policy network θ' is also synchronized with θ using the soft update rule θ'=τθ+(1-τ)θ'.

5. The method according to claim 4, characterized in that The step 5 specifically includes: Step 5.1: Calculate the curvature of the three optimization objective functions based on Taylor’s formula to obtain the optimization function value sequence η j The reward value at any point i in the upper i∈[2,κ-2] window The second derivative of Step 5.2, for the second-order derivative Perform smoothing; Step 5.3, according to the second-order derivative after smoothing Calculate the importance weight of each objective in the utility function and substitute it into the scheduling model to output the strategic action.

Citation Information

Patent Citations

  • Multi-objective optimization method, system and equipment for RIS auxiliary communication

    CN116436512A

  • Intelligent computing network scheduling method for computing and communication fusion of large model task

    CN117667360A