A Joint Regulation Method and System for Dynamic Frequency and Deep Learning Model Offloading Based on Deep Reinforcement Learning

Through the joint adjustment method of dynamic frequency and deep learning model offloading based on deep reinforcement learning, the calculation frequency and feature map offload ratio of edge devices are optimized, and the problems of power consumption and inference delay balance of edge devices are solved, and the significant reduction of energy consumption and delay is achieved.

CN115827239BActive Publication Date: 2025-06-10HARBIN INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211603470.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-06-10
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively balance the power consumption and inference delay of edge devices, especially in dynamic voltage and frequency regulation (DVFS) technology, there is a bottleneck problem of feature map offloading, affecting the end-to-end delay of edge-cloud collaborative inference.

Method used

Using a joint adjustment method for dynamic frequency and deep learning model offloading based on deep reinforcement learning, a edge-cloud collaborative inference framework based on DVFS is built by establishing an edge device energy consumption minimization model, and using DRL to optimize the calculation frequency and feature map offloading ratio of edge devices to reduce the overall energy consumption of edge devices.

Benefits of technology

In the experimental results on different data sets, the average energy consumption of edge devices was reduced by 33%, and the end-to-end delay was reduced by more than 54%, which significantly improved the efficiency of edge-cloud collaborative reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115827239B_ABST
    Figure CN115827239B_ABST
Patent Text Reader

Abstract

The present invention proposes a joint regulation method and system for dynamic frequency and deep learning model offloading based on deep reinforcement learning. First, an edge device energy consumption minimization model is established; then, a DVFS-based edge-cloud collaborative inference framework is constructed to jointly optimize the energy consumption and end-to-end latency of edge devices; next, an enhanced DVFS optimization algorithm DVFO based on DRL is constructed to reduce the overall energy consumption of edge devices by jointly optimizing the computing frequency of edge devices and the offloading ratio of feature maps; finally, an offloading mechanism is established to solve the offloading bottleneck problem of feature maps and avoid the direct impact of the size of feature maps on the end-to-end latency of edge-cloud collaborative inference. The present invention is used for DNN feature map offloading in edge-cloud collaborative inference, and a reinforcement learning algorithm using the "think while moving" concurrent strategy calculates the optimal offloading ratio of feature maps and the computing frequency of edge devices for each task to minimize the energy consumption of edge devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of edge computing, and specifically, relates to a joint regulation method and system for dynamic frequency and deep learning model offloading based on deep reinforcement learning. Background Art

[0002] Edge-cloud collaborative inference and traditional dynamic voltage and frequency scaling (DVFS) technologies can effectively balance the power consumption and inference latency of edge devices. The edge-cloud collaborative paradigm offloads some feature maps of a deep neural network (DNN) from a local device to a remote server.

[0003] Among them, the local device executes part of the DNN, and the remote server processes the remaining DNN. Finally, the edge device uses a small neural network to fuse the two parts of the inference results to obtain the final output. DVFS is a widely used low-power technology that dynamically adjusts the operating voltage and clock frequency according to the real-time power consumption of the device, and reduces the system energy consumption by extending the execution time of tasks. Summary of the Invention

[0004] The present invention proposes a joint regulation method and system for dynamic frequency and deep learning model offloading based on deep reinforcement learning, calculates the optimal offloading ratio of feature maps and the computing frequency of edge devices for each task, so as to minimize the energy consumption of edge devices.

[0005] The present invention is realized through the following technical solutions:

[0006] A joint regulation method for dynamic frequency and deep learning model offloading based on deep reinforcement learning, the method specifically includes the following steps:

[0007] Step 1, establish an edge device energy consumption minimization model;

[0008] Step 2, based on the edge device energy consumption minimization model established in Step 1, construct an edge-cloud collaborative inference framework based on DVFS to jointly optimize the energy consumption and end-to-end latency of edge devices;

[0009] Step 3, according to the edge-cloud collaborative inference framework based on DVFS constructed in Step 2, construct an enhanced DVFS optimization algorithm DVFO based on DRL, and reduce the overall energy consumption of edge devices by jointly optimizing the computing frequency of edge devices and the offloading ratio of feature maps;

[0010] Step 4, establish an offloading mechanism to solve the offloading bottleneck problem of feature maps and avoid the size of feature maps directly affecting the end-to-end latency of edge-cloud collaborative inference.

[0011] Further, in Step 1,

[0012] Let χ be a set of tasks consisting of N independent and non-preemptive tasks, denoted as χ = {x 1 , x 2 , …, x N};

[0013] The overall energy consumption of the edge device for task x i is

[0014] where is the computing energy consumption, is the offloading energy consumption;

[0015] The computing energy consumption i of task x depends on the local computing time c i and the computing power and is expressed as

[0016] where is proportional to the square of the voltage V 2 and the computing frequency f i , that is,

[0017] The frequency of the CPU is The frequency of the GPU is The frequency of the memory is Then the computing frequency

[0018] The offloading power consumption i of task x is related to the network bandwidth B, the size m i of the offloaded feature map, and the offloading power and is denoted as Similarly,

[0019] Then minimize the overall energy consumption of the edge device while satisfying the minimum frequency of system operation:

[0020]

[0021] s.t.l i ≤d i

[0022] where l i = c i + o i , representing the end-to-end delay of task x i .

[0023] Furthermore, in step 2,

[0024] Design an edge-cloud collaborative inference framework for DVFS to jointly optimize the energy consumption and end-to-end latency of edge devices

[0025] The DVFS-based edge-cloud collaborative inference framework includes local inference of edge devices and remote inference of cloud servers;

[0026] Use WiFi on the local area network (LAN) as the communication medium to offload the feature maps of the DNN from the edge device to the cloud server.

[0027] Furthermore, in step 2, it specifically includes the following steps

[0028] Step 2.1, use a negligible-overhead feature evaluator on the edge device to judge the current DNN features; that is, use the computational density to evaluate whether the current DNN belongs to a compute-intensive or memory-intensive workload;

[0029] Step 2.2, the DVFS module based on DRL, namely DVFO, learns the DNN characteristics and network bandwidth, and sets the optimal computing frequency of the edge device and the offloading ratio of the feature maps for each task;

[0030] Step 2.3, after obtaining the offloading ratio of the feature maps, it is necessary to partition the DNN and deploy it to the edge side and the cloud respectively;

[0031] Combine the remote DNN to predict the remaining feature maps to obtain the final inference output.

[0032] Furthermore, in step 3, it specifically includes the following steps:

[0033] Step 3.1, transform the optimization objective in formula (1) into the reward function in DRL and model it as a Markov decision process (MDP);

[0034] The agent in the DRL consists of three components: state, action, and reward,

[0035] The state space is specifically: at each moment t, the agent in the DRL constructs a state space Define the type φ of the workload x i and the currently measured network bandwidth B as the state; i The type φ of the workload x

[0036] The type φ of the workload x i and the currently measured network bandwidth B these three discrete metrics constitute the state space i Denoted as Denoted as

[0037] The action space is specifically: set the computing frequency f of the task x i of the task xi Taking the computing frequency \(f\) of the edge device and the offloading ratio \(\xi\) of the feature map as actions, the action space is represented as where

[0038] The reward transforms the energy consumption optimization objective into a reward function, and defines the reward function \(r\) of the agent as follows:

[0039]

[0040] Step 3.2: Use the DQN reinforcement learning algorithm to control the computing frequency and offloading ratio of the edge device to reduce the decision-making overhead of DRL;

[0041] The agent in the DRL observes the state \(s\) of the environment at time \(t\) i , and when it selects an action at the same time, the previous action has slid to a new unobserved state That is, the state capture and policy decision in the concurrent environment can be executed concurrently; where \(H\) is the duration of the action trajectory from state \(s\) t to ;

[0042] Step 3.3: Modify the standard DQN to implement policy decision-making in a concurrent environment, and the DQN in the discrete-time concurrent environment is still convergent;

[0043] Therefore, the concurrent Q-value function of DQN can be reformulated as follows:

[0044]

[0045] Furthermore, in Step 3.3, it specifically includes:

[0046] Step 3.3.1: Initialize the neural network parameters and experience buffer in the DRL; then use the DNN feature map type and the current network bandwidth obtained by the feature evaluator as the initial state of the DRL;

[0047] Step 3.3.2: At the beginning of training, the agent in the DRL randomly selects an action;

[0048] At time \(t\), the agent captures the state \(s\) in the discrete-time concurrent environment t , and at the same time the agent selects an action \(a\) according to the concurrent mechanism t ;

[0049] Step 3.3.3: Use the \(\epsilon\)-greedy strategy to explore the environment, feedback the frequencies and offloading rates selected by the actions to the frequency and offloading controllers respectively, and obtain an immediate reward \(r\);

[0050] Step 3.3.4, set the state from s t to s t+1 , and store the current state, action, reward, and the state at the next moment as an action trajectory in the experience buffer;

[0051] Step 3.3.5, at each gradient step, first randomly sample a small batch of historical trajectories from the experience buffer, then calculate the Q value in the concurrent environment using formula (3), update the network parameters according to gradient descent, and finally deploy the offline-trained DVFO online to evaluate the performance.

[0052] Furthermore, in Step 4, it specifically includes the following steps:

[0053] Step 4.1, determine the number of DNN partition points; for a DNN with N layers, there are a total of 2 N-1 partitioning strategies, and only select one partition point;

[0054] Step 4.2, determine which layer in the DNN should be partitioned, and perform partitioning according to the DNN type given by the feature evaluator;

[0055] For a compute-intensive DNN, select the layer with the highest computational workload for partitioning;

[0056] Similarly, for a memory-intensive DNN, select the layer with the highest memory access volume for partitioning;

[0057] Step 4.3, adopt the precision quantization technology, that is, convert the 32-bit floating-point number of the feature map into an 8-bit fixed-length number to reduce the latency.

[0058] A joint regulation system for dynamic frequency and deep learning model offloading based on deep reinforcement learning,

[0059] The joint regulation system includes an edge device energy consumption optimization module, an edge-cloud collaborative inference framework construction module, a DVFS optimization algorithm construction module, and an offloading module;

[0060] The edge device energy consumption optimization module is used to establish an edge device energy consumption minimization model;

[0061] The edge-cloud collaborative inference framework construction module constructs an edge-cloud collaborative inference framework based on DVFS based on the edge device energy consumption minimization model established by the edge device energy consumption optimization module to jointly optimize the energy consumption and end-to-end latency of the edge device;

[0062] The DVFS optimization algorithm construction module constructs an enhanced DVFS optimization algorithm DVFO based on DRL according to the edge-cloud collaborative inference framework of DVFS constructed by the edge-cloud collaborative inference framework construction module, and reduces the overall energy consumption of the edge device by jointly optimizing the computing frequency of the edge device and the offloading ratio of the feature map;

[0063] The offloading module is used to establish an offloading mechanism to solve the offloading bottleneck problem of the feature map and avoid the end-to-end delay of edge-cloud collaborative inference being directly affected by the size of the feature map.

[0064] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0065] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.

[0066] Advantages of the present invention

[0067] The present invention is used for DNN feature map offloading in cloud-edge collaborative inference to minimize the energy consumption of edge devices. A reinforcement learning algorithm using the "think while moving" concurrent strategy calculates the optimal offloading ratio of the feature map and the computing frequency of the edge device for each task. Experimental results on different datasets show that compared with the state-of-the-art offloading solutions, DVFO reduces the average energy consumption of edge devices by 33%. In addition, it also achieves an end-to-end delay reduction of up to more than 54%. Brief description of the drawings

[0068] Figure 1 It is the edge-cloud collaborative inference framework based on DVFS of the present invention;

[0069] Figure 2 It is the action trajectory under the blocking environment and concurrent environment of the present invention;

[0070] Figure 3 It is the feature map offloading module of the present invention;

[0071] Figure 4 It is a comparison graph of energy consumption and end-to-end delay on different datasets, where (a) is the comparison graph of average energy consumption and average delay of EfficientNet-B0 on the CIFAR-10 dataset, (b) is the comparison graph of average energy consumption and average delay of ViT on the CIFAR-10 dataset, (c) is the comparison graph of average energy consumption and average delay of EfficientNet-B0 on the PASCAL VOC-2012 dataset, and (d) is the comparison graph of average energy consumption and average delay of ViT on the PASCAL VOC-2012 dataset;

[0072] Figure 5 The figure shows the comparison of end-to-end delays under different network bandwidths. Among them, (a) is the comparison chart of the average delay of CIFAR-10 under different network bandwidths, and (b) is the comparison chart of the average delay of PASCAL VOC-2012 under different network bandwidths;

[0073] Figure 6 The figure shows the runtime overhead under different datasets. Among them, (a) is the comparison chart of the normalized overhead of the CIFAR-10 dataset, and (b) is the comparison chart of the normalized overhead of the PASCAL VOC-2012 dataset. Specific implementation manners

[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0075] Combined with Figures 1 to 6 。

[0076] A joint regulation method for dynamic frequency and deep learning model offloading based on deep reinforcement learning,

[0077] The method specifically includes the following steps:

[0078] Step 1, establish an edge device energy consumption minimization model;

[0079] Step 2, based on the edge device energy consumption minimization model established in Step 1, construct an edge-cloud collaborative inference framework based on DVFS to jointly optimize the energy consumption and end-to-end delay of the edge device;

[0080] Step 3, according to the edge-cloud collaborative inference framework based on DVFS constructed in Step 2, construct an enhanced DVFS optimization algorithm DVFO based on DRL, and reduce the overall energy consumption of the edge device by jointly optimizing the computing frequency of the edge device and the offloading ratio of the feature map;

[0081] Step 4, establish an offloading mechanism to solve the offloading bottleneck problem of the feature map and avoid the direct impact of the size of the feature map on the end-to-end delay of edge-cloud collaborative inference.

[0082] Let χ be a task set composed of N independent and non-preemptive tasks, expressed as χ = {x 1 , x 2 , …, x N};

[0083] Task x iOverall energy consumption of the edge device

[0084] Among them is the computing energy consumption, is the offloading energy consumption;

[0085] Task x i The computing energy consumption of depends on the local computing time c of the edge device i and the computing power is expressed as

[0086] Among them is proportional to the square of the voltage V 2 and the computing frequency f i , that is,

[0087] Different from only considering the CPU frequency in the prior art, the present invention considers the frequencies of the CPU, GPU, and memory,

[0088] The frequency of the CPU is The frequency of the GPU is The frequency of the memory is Then the computing frequency

[0089] Task x i The offloading power consumption of is related to the network bandwidth B, the size m of the offloading feature map i and the offloading power , denoted as Similarly,

[0090] Then minimize the overall energy consumption of the edge device while meeting the minimum frequency for system operation:

[0091]

[0092] s.t.l i ≤d i

[0093] Among them l i =c i +o i , representing the end-to-end delay of task x i

[0094] In addition, in the present invention, the cloud server has sufficient computing resources to ensure the real-time performance of remote inference. For each x in the task set i ​, it is necessary to calculate the frequencies of the CPU, GPU, and memory corresponding to this task, as well as the offloading ratio of the feature map, which are respectively represented as and ξ. The present invention also assumes that the edge device immediately turns into an idle state with the minimum energy consumption after completing inference and offloading, so as to further save the runtime energy consumption of the edge device.

[0095] In step 2, as Figure 1 shown,

[0096] Design an edge-cloud collaborative inference framework for DVFS to jointly optimize the energy consumption and end-to-end latency of edge devices

[0097] The DVFS-based edge-cloud collaborative inference framework described above includes local inference on the edge device and remote inference on the cloud server;

[0098] Use WiFi on the local area network LAN as the communication medium to offload the feature map of the DNN from the edge device to the cloud server.

[0099] In step 2, it specifically includes the following steps,

[0100] Step 2.1, use a negligible-overhead feature evaluator on the edge device to judge the current features of the DNN; this is implemented based on the roofline model, that is, use the computational density to evaluate whether the current DNN belongs to a compute-intensive or memory-intensive workload;

[0101] Step 2.2, the DVFS module based on DRL, that is, DVFO, learns the DNN characteristics and network bandwidth, and sets the optimal computing frequency of the edge device and the offloading ratio of the feature map for each task;

[0102] After obtaining the offloading ratio of the feature map in step 2.3, it is necessary to partition the DNN to be deployed to the edge side and the cloud respectively;

[0103] Specifically, the present invention uses the DNN on the edge device to predict the local feature map obtained by the offloading controller.

[0104] Combine the remote DNN to predict the remaining feature maps to obtain the final inference output.

[0105] By reducing the complexity of the local DNN, the overall energy consumption of the edge device can be effectively reduced.

[0106] In step 3, it specifically includes the following steps:

[0107] Step 3.1, transform the optimization objective in formula (1) into a reward function in DRL and model it as a Markov decision process MDP;

[0108] The agent in the DRL consists of three components: state, action, and reward.

[0109] Specifically, the state space is as follows: at each moment t, the agent in the DRL constructs a state space. In the present invention, the type φ i of the workload x i and the currently measured network bandwidth B are defined as the state.

[0110] The type φ i of the workload x i and the currently measured network bandwidth B, these three discrete metrics constitute the state space, denoted as

[0111] Specifically, the action space is as follows: in the present invention, the computing frequency f i of the task x i and the offloading ratio ξ of the feature map are taken as actions, and the action space is denoted as where are respectively combinations of CPU, GPU, and memory computing frequencies;

[0112] For example, (1500, 900, 1200, 50) means that when the clocks of the CPU, GPU, and memory are set to 1500 MHz, 900 MHz, and 1200 MHz respectively, 50% of the feature maps are locally inferred on the edge device, and the remaining 50% of the feature maps are offloaded to the cloud server for computing. To reduce the complexity of the action space and accelerate convergence, the present invention sets both the computing frequency and the offloading ratio to discrete values. Specifically, the present invention uniformly sets 10 frequency levels for the three computing frequencies respectively between the minimum frequency and the maximum frequency that satisfy the system operation.

[0113] Specifically, the reward is as follows: in the problem of optimizing the energy consumption of edge devices based on DVFS and edge-cloud collaborative inference, the objective of the present invention is to minimize the energy consumption of edge devices under the constraint of the lowest frequency for system operation. Based on the fact that the objective of the agent is to maximize the cumulative expected reward, the energy consumption optimization objective can be transformed into a reward function. To achieve this transformation, the reward function r of the agent is defined as follows:

[0114]

[0115] As Figure 2 (a) shows, most DRLs usually assume that the environmental state in which the agent makes decisions is static, meaning that the agent first observes the state and then makes a policy decision.

[0116] However, this sequential execution blocking method is not applicable to the real-world dynamic real-time environment.

[0117] Because after the agent observes the environmental state and executes an action, the state has "shifted", that is, the previous state has been transformed into a new unobserved state.

[0118] In particular, in a DVFS-based edge-cloud collaborative environment with strict time constraints, it is necessary to use DRL to adjust the computing frequency of edge devices and the offloading ratio of feature maps in real time according to the workload characteristics and network bandwidth. Therefore, reducing the decision-making overhead of DRL is crucial.

[0119] Step 3.2, the present invention uses the DQN reinforcement learning algorithm to control the computing frequency and offloading ratio of edge devices to reduce the decision-making overhead of DRL;

[0120] Based on the concurrent control mechanism of "thinking while moving at the edge" to reduce the decision-making overhead of DQN in discrete time. Figure 2 (b) depicts the main idea behind this method.

[0121] Specifically, the agent in the DRL observes the state s of the environment at time t i , when it selects an action , at the same time, the previous action has shifted to a new unobserved state (the dotted line in the figure), that is, the state capture and policy decision-making in the concurrent environment can be executed concurrently; where H is the duration of the action trajectory from state s t to ;

[0122] Step 3.3, the present invention realizes policy decision-making in a concurrent environment by modifying the standard DQN, and the DQN in the concurrent environment in discrete time is still convergent;

[0123] Therefore, the concurrent Q-value function of DQN can be reformulated as follows:

[0124]

[0125] In the said step 3.3, it specifically includes:

[0126] Algorithm 1 details the optimization process of DVFO.

[0127] Step 3.3.1, specifically, first initialize the neural network parameters and experience buffer in the DRL; then use the DNN feature map type obtained by the feature evaluator and the current network bandwidth as the initial state of the DRL;

[0128] Step 3.3.2, at the beginning of training, the agent in the DRL will randomly select an action;

[0129] At time t, the agent captures state s in a concurrent environment of discrete time t , and at the same time, the agent selects an action a according to the concurrent mechanism of "thinking while moving" t ;

[0130] Step 3.3.3, the present invention uses an ∈-greedy strategy to explore the environment, feeds back the frequencies of action selection and offloading rates to the frequency and offloading controllers respectively, and obtains an immediate reward r;

[0131] Step 3.3.4, next, set the state from s t to s t+1 , and store the current state, action, reward, and the state at the next time as an action trajectory in the experience buffer;

[0132] Step 3.3.5, at each gradient step, first randomly sample a mini-batch of historical trajectories from the experience buffer, then calculate the Q value in the concurrent environment using formula (3), update the network parameters according to gradient descent, and finally deploy the offline-trained DVFO online to evaluate the performance.

[0133]

[0134]

[0135] In step 4, it specifically includes the following steps:

[0136] Figure 3 Shows the offloading process of the feature map.

[0137] Step 4.1, the present invention first determines the number of DNN partition points; for a DNN with N layers, there are a total of 2 N-1 partitioning strategies. However, considering that the computing resources of the cloud server are sufficient and the communication cost between the edge and the cloud cannot be ignored, it is unnecessary to partition the DNN multiple times. The present invention only selects one partition point;

[0138] Step 4.2, secondly, it is necessary to determine which layer in the DNN should be partitioned. Here, partitioning is performed according to the DNN type given by the feature evaluator;

[0139] Specifically, for a compute-intensive DNN, select the layer with the highest computational load for partitioning;

[0140] Similarly, for a memory-intensive DNN, select the layer with the highest memory access volume for partitioning;

[0141] The above design can automatically select appropriate partitioning strategies for different types of DNNs to alleviate the high load on network bandwidth caused by offloaded feature maps.

[0142] Step 4.3, although model partitioning in edge-cloud collaboration reduces network bandwidth transmission, the original unoptimized feature maps still pose a challenge to network bandwidth. To further reduce offloading latency, the present invention adopts a precision quantization technique, that is, converting the 32-bit floating-point numbers of the feature maps into 8-bit fixed-length numbers. This method effectively reduces the transmission volume of the feature maps on the premise of ensuring that not too much information is lost.

[0143] A joint regulation system for dynamic frequency and deep learning model offloading based on deep reinforcement learning

[0144] The joint regulation system includes an edge device energy consumption optimization module, an edge-cloud collaborative inference framework construction module, a DVFS optimization algorithm construction module, and an offloading module;

[0145] The edge device energy consumption optimization module is used to establish an edge device energy consumption minimization model;

[0146] Based on the edge device energy consumption minimization model established by the edge device energy consumption optimization module, the edge-cloud collaborative inference framework construction module constructs an edge-cloud collaborative inference framework based on DVFS to jointly optimize the energy consumption and end-to-end latency of the edge device;

[0147] According to the edge-cloud collaborative inference framework based on DVFS constructed by the DVFS optimization algorithm construction module, the DVFS optimization algorithm construction module constructs an enhanced DVFS optimization algorithm DVFO based on DRL, and reduces the overall energy consumption of the edge device by jointly optimizing the computing frequency of the edge device and the offloading ratio of the feature maps;

[0148] The offloading module is used to establish an offloading mechanism to solve the offloading bottleneck problem of the feature maps and avoid the direct impact of the size of the feature maps on the end-to-end latency of edge-cloud collaborative inference.

[0149] Embodiment:

[0150] Datasets and DNN models.

[0151] The present invention evaluated DVFO on the CIFAR-10 and PASCAL VOC-2012 datasets respectively. These two datasets have images of different sizes, which can more comprehensively reflect the diversity of input data. Considering the limited computing resources of edge devices, the batch size was set to 1 during the inference process. In addition, the present invention selected EfficientNet-b0 and Visual Transformer (ViT-B16) to represent memory-intensive and compute-intensive DNNs respectively.

[0152] Baseline method.

[0153] The present invention compares DVFO with edge-only inference, cloud-only inference, and the current state-of-the-art edge-cloud collaboration methods (ApplealNet and DRLDO) to evaluate performance. All experimental results are averaged over the entire test dataset.

[0154] AppealNet: An edge-cloud collaborative inference framework for binary offloading that determines whether to use a lightweight model on an edge device or a complex model in a cloud server by judging the difficulty of the input data. Here, two more complex DNNs, EfficientNet-b1 and Visual Transformer (ViT-B32), are deployed in the cloud, corresponding to the lightweight DNNs on the edge side respectively.

[0155] DRLDO: A DVFS-based offloading framework that jointly optimizes the CPU frequency of a local device and the data offloading volume using DRL to reduce device energy consumption.

[0156] Cloud-only: All feature maps are offloaded to the cloud for inference. Here, the original feature maps are also compressed using precision quantization.

[0157] Edge-only: Tasks are only inferred on local edge devices without considering computational offloading.

[0158] DVFO performance comparison:

[0159] First, the end-to-end latency and runtime energy consumption of DVFO and the baseline method were compared on different datasets. The present invention uses the lightweight bandwidth control tool trickle to set the transmission rate of the network bandwidth to 5 Mbps. Figure 4 The performance comparison of inferring EfficientNet-B0 and Visual Transformer (ViT) using local and edge-cloud collaboration methods respectively on different datasets is shown. It can be seen that DVFO is always better than all baseline methods. More specifically, the average runtime energy consumption of DVFO for inferring the two DNNs is reduced by 18%, 31%, 39%, and 43% compared to DRLDO, AppealNet, Cloud-only, and Edge-only respectively, while reducing the end-to-end latency by 15%-36%.

[0160] Impact of network bandwidth:

[0161] Network bandwidth dominates the offloading latency of feature maps. Therefore, it is necessary to study the reliability of DVFO under different network bandwidths. Different from cloud servers, due to power consumption and cost limitations, edge devices use WiFi modules with lower transmission rates. In addition, considering the edge computing scenario, the present invention limits the network bandwidth between 0.5 Mbps and 4 Mbps.

[0162] Figure 5 The results show that the end-to-end latency of all edge-cloud collaboration methods decreases with the increase of network bandwidth, which means that in the case of poor network bandwidth, these methods are restricted by poor communication conditions and tend to infer DNN on edge devices. Thanks to the precision quantization technology, even when the available network bandwidth is only 0.5 Mbps, the end-to-end latency of DVFO is better than the baseline method, and the average latency can be reduced by 36%. In addition, when the network bandwidth increases, the performance improvement of DVFO decreases. This means that network bandwidth is no longer the decisive factor, and the performance difference mainly comes from the adjustment of the computing frequency of local devices. In contrast, the performance of DRLDO and AppealNet depends largely on network bandwidth. This shows that DVFO can make better adaptive adjustments to the offloading rate and the computing frequency of local devices under network bandwidth fluctuations, so it has higher reliability.

[0163] Runtime overhead analysis:

[0164] DVFO introduces additional runtime overhead (i.e., using DRL to determine the appropriate computing frequency and offloading ratio), which cannot be ignored for dynamic real-time environments. As Figure 6 shown, the present invention compares the average running time of the proposed DVFO with the baseline method, and normalizes all results for easy comparison. The evaluation results show that DVFO has less runtime overhead than other methods. Especially when used on larger datasets such as PASCAL VOC-2012, its average runtime overhead is reduced by 50% and 29.5% compared with DRLDO and AppealNet respectively.

[0165] Table 1 shows the specific parameters of edge devices and cloud servers. The present invention uses an NVIDIA Xavier NX as the edge device and a workstation equipped with an NVIDIA RTX 3080 GPU as the cloud server. Since the present invention evenly sets 10 levels between the maximum and minimum computing frequencies of the CPU, GPU, and memory of the edge device to meet the system operation, there are a total of 1000 CPU-GPU-memory pairs.

[0166] Table 1 Specific parameters of edge-cloud collaborative inference devices

[0167]

[0168] The present invention implements DQN in a concurrent environment using PyTorch. Specifically, both the network and the target network of DQN with prioritized experience replay and ∈-greedy strategy are trained using the Adam optimizer. The entire network consists of three hidden layers and an output layer, with each hidden layer having 128, 64, and 32 ReLU activation units respectively. In addition, the learning rate, buffer size, and batch size are set to 0.0001, 1e6, and 256 respectively.

[0169] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0170] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.

[0171] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory for the method described in the present invention is intended to include but not be limited to these and any other suitable types of memory.

[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-definition digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0173] In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor, or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0174] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0175] The above has introduced in detail a method and system for jointly adjusting dynamic frequency and deep learning model offloading based on deep reinforcement learning proposed by the present invention, and expounded the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A joint regulation method for dynamic frequency and deep learning model offloading based on deep reinforcement learning, characterized in that: The method specifically includes the following steps: Step 1, establish an edge device energy consumption minimization model; In Step 1, Let be a task set consisting of N independent and non-preemptive tasks, denoted as ; Task Overall energy consumption of edge devices ; Among them For calculating the energy consumption is the offloading energy consumption Task The calculated energy consumption depends on the local computing time of the edge device and the computing power , expressed as = , wherein is proportional to the square of the voltage and the calculated frequency , that is ; The frequency of the CPU is , the frequency of the GPU is , the frequency of the memory is , then the computing frequency is ; Task offloading power consumption and network bandwidth as well as the size of the offloaded feature map and offloading power are related and denoted as Similarly ; minimize the overall energy consumption of the edge device while meeting the minimum frequency of system operation: (1) Among them , represented as the task 's end-to-end delay; Step 2, based on the edge device energy consumption minimization model established in Step 1, construct a DVFS-based edge-cloud collaborative inference framework to jointly optimize the energy consumption and end-to-end latency of the edge device; Step 3, according to the DVFS-based edge-cloud collaborative inference framework constructed in Step 2, construct an enhanced DVFS optimization algorithm DVFO based on DRL to reduce the overall energy consumption of the edge device by jointly optimizing the computing frequency of the edge device and the offloading ratio of the feature map; In Step 3, it specifically includes the following steps: Step 3.1, transform the optimization objective in the edge device energy consumption minimization model established in Step 1 into a reward function in DRL and model it as a Markov decision process MDP; The agent in the DRL consists of three components: state, action, and reward, The state space is specifically as follows: at each moment t, the agent in the DRL constructs a state space ; define the type of the workload and the currently measured network bandwidth B as the state; ​ Workload type and the current measured network bandwidth B, these three discrete metrics constitute the state space , denoted as ; The action space specifically is: the task computing frequency and the offloading ratio of the feature map are used as actions, and the action space is represented as , where ; The reward converts the energy consumption optimization goal into a reward function and defines the reward function of the agent as follows: (2) Step 3.2, use the DQN reinforcement learning algorithm to control the computing frequency and offloading ratio of the edge device to reduce the decision-making overhead of DRL; The agent in the DRL observes the state of the environment at time t , and when it selects an action , at the same time, the previous action has slid to a new unobserved state , that is, the state capture and policy decision in the concurrent environment can be executed concurrently; where H is the duration of the action trajectory from state to ; Step 3.3, modify the standard DQN to achieve policy decision-making in a concurrent environment, and the DQN in the discrete-time concurrent environment is still convergent; Therefore, the concurrent Q-value function of DQN can be reformulated as follows: (3); Step 4, establish an offloading mechanism to solve the offloading bottleneck problem of the feature map and avoid the size of the feature map directly affecting the end-to-end latency of edge-cloud collaborative inference.

2. The method according to claim 1, characterized in that: In Step 2, design a DVFS-based edge-cloud collaborative inference framework to jointly optimize the energy consumption and end-to-end latency of the edge device The DVFS-based edge-cloud collaborative inference framework includes local inference of the edge device and remote inference of the cloud server; Use WiFi on the local area network LAN as the communication medium to offload the feature map of the DNN from the edge device to the cloud server.

3. The method according to claim 2, characterized in that: In Step 2, it specifically includes the following steps, Step 2.1, use a negligible-overhead feature evaluator on the edge device to judge the current DNN features; that is, use the computational density to evaluate whether the current DNN belongs to a compute-intensive or memory-intensive workload; Step 2.2, the DRL-based DVFS module, namely DVFO, learns the DNN characteristics and network bandwidth to set the optimal computing frequency of the edge device and the offloading ratio of the feature map for each task; Step 2.3, after obtaining the offloading ratio of the feature map, it is necessary to partition the DNN to be deployed to the edge side and the cloud respectively; Combine the remote DNN to predict the remaining feature map to obtain the final inference output.

4. The method according to claim 3, characterized in that: In the said Step 3.3, it specifically includes: Step 3.3.1: Initialize the neural network parameters and the experience buffer in DRL; then use the DNN feature map type and the current network bandwidth obtained by the feature evaluator as the initial state of DRL. Step 3.3.2: At the beginning of training, the agent in DRL randomly selects an action. At time t, the agent captures the state in a concurrent environment with discrete time , and at the same time, the agent selects an action according to the concurrent mechanism ; Step 3.3.3: Use the ε-greedy policy to explore the environment, feedback the frequencies of action selection and offloading rates to the frequency and offloading controllers respectively, and obtain an immediate reward r. Step 3.3.4, set the status from to , and store the current status, action, reward, and the status at the next moment as an action trajectory in the experience buffer; Step 3.3.5: At each gradient step, first randomly sample a small batch of historical trajectories from the experience buffer, then calculate the Q value in the concurrent environment using formula (3), update the network parameters according to gradient descent, and finally deploy the offline-trained DVFO online to evaluate the performance.

5. The method according to claim 4, wherein: In step 4, it specifically includes the following steps: Step 4.1, determine the number of DNN partition points; for a DNN with N layers, there are a total of 2 N-1 partitioning strategies, and only select one partition point; Step 4.2: Determine which layer in the DNN should be partitioned, and perform partitioning according to the DNN type given by the feature evaluator. For a compute-intensive DNN, select the layer with the highest computational amount for partitioning. For a memory-intensive DNN, select the layer with the highest memory access amount for partitioning. Step 4.3: Adopt the precision quantization technique, that is, convert the 32-bit floating-point numbers of the feature map into 8-bit fixed-length numbers to reduce the latency.

6. A joint regulation system for dynamic frequency and deep learning model offloading based on deep reinforcement learning, wherein: It is used to execute the steps of the method described in any one of claims 1 to 5. The joint regulation system includes an edge device energy consumption optimization module, an edge-cloud collaborative inference framework construction module, a DVFS optimization algorithm construction module, and an offloading module. The edge device energy consumption optimization module is used to establish a model for minimizing the energy consumption of edge devices. The edge-cloud collaborative inference framework construction module constructs a DVFS-based edge-cloud collaborative inference framework based on the edge device energy consumption minimization model established by the edge device energy consumption optimization module to jointly optimize the energy consumption and end-to-end latency of edge devices. The DVFS optimization algorithm construction module constructs an enhanced DVFS optimization algorithm DVFO based on DRL according to the DVFS-based edge-cloud collaborative inference framework constructed by the edge-cloud collaborative inference framework construction module, and reduces the overall energy consumption of edge devices by jointly optimizing the computing frequency of edge devices and the offloading ratio of feature maps. The offloading module is used to establish an offloading mechanism to solve the offloading bottleneck problem of feature maps and avoid the direct impact of the size of feature maps on the end-to-end latency of edge-cloud collaborative inference.

7. An electronic device, including a memory and a processor, the memory stores a computer program, wherein, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.

8. A computer-readable storage medium for storing computer instructions, wherein, When the computer instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-objective optimization unloading strategy based on deep reinforcement learning under edge cloud architecture

    CN113626104A

  • Deep reinforcement learning-based information processing method and apparatus for edge computing server

    US20220255790A1