Profit-aware offloading framework towards prediction-assisted MEC network slicing

The SliceOff framework addresses the challenges of dynamic user traffic and service demands in MEC by decoupling optimization problems into edge network slicing and computation offloading, using predictive models and deep reinforcement learning for optimal resource allocation, thereby enhancing ESP profits and QoS.

WO2025107095A1PCT designated stage expired Publication Date: 2025-05-30FUZHOU UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/132499
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing solutions for Mobile Edge Computing (MEC) network slicing and computation offloading struggle to adapt to dynamic user traffic and service demands, leading to under-supply or over-supply of slice resources, which degrades Quality of Service (QoS) and profits for Edge Service Providers (ESPs).

Method used

The proposed SliceOff framework decouples the optimization problem of maximizing long-term ESP profits into sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA). It uses a gated recurrent neural network for prediction-assisted slice partitioning and an improved deep reinforcement learning with twin critic-networks and delay mechanism for optimal offloading and resource allocation.

Benefits of technology

SliceOff enhances ESP profits and achieves superior performance in different scenarios by accurately predicting user requests, optimizing slice partitioning, and addressing Q-value overestimation and high variance in computation offloading, resulting in improved resource utilization and reduced delay violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023132499_30052025_PF_FP_ABST
    Figure CN2023132499_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing is provided. Towards MEC network slicing, formulate the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA); For the slicing sub-problem, use a gated recurrent neural network (GRNN) to accurately predict user requests in different regions, and then use the optimal partitioning of network slices with the predicted requests and expected demands; For the offloading sub-problem, incorporating results from slice parti-tioning, use an improved deep reinforcement learning with twin ritic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation.
Need to check novelty before this filing date? Find Prior Art

Description

PROFIT-AWARE OFFLOADING FRAMEWORK TOWARDS PREDICTION-ASSISTED MEC NETWORK SLICINGTECHNICAL FIELDThe present invention belongs to the technical field of Artificial Intelligence (AI) and Mobile Edge Computing (MEC) , in particular relates to A Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing.BACKGROUNDWith the rapid development of Artificial Intelligence (AI) and 5G communication technologies, a range of intelligent applications have been created and penetrated into every aspect of our society such as image recognition, semantic segmentation, and autonomous driving [1] , exhibiting computation-intensive and latency-sensitive features. However, the limited computational capabilities of intelligent end devices seriously restrain their further development and popularity. To relieve this problem, Mobile Edge Computing (MEC) , deploying resources at the network edge close to users, has been deemed as a promising solution. Compared to Cloud Computing, MEC significantly reduces data transmission latency and thus improves the Quality of Service (QoS) [2] .For diverse applications, there are considerable differences in user service demands with respect to communication rate, response delay, and reliability [3] . Traditional network architectures with fixed configurations struggle to meet such demand variability. To address this issue, the emerging network slicing [4] , based on virtualization technologies including Network Function Virtualization (NFV) and Software Defined Networks (SDN) , divides physical network resources into multiple logically-isolated slices bringing better network management and orchestration. Using the network slicing technique, a multi-tenant ecosystem is created in MEC environments. Thus, Edge Service Providers (ESPs) are able to deploy services to proper slices based on system states and user demands [5] , offering customized configurations for network resources. Specifically, ESPs request slice resources from edge Infrastructure Providers (InPs) aligned with the demands of users' offloaded tasks to effectively alleviate the resource constraints of end devices. Empowered by partitioned slicing resources, ESPs can assist users in processing their offloaded tasks to enhance QoS. Despite being such a promising technology, most existing studies focus on slice partitioning in static environments [6] , [7] , neglecting the spatio-temporal variability of user traffic and service demands in real-world scenarios. This oversight can lead to under-supply or over-supply situations that significantly degrade the QoS and profits of ESPs. Therefore, ESPs are required to dynamically orchestrate the slice resources to rapidly response to such variable user traffic and service demands.Designing a reliable framework that effectively integrates network slicing and computation offloading is critically challenging. Existing studies on network slicing typically relied on predicting resources [8] , [9] . However, the changeable number of users and unknown task attributes in real-world complex MEC environments make accurate prediction extremely difficult, and thus these studies struggle to achieve adequate adaptiveness. Moreover, the computation offloading access under dynamic MEC environments is also a challenging problem to be solved

[0010] . To address this issue, existing studies of computation offloading access usually employed control theories or iterative algorithms

[0011] ,

[0012] . However, when the problem scale rises, the increasing computational complexity becomes unacceptable and causes excessive system overheads. Deep Reinforcement Learning (DRL) , an emerging branch of Machine Learning (ML) , has been preliminary applied to cope with the optimization problem of network slicing or computation offloading

[0013] ,

[0014] ,

[0015] . Through interacting with unknown environments, the DRL demonstrates great potential for making appropriate decisions on dynamic and uncertain optimization problems, ultimately maximizing long term rewards. A few DRL-based studies attempted to address network slicing and computation offloading simultaneously

[0016] ,

[0017] , but they struggle to effectively address the issues of Q-value overestimation and high variance, causing unstable convergence or sub-optimum. Moreover, these studies fail to effectively integrate user traffic fluctuations with MEC network slicing.RELATED WORKNetwork Slicing. Cheng et al. [6] proposed a two-stage network slicing model by predicting link traffic and correcting errors. Papa et al. [7] designed a Lyapunov optimization based slicing approach to satisfy individual throughput while ensuring slice isolation. Zhao et al. [8] developed a slicing algorithm based on information prediction and dynamic programming, aiming to maximize profits while realizing inter-slice isolation and intra-slice customization. Chiariotti et al. [9] proposed a frame-size prediction model for virtual-reality applications to optimize network slicing strategies, ensuring QoS in multi-party interactions. Cui et al.

[0013] designed a QoS-aware network slicing orchestration for Internet-of-Vehicles (IoV) , guaranteeing stable QoS for vehicles. Most of these studies relied on prior user demands and static resource provisioning, but they neglect dynamic user locations and uncertain resource demands. Therefore, the existing studies commonly encounter huge prediction difficulties, leading to the under-supply or over-supply slice resources in real-world MEC environments.Computation Offloading. Ren et al. [2] proposed a two layer task offloading collaboration model, optimizing internal load balancing and external task offloading based on the game theory to maximize the revenue of ESPs. Wang et al.

[0011] formulated the dynamic task offloading as a multi-armed bandit process and then designed a decentralized offloading method to optimize user rewards. Wu et al.

[0012] developed an offloading access algorithm based on the Lyapunov optimization, balancing the energy and delay in MEC systems with dynamically-changing network conditions. Duan et al.

[0014] proposed a deep Q-network based server grouping method for adaptive task offloading and load balancing in a scenario with unknown user movement. Hwang et al.

[0015] designed an improved DRL-based approach for energy-efficient offloading in the UAV-assisted MEC network. However, these studies did not consider some important factors in real-world offloading scenarios such as the diversity of user request patterns and the dynamics of MEC resources.Network Slicing with Computation Offloading. et al.

[0018] proposed a network slicing method for low-latency offloading with game theory, jointly managing the radio and computing resources for slices. Feng et al.

[0019] designed a network slicing framework for MEC systems, where the Liapunov optimization was used to make customized slicing to enhance the revenue of operators. These studies may perform well when facing static or stable scenarios, but it is hard for them to achieve the optimal network slicing and computation offloading in highly-dynamic MEC environments. Moreover, these studies required excessive iterations, causing huge computational overheads. Therefore, it is difficult for them to efficiently handle the large-scale problem of network slicing and computation offloading in complex MEC environments. Few studies utilized DRL to tackle the joint problem of network slicing and computation offloading. Shen et al.

[0016] designed a slicing-enabled task offloading framework for space-air-ground vehicular networks by using a service-oriented slicing and Double DQN algorithm. Chiang et al.

[0017] proposed a DQN-based network slicing framework to optimize slice scaling and task offloading, aiming to enhance QoS and the profits of service providers in edge systems. However, these studies evaluated state-action pairs by the maximized Q-value, leading to the overestimation, and the policy may fall into the sub-optimum or cannot achieve stable convergence due to the accumulated estimation errors and high variance.REFERENCES[1] R. Zhou, X. Wu, H. Tan, and R. Zhang, “Two time-scale joint service caching and task offloading for uav-assisted mobile edge computing, ” in IEEE Conference on Computer Communications (INFOCOM) , pp. 1189–1198, IEEE, 2022.[2] J. Ren, J. Liu, Y. Zhang, Z. Li, F. Lyu, Z. Wang, and Y. Zhang, “An efficient two-layer task offloading scheme for mec system with multiple services providers, ” in IEEE Conference on Computer Communications (INFOCOM) , pp. 1519–1528, IEEE, 2022.[3] B. Yin, J. Tang, and M. Wen, “Connectivity maximization in nonorthogonal network slicing enabled industrial internet-of-things with multiple services, ” IEEE Transactions on Wireless Communications (TWC) , 2023.[4] P. Promponas and L. Tassiulas, “Network slicing: Market mechanism and competitive equilibria, ” in IEEE Conference on Computer Communications (INFOCOM) , IEEE, 2023.[5] Y. Wu, H. Dai, H. Wang, Z. Xiong, and S. Guo, “A survey of intelligent network slicing management for industrial iot: Integrated approaches for smart transportation, smart energy, and smart factory, ” IEEE Communications Surveys&Tutorials, vol. 24, no. 2, pp. 1175–1211, 2022.[6] X. Cheng, Y. Wu, G. Min, A. Y. Zomaya, and X. Fang, “Safeguard network slicing in 5g: Alearning augmented optimization approach, ” IEEE Journal on Selected Areas in Communications (JSAC) , vol. 38, no. 7, pp. 1600–1613, 2020.[7] A. Papa, A. Jano, S. O. Ayan, H. M. G¨ursu, and W. Kellerer, “User-based quality of service aware multi-cell radio access network slicing, ” IEEE Transactions on Network and Service Management (TNSM) , vol. 19, no. 1, pp. 756–768, 2021.[8] P. Zhao, H. Tian, S. Fan, and A. Paulraj, “Information prediction and dynamic programming-based ran slicing for mobile edge computing, ” IEEE Wireless Communications Letters, vol. 7, no. 4, pp. 614–617, 2018.[9] F. Chiariotti, M. Drago, P. Testolina, M. Lecci, A. Zanella, and M. Zorzi, “Temporal characterization and prediction of vr traffic: A network slicing use case, ” IEEE Transactions on Mobile Computing (TMC) , 2023.

[0010] Q. Luo, S. Hu, C. Li, G. Li, and W. Shi, “Resource scheduling in edge computing: A survey, ” IEEE Communications Surveys&Tutorials, vol. 23, no. 4, pp. 2131–2165, 2021.

[0011] X. Wang, J. Ye, and J. C. Lui, “Decentralized task offloading in edge computing: A multi-user multi-armed bandit approach, ” in IEEE Conference on Computer Communications (INFOCOM) , pp. 1199–1208, IEEE, 2022.

[0012] H. Wu, Y. Sun, and K. Wolter, “Energy-efficient decision making for mobile cloud offloading, ” IEEE Transactions on Cloud Computing (TCC) , vol. 8, no. 2, pp. 570–584, 2018.

[0013] Y. Cui, X. Huang, P. He, D. Wu, and R. Wang, “Qos guaranteed network slicing orchestration for internet of vehicles, ” IEEE Internet of Things (IoT) Journal, vol. 9, no. 16, pp. 15215–15227, 2022.

[0014] S. Duan, F. Lyu, H. Wu, W. Chen, H. Lu, Z. Dong, and X. Shen, “Moto: Mobility-aware online task offloading with adaptive load balancing in small-cell mec, ” IEEE Transactions on Mobile Computing (TMC) , 2022.

[0015] S. Hwang, J. Park, H. Lee, M. Kim, and I. Lee, “Deep reinforcement learning approach for uav-assisted mobile edge computing networks, ” in IEEE Global Communications Conference (GLOBECOM) , pp. 3839–3844, IEEE, 2022.

[0016] H. Shen, Y. Tian, T. Wang, and G. Bai, “Slicing-based task offloading in space-air-ground integrated vehicular networks, ” IEEE Transactions on Mobile Computing (TMC) , 2023.

[0017] Y. Chiang, C. Hsu, G. Chen, and H. Wei, “Deep q-learning-based dynamic network slicing and task offloading in edge network, ” IEEE Transactions on Network and Service Management (TNSM) , vol. 20, no. 1, pp. 369–384, 2022.

[0018] S. Joˇsilo and G. D′an, “Joint wireless and edge computing resource management with dynamic network slice selection, ” IEEE / ACM Transactions on Networking (ToN) , vol. 30, no. 4, pp. 1865–1878, 2022.

[0019] J. Feng, Q. Pei, F. R. Yu, X. Chu, J. Du, and L. Zhu, “Dynamic network slicing and resource allocation in mobile edge computing systems, ” IEEE Transactions on Vehicular Technology (TVT) , vol. 69, no. 7, pp. 7863–7878, 2020.

[0020] C. Zhang, H. Zhang, J. Qiao, D. Yuan, and M. Zhang, “Deep transfer learning for intelligent cellular traffic prediction based on cross-domain big data, ” IEEE Journal on Selected Areas in Communications (JSAC) , vol. 37, no. 6, pp. 1389–1401, 2019.

[0021] F. Zhou, Y. Wu, R. Q. Hu, and Y. Qian, “Computation rate maximization in uav-enabled wireless-powered mobile-edge computing systems, ” IEEE Journal on Selected Areas in Communications (JSAC) , vol. 36, no. 9, pp. 1927–1941, 2018.

[0022] Y. Mao, J. Zhang, and K. B. Letaief, “Dynamic computation offloading for mobile-edge computing with energy harvesting devices, ” IEEE Journal on Selected Areas in Communications (JSAC) , vol. 34, no. 12, pp. 3590–3605, 2016.

[0023] S. Burer and A. N. Letchford, “Non-convex mixed-integer nonlinear programming: Asurvey, ” Surveys in Operations Research and Management Science, vol. 17, no. 2, pp. 97–106, 2012.

[0024] W. Wu, C. Zhou, M. Li, H. Wu, H. Zhou, N. Zhang, X. S. Shen, and W. Zhuang, “Ai-native network slicing for 6g networks, ” IEEE Wireless Communications, vol. 29, no. 1, pp. 96–103, 2022.

[0025] T. Afrin and N. Yodo, “A long short-term memory-based correlated traffic data prediction framework, ” Knowledge-Based Systems (KBS) , vol. 237, p. 107755, 2022.

[0026] S. Chen, L. Wang, and F. Liu, “Optimal admission control mechanism design for time-sensitive services in edge computing, ” in IEEE Conference on Computer Communications (INFOCOM) , pp. 1169–1178, IEEE, 2022.

[0027] S. Khuller, A. Moss, and J. S. Naor, “The budgeted maximum coverage problem, ” Information Processing Letters, vol. 70, no. 1, pp. 39–45, 1999.

[0028] G. Barlacchi, M. De Nadai, R. Larcher, A. Casella, C. Chitic, G. Torrisi, F. Antonelli, A. Vespignani, A. Pentland, and B. Lepri, “A multi-source dataset of urban life in the city of milan and the province of trentino, ” Scientific Data, vol. 2, no. 1, pp. 1–15, 2015.

[0029] T. Liu, L. Fang, Y. Zhu, W. Tong, and Y. Yang, “A near-optimal approach for online task offloading and resource allocation in edge-cloud orchestrated computing, ” IEEE Transactions on Mobile Computing (TMC) , vol. 21, no. 8, pp. 2687–2700, 2020.

[0030] S. Guo, J. Liu, Y. Yang, B. Xiao, and Z. Li, “Energy-efficient dynamic computation offloading and cooperative task scheduling in mobile cloud computing, ” IEEE Transactions on Mobile Computing (TMC) , vol. 18, no. 2, pp. 319–333, 2018.

[0031] Alibaba cloud edge node service pricing. Accessed Mar. 12, 2023. [Online] . https:  / / www. alibabacloud. com / help / en / ens / latest / 8caea3.

[0032] S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods, ” in International Conference on Machine Learning (ICML) , pp. 1587–1596, PMLR, 2018.SUMMARYThe purpose of the present invention is to provide a Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing.Considering in Mobile Edge Computing (MEC) , emerging network slicing and computation offloading enable Edge Service Providers (ESPs) to respond to diverse spatio-temporal distributions of user requests, with the aim of improving Quality-of-Service (QoS) and resource efficiency. However, dynamic system states and various request patterns seriously hinder their broader implementation in MEC systems. Existing solutions commonly rely on static resource slicing or prior system knowledge, lacking adaptability and thus causing unsatisfying QoS and resource provisioning.To address these important challenges, we propose SliceOff, a novel profit-aware offloading framework towards prediction-assisted MEC network slicing. For the slicing sub-problem, we design a gated recurrent neural network to accurately predict user requests in different regions, and then theoretically derive the optimal partitioning of network slices with the predicted requests and expected demands. For the offloading sub-problem, incorporating results from slice partitioning, we develop an improved deep reinforcement learning with twin critic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation. Using real-world testbed and datasets of user traffic, extensive experiments are conducted to validate the effectiveness of the proposed SliceOff. Compared to benchmark methods, the SliceOff enhances ESP profits and exhibits superior performance in different scenarios.To realize the above purpose, the technical solution of the present invention is as follows:A Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing:Towards MEC network slicing, formulating the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA) ; For the slicing sub-problem, use a gated recurrent neural network (RNN) to accurately predict user requests in different regions, and then use the optimal partitioning of network slices with the predicted requests and expected demands; For the offloading sub-problem, incorporating results from slice partitioning, use an improved deep reinforcement learning with twin critic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation.Furthermore, wherein formulate the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA) is as follows:When assessing the profits of the ESP, it is essential to take into consideration both the revenues and costs of processing tasks; On the one hand, the ESP receives revenues from users according to the provided services; If a user task can be completed within its maximum tolerable delay Tmax, the ESP will receive revenue Φ; Otherwise, there is no revenue; The revenue received from ui within t defined asAfter completing tasks with different priorities, the ESP receives various revenues. The revenues within h are defined asOn the other hand, the ESP needs to pay for the rented resources; The costs of renting resources within h are defined aswhereandare bandwidth and computing resources rented by the ESP, respectively; ζ b and ζ f are the unit price of bandwidth and computing resources, respectively;Considering our goal is to maximize the long-term ESP profits, and thus the optimization problem is defined aswhere C1 and C2 represent that the bandwidth and computing resources allocated to the ESP cannot exceed total system resources. C3 represents that the offloading request can onlybe accepted or rejected by the ESP; C4 and C5 represent that the bandwidth and computing resources allocated to users cannot exceed the available resources of the ESP;Because the allocation of bandwidth and computing resources is a continuous decision-making process and the offloading decision is an integer variable, P1 is a mixed integer nonlinear programming (MINLP) problem; To effectively relieve this issue, we decouple P1 into the sub-problems of Edge network slicing (EnS) and Computation offloading Access (CoA) as follows:· P1.1 (EnS) : Maximize ESP profits in long time slots by partitioning network slices; This sub-problem is defined as:· P1.2 (CoA) : Maximize the ESP revenues in short time slots by conducting computation offloading and resource allocation; This sub-problem is defined asFurthermore, wherein use a gated recurrent neural network (GRNN) to accurately predict user requests in different regions, and then use the optimal partitioning of network slices with the predicted requests and expected demands is as follows:As different regions have varying bandwidth demands, based on the predicted user request traffic, adopt the historical average value (HAV) to calculate the expected resource demands, and use Long Short-Term Memory (LSTM) to predict user request traffic in these regions; Finally, theoretically derive the optimal slice partitioning through combining the predicted user request traffic and expected resource demands.Furthermore, execute Algorithm 2: After inputting the historical request traffic Z, initialize the learning rate γ, input length Lc, and prediction length Lp, where Lp ≥ T; For each prediction window, inputs the historical user traffic Z τ into the LSTM cell to predict the traffic for future Lp time slots; Specifically, the LSTM cell controls the information flow into neural networks through the forget, input, and output gates;First, Z τ is used to update the forget gate f τ and the input gate i τ, where f τ determines the information that was forgotten at the previous moment and i τ determines the new informationthat will be stored in the current cell state; Then, the cell candidate stateτ is defined asNext, cell state Cτ and the output gateare updated, and then the output of the hidden layer Hτ is updated; After analyzing all the data in historical windows, the prediction of user traffic for m regions in Lp short time slots can be obtained, denotedbywhereFor each historical short time slot t, the HAV is used to calculate the expected bandwidth E [bj] and user priority E [ρj] for the completed tasks, and to calculate the bandwidth demandcomputing demandand expected revenue Rt forthe ESP based onE [bj] , E [ρj] , andfedge; Finally, the optimalandare derived by Lemma 1 and used for slice partitioning;Lemma 1. ESPprofits can be maximized whenandFurthermore, where in for the offloading sub-problem, incorporating results from slice partitioning, use an improved deep reinforcement learning with twin critic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation is as follows:ESP profits while meeting QoS, regard the problem model of P1.2 as the environment, the DRL agent optimizes policies through continuously interacting with the environment, which can be formulated as a Markov decision process; Specifically, the state space, action space, and reward function are defined as follows:State Space: The state space contains the available resources of the ESP and task attributes in the current time slot; To better capture demand features, convert the data size and computing density of tasks into the demands for uploading rate and computational frequency; Thus, the system state at t is defined asAction Space: design a probability distribution function for the discrete action space of offloading access; Specifically, the action space at t indicates the bandwidth allocation for each user, and it is defined asat=bt,whereNext, the action of offloading access is defined asIf the allocated bandwidth is positive, a request for offloading access will be accepted along with the corresponding bandwidth allocation; Otherwise, the request will be rejected and the task will be executed locally;Reward Function: The optimization objective of P1.2 is to maximize the cumulative ESP revenues from users; Therefore, the reward function indicates ESP revenues, and it is defined as Based on the above definitions, execute Algorithm 3: First, we initialize the online networks including two critic-networks Q1 and Q2 and the actor-network μ, and two target critic-networks Q1' and Q2' and the target actor-network μ'; For each training epoch, the environment is first initialized; At each short time slot t, the users send their requests of computation offloading access to the ESP; For these requests, the state st is fed into the actor-network μ, and then the DRL agent explores the action of resource allocation at in the current state according to μ and exploration noise; Next, the actions of bandwidth allocation bt and computation offloading access xt are obtained according to Eqs at=bt andAfter completing bandwidth allocation and computation offloading, the environment provides feedback in the form of immediate reward and the next state; Next, the samples of state transition are stored in the replay buffer, where K training samples are randomly selected for updating network parameters; When updating the critic-network, the actionat st+1 is first obtained by the target actor-network; This process can be described aswhere the network noise ε is a regularization that makes similar actions own comparable rewards;Then, the target Q-value is obtained by using the reward and comparing two critic-networks; This process can be described asNext, the two critic-networks are updated; To reduce the updating frequency of low-quality policies, we design a delay mechanism to update the actor-network and target networks; If t mod 2=0, the actor-network is updated using gradient ascent, and the target networks are updated using soft updates.Furthermore, for each long time slot, first invokes Algorithm 2 to predict the demands of bandwidth and computing resources for task execution (i.e., and) , and then performs network slicing and calculate resource costs based on the partitioned slice resources; For each short time slot, after users send requests of computation offloading access to the ESP, invokes Algorithm 3 to generate offloading access and bandwidth allocation decisions (i.e., xt and bt) based on the current system state and user demands; Next, users' tasks are processed according to offloading decisions and the ESP receives revenues; Then, calculates ESP profits based on the costs and revenues; Finally, goes to the next long time slot.Compared with the prior art, the present invention has the following beneficial effects:1. We propose a new two-timescale based computation offloading access model towards MEC network slicing. Specifically, we formulate the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA) .2. For EnS, we design a novel prediction-assisted slice partitioning method. First, we adopt a gated recurrent neural network to accurately predict future user requests in different regions. Next, we calculate expected demands via the historical average value (HAV) . Finally, we theoretically derive the optimal partitioning of network slices based on user requests and expected demands.3. For CoA, we develop an improved DRL-based offloading method. First, we prove that the CoA is NP-hard. Next, we design a dedicated Markov model for the action space of CoA. Finally, incorporating results from slice partitioning in EnS, twin critic-networks and a delay mechanism are designed in the improved DRL to address the Q-value overestimation and high variance, enabling near-optimal offloading and resource allocation.4. Using real-world testbed and datasets of user traffic, extensive experiments are conducted to verify the effectiveness of the proposed SliceOff. Compared to benchmark methods, the SliceOff improves ESP profits while achieving higher resource utilization and lower delay violation. Further, testbed experiments validate the superiority of the SliceOff, which mitigates the unbalanced performance caused by diverse spatio-temporal distributions of user requests and reduces task execution time.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 is a computation offloading access system towards MEC network slicing;FIG. 2 is an improved DRL for offloading access and resource allocation;FIG. 3 shows CDF of user traffic;FIG. 4 shows prediction of user traffic;FIG. 5 shows a convergence comparison;FIG. 6 shows changes of different indexes;FIG. 7 shows profit with various prices;FIG. 8 shows profit with various costs;FIG. 9 shows RU with various frequencies;FIG. 10 shows DVR with various tolerances;FIG. 11 shows Task execution time with different bandwidth allocation strategies;FIG. 12 shows Task execution time with different offloading methods.DETAILED DESCRIPTION OF THE EMBODIMENTSThe technical solution of the present invention is described in detail in combination with the accompany drawings.SYSTEM MODEL AND PROBLEM FORMULATIONFig. 1 illustrates the proposed computation offloading access system towards MEC network slicing. In this system, the Base Station (BS) and MEC server offer network and computing resources for processing the offloaded tasks from the intelligent applications of users. Users are randomly distributed within the communication coverage of the BS, which is divided into several regions based on the distance to the BS. The ESP first sends slicing requests with demanded resources to the InP and makes payment. Next, the ESP deploys the offloading service on the MEC server and manages the owned resources and system states. Finally, by paying fees, users can access the offloading service for processing their tasks.The total bandwidth of the BS and the total computational capability of the MEC server are denoted as Bmax and Fmax, respectively. Users within the coverage of BS are denoted as the set U= {u1, u2, . . ., un} , and regions are denoted as the set Reg= {reg1, reg2, . . ., regm} . Due to user mobility, the number of user requests in different regions and time slots experiences fluctuations, leading to an uneven spatio-temporal distribution of service demands

[0020] . To avoid frequent slice adjustment and simulate real-world scenarios, we adopt two scales of time slots to cope with the problems with different dimensions. Specifically, a long time slot is denoted as h∈ {1, 2, . . ., H} . At the beginning of h, the ESP evaluates the required resources in the current system and sends a request for network slicing to the InP. h is divided into several short time slots, denoted as t∈ {1, 2, . . ., T} . At the beginning of t, users send requests for computation offloading access to the ESP. The ESP evaluates the resource demands and priorities of users and then makes proper policies for computation offloading access and resource allocation.A. Computation ModelA task from ui is defined as a 5-tuple<di, ηi, ρi, regi, li>, where di is the data size, ηi is the computing density, ρi is the priority of ui, regi is the region where ui stays, and li is the distance from ui to the BS. When ui sends an offloading request, the ESP will decide whether to accept the request.Local Mode. When the request is rejected, ui executes the task locally, and the task execution time is defined aswhere floc is the computational capability of ui.Edge Mode. When the request is accepted, ui offloads the task to the MEC server for execution, and the input data should also be uploaded. When the ESP allocates the bandwidthto ui, according to Shannon’s formula

[0021] , the upload rate is defined aswherep is the uploadpower, σ2 is the noise power, is the channel power gain between ui and the BS.Compared to the input data, the output data is typically small and negligible

[0022] . Therefore, the task execution time in the edge mode is defined aswhere f edge indicates the computational capability allocated by the MEC server.By considering the above two modes, the task execution time can be described aswhereis the offloading access decision of ui at t.B. Profit ModelWhen assessing the profits of the ESP, it is essential to take into consideration both the revenues and costs of processing tasks. On the one hand, the ESP receives revenues from users according to the provided services. If a user task can be completed within its maximum tolerable delay Tmax, the ESP will receive revenueΦ. Otherwise, there is no revenue. The revenue received from ui within t defined asAfter completing tasks with different priorities, the ESP receives various revenues. The revenues within h are defined asOn the other hand, the ESP needs to pay for the rented resources. The costs of renting resources within h are defined aswhereandare bandwidth and computing resources rented by the ESP, respectively; ζb and ζf are the unit price of bandwidth and computing resources, respectively.C. Problem FormulationOur goal is to maximize the long-term ESP profits, and thus the optimization problem is defined aswhere C1 and C2 represent that the bandwidth and computing resources allocated to the ESP cannot exceed total system resources. C3 represents that the offloading request can onlybe accepted or rejected by the ESP; C4 and C5 represent that the bandwidth and computing resources allocated to users cannot exceed the available resources of the ESP;Because the allocation of bandwidth and computing resources is a continuous decision-making process and the offloading decision is an integer variable, P1 is a mixed integer nonlinear programming (MINLP) problem

[0023] ; To effectively relieve this issue, we decouple P1 into the sub-problems of Edge network slicing (EnS) and Computation offloading Access (CoA) as follows:· P1.1 (EnS) : Maximize ESP profits in long time slots by partitioning network slices; This sub-problem is defined as:· P1.2 (CoA) : Maximize the ESP revenues in short time slots by conducting computation offloading and resource allocation; This sub-problem is defined asTHE PROPOSED SliceOffA. Overview of the SliceOffWe propose SliceOff, a novel profit-aware offloading framework towards prediction-assisted MEC network slicing, with the aim of maximizing ESP profits by addressing the sub problems P1.1 and P1.2. For P1.1, the SliceOff first conducts in-depth analysis and accurate prediction of user traffic and resource demands, and then theoretically derives the optimal slice partitioning for long time slots. For P1.2, the SliceOff generates proper policies of computation offloading access and resource allocation for short time slots.The main workflow of the SliceOff is outlined in Algorithm 1. For each long time slot, the SliceOff first invokes Algorithm 2 to predict the demands of bandwidth and computing resources for task execution (i.e., and) , and then performs network slicing and calculate resource costs based on the partitioned slice. For each short time slot, after users send requests of computation offloading access to the ESP, the SliceOff invokes Algorithm 3 to generate offloading access and bandwidth allocation decisions (i.e., xt and bt) based on the current system state and user demands. Next, users' tasks are processed according to offloading decisions and the ESP receives revenues; Then, calculates ESP profits based on the costs and revenues. Finally, the SliceOff goes to the next long time slot.B. Prediction-assisted Slice PartitioningCommonly, the performance of slice resource allocation can be greatly improved by analyzing features of user requests and predicting future resource demands

[0024] . In light of this idea, we will solve P1.1 by conducting slice partitioning based on the prediction of future resource demands. As different regions have varying bandwidth demands, we first analyze the historical data and use Long Short-Term Memory (LSTM) , an improved Recurrent Neural Network (RNN) , to predict user request traffic in these regions. LSTM extracts temporal correlations in sequences, solving gradient vanishing and has proven useful for traffic prediction

[0025] . However, resource demands cannot be directly obtained from tasks in real-world scenarios due to their uncertain attributes. To address this problem, based on the predicted user request traffic, we further adopt the historical average value (HAV) to calculate the expected resource demands. Finally, we theoretically derive the optimal slice partitioning through combining the predicted user request traffic and expected resource demands. The key steps of the proposed prediction-assisted slice partitioning method are outlined in Algorithm 2.After inputting the historical request traffic Z, initialize the learning rate γ, input length Lc, and prediction length Lp, where Lp≥T; For each prediction window, inputs the historical user traffic Zτ into the LSTM cell to predict the traffic for future Lp time slots. Specifically, the LSTM cell controls the information flow into neural networks through the forget, input, and output gates;First, Zτ is used to update the forget gate fτ and the input gate iτ, where fτ determines the information that was forgotten at the previous moment and iτ determines the new information that will be stored in the current cell state. Then, the cell candidate stateis defined asNext, cell state Cτ and the output gateare updated, and then the output of the hidden layer Hτ. After analyzing all the data in historical windows, the prediction of user traffic for m regions in Lp short time slots can be obtained, denoted bywhereFor each historical short time slot t, the HAV is used to calculate the expected bandwidth E [bj] and user priority E [ρj] for the completed tasks, and to calculate the bandwidth demandcomputing demandand expected revenue Rt for the ESP based onE [bj] , E [ρj] , andfedge; Finally, the optimalandare derived by Lemma 1 and used for slice partitioning;Lemma 1. ESP profits can be maximized whenandProof 1. For clarity, andare replaced by B and Bt, respectively. Due to common rationality

[0026] , the revenue obtained by ESP for providing offloading services should be more than the cost of resources, thusalways holds. According to Eq. (5) , the ESP can obtain revenue only when the resources allocated to the task enable it to be completed within the maximum tolerable delay. Therefore, the revenues that the ESP obtains within a short time slot are equal to the ratio between total revenues and resource demands.Let B1≤B2≤...≤BT, the ESP profits within long time slots are defined asWhenWhen B∈ [B1, B2] , for B=B2,LetBecausethusSimilarly, when B∈ [BT-1, BT] , for B=BT,We can derive thatThus, when B=BT (i.e., ) , ESP profits can be maximized. Similarly, whenESP profits can be maximized. In summary, Lemma 1 is proved.C. Improved DRL for Computation Offloading Access and Resource AllocationP1.2 can be transferred into a classic budgeted maximum coverage problem (BMCP) that has been proven to be NP-hard

[0027] . The BMCP considers a set E= {e1, e1, . . ., en} , where elements are with costs and values. The objective of solving the BMCP is to select a subsetso that the total values are maximized without exceeding the cost budget. For P1.2, offloading requests can be regarded as elements in E, costs can be regarded as allocated resources, and values of elements can be regarded as revenues obtained from completing tasks. Therefore, we aim to seek the setthat can maximize the total revenue Rh without exceeding the costsandBased on the above analysis and derivation, we prove that P1.2 is an NP-hard problem.To address this problem, we propose an improved DRL that can adaptively make the optimal policy of computation offloading access and resource allocation, aiming to maximize ESP profits while meeting QoS. As shown in Fig. 2, we regard the problem model of P1.2 as the environment, the DRL agent optimizes policies through continuously interacting with the environment, which can be formulated as a Markov decision process. Specifically, the state space, action space, and reward function are defined as follows.State Space. The state space contains the available resources of the ESP and task attributes in the current time slot. To better capture demand features, we convert the data size and computing density of tasks into the demands for uploading rate and computational frequency. Thus, the system state at t is defined asAction Space. We design a probability distribution function for the discrete action space of offloading access. Specifically, the action space at t indicates the bandwidth allocation for each user, and it is defined asat=bt,          (18)whereNext, the action of offloading access is defined asIf the allocated bandwidth is positive, a request for offloading access will be accepted along with the corresponding bandwidth allocation. Otherwise, the request will be rejected and the task will be executed locally.Reward Function. The optimization objective of P1.2 is to maximize the cumulative ESP revenues from users. Therefore, the reward function indicates ESP revenues, and it is defined as Based on the above definitions, we propose an improved DRL-based computation offloading access and resource allocation method. The key steps are outlined in Algorithm 3.First, we initialize the online networks including two critic-networks Q1 and Q2 and the actor-network μ, and two target critic-networks Q1' and Q2' and the target actor-network μ'. Different from the classic DRL that uses the maximized Q-values for evaluation, the proposed method introduces two independent critic-networks to approximate the Q-value function with smaller values. This design helps alleviate the Q-value overestimation and prevents the algorithm from getting stuck in sub-optimal solutions due to excessive cumulative errors. For each training epoch, the environment is first initialized. At each short time slot t, the users send their requests of computation offloading access to the ESP. For these requests, the state st is fed into the actor-network μ, and then the DRL agent explores the action of resource allocation at in the current state according to μ and exploration noise. Next, the actions of bandwidth allocation bt and computation offloading access xt are obtained according to Eqs. (18) and (19) . After completing bandwidth allocation and computation offloading, the environment provides feedback in the form of immediate reward and the next state. Next, the samples of state transition are stored in the replay buffer, where K training samples are randomly selected for updating network parameters. When updating the critic-network, the actionat st+1 is first obtained by the target actor-network. This process can be described aswhere the network noise ε is a regularization that makes similar actions own comparable rewards.Then, the target Q-value is obtained by using the reward and comparing two critic-networks. This process can be described asNext, the two critic-networks are updated. To reduce the updating frequency of low-quality policies, we design a delay mechanism to update the actor-network and target networks. If t mod 2=0, the actor-network is updated using gradient ascent, and the target networks are updated using soft updates. Thus, the actor-network is updated more frequently than the critic-network. Compared to frequent network updates, this manner reduces cumulative errors and thus improves training stability.D. Complexity AnalysisThe proposed SliceOff consists of two main phases including offline training and online decision-making.Offline Training. There are M iterations for training the LSTM in Algorithm 2, each iteration contains H long time slots, the length of the prediction window is Lp, and thus the complexity of training the LSTM is O (MHLp) . There are Ntraining epochs in Algorithm 3, each training epoch contains H long time slots, each long time slot contains T short time slots, and thus the training complexity is O (NH T) .Online Decision-making. For each long time slot, the complexity of predicting traffic is O (Lp) , the complexity of slice partitioning is O (T) , the complexity of the improved DRL is O (T) , and thus the complexity of the online decision-making is O (Lp+T) .Through the above analysis, it is noted that the SliceOff owns low complexity. Therefore, it is able to quickly adjust the policy of network slicing and computation offloading and well fit in complex MEC scenarios with different problem scales.PERFORMANCE EVALUATIONIn this section, we evaluate the proposed SliceOff by conducting extensive simulation and testbed experiments.A. Experiment SetupExperimental Environment and Datasets. Based on a workstation equipped with an 8-core Intel (R) Xeon (R) Silver 4208 CPU@3.2GHz, 2 NVIDIA GeForce RTX 3090 GPUs, and 32GB RAM, we build a simulation environment for the proposed system and implement the SliceOff by using PyTorch. The real-world datasets of Milan cellular traffic

[0028] are used to construct dynamic user requests. The datasets contain three types of services including message, call, and Internet. The user traffic over two months was recorded with a sampling frequency of 10 min. Specifically, we select 3 regions (ID=4259, 4456, and 5060) and regard the traffic of Internet service recorded each time as the number of user requests in a short time slot. Moreover, an epoch contains 24 long time slots and each long time slot contains 6 short time slots.Parameter Settings. The communication coverage of the BS is with a radius of 3 km, which is divided into 3 regions with different distances to the BS (i.e., 0~1.5 km, 1.5~2.5 km, and 2.5~3 km) , corresponding to the selected 3 regions in the datasets. Moreover, Bmax=15 MHz, Fmax=30 GHz, floc=1.0 GHz, fedge=2.0 GHz, p=100 mW, β0=-60 dB, θ=2, and σ2=-110 dBm

[0029] . For a task, di∈ [200, 500] KB, ηi∈ [1, 10]

[0030] , ρi∈ {1, 2, 3} , T max=0.8 s, Φ=1.0$, ζb=0.28$ / Mbps, ζf=0.67$ / GHz

[0031] . For SliceOff, γ=0.001, Lc=12, Lp=6, and the learning rate is 0.001

[0032] .Performance Metrics. Except for ESP profits, we use the following metrics to further evaluate the SliceOff.· Resource Utilization (RU) : The rate between the actually used resources of executing offloaded tasks and the resources allocated to the ESP.· Deadline Violate Rate (DVR) : The rate of the number of tasks whose execution time exceeds the maximum tolerable delay and the total number of tasks.Comparison Methods. We compare the SliceOff with the following benchmark methods to verify its superiority.· SliceDDPG

[0013] : The LSTM-based bandwidth prediction is used for network slicing and the DDPG is used to handle computation offloading and resource allocation.· Off

[0015] : An improved DDPG is used for computation offloading under fixed slice resources, where half of the total resources are allocated to the ESP.· MEC: Half of the total resources are allocated to the ESP that accepts all offloading requests, and the bandwidth is averagely allocated to users.· Local: All tasks are executed locally without considering the costs of MEC resources.B. Experiment Results and AnalysisTraffic Distribution and Prediction. Fig. 3 illustrates the distribution of user traffic in different regions. The three regions exhibit various cumulative probabilities of user traffic, indicating significant differences in the spatial distribution of user traffic. The comparison of the predicted values and real values in different regions is shown in Fig. 4, where the user traffic dynamically changes over time and presents a certain periodicity. This indicates the uneven temporal distribution of user traffic. The results show that the SliceOff can effectively capture the fluctuating trend of user traffic, resulting in excellent prediction performance within different short time slots for various regions.Convergence Analysis. Fig. 5 presents the convergence comparison of the SliceOff with other benchmark methods. The Local and MEC perform worse than the other three DRL based methods because they do not well consider changeable system states and diverse task attributes, resulting in the failure of many tasks due to exceeding the maximum tolerance delay. Compared to the Off that uses fixed slice resources, the SliceOff and SliceDDPG significantly increase ESP profits by dynamically adjusting slice resources based on the accurate prediction of user traffic. Compared to the SliceDDPG, the SliceOff demonstrates more stable convergence and higher rewards during the training process. This is because the SliceOff employs a delayed-update mechanism that avoids frequent network updates, reducing cumulative errors and thus improving training stability. Meanwhile, the SliceOff adopts two independent critic-networks, solving the problem of Q-value overestimation that exists in the SliceDDPG.Revenue-cost Trend. Fig. 6 depicts the trends of revenues and costs with the change of traffic in different long time slots. The ESP revenues decline with the decrease of user traffic. This is because the SliceOff reduces the resources allocated to the ESP to save costs and improve profits. In contrast, as the user traffic grows, the SliceOff increases the resources allocated to the ESP, enabling more tasks to be completed, and thus increasing ESP revenue. It is worth noting that the difference between revenues and costs is considered as profits. The variation of ESP profits along with user traffic demonstrates that the SliceOff can make proper offloading decisions and maintain high profits across a variety of scenarios.Profit Evaluation. Figs. 7 and 8 illustrate the impact of task price and resource cost on ESP profits achieved by different methods. As shown in Fig. 7, when the task price is low (e.g., the multiple is 0.5x) , the MEC ends up with negative profits because the cost of resources outweighs the revenue generated. As the task price increases, the ESP can earn more profits. The MEC outperforms the Local because the Local can only complete a few tasks, and thus the increase of task price has less impact on the Local. Compared to other benchmark methods, the SliceOff achieves higher profits in different scenarios, which demonstrates the superiority of the SliceOff in addressing the problems of slice partitioning and computation offloading. As shown in Fig. 8, the Local does not use MEC resources, and thus any change in resource costs does not affect profits. For all methods, ESP profits decrease with the increase of resource costs. The MEC exhibits the most obvious reduction in profits as the multiple grows to 1.5x, even leading to lower profits than the Local. In different resource cost scenarios, the SliceOff always achieves the best performance, verifying the superiority of the SliceOff in enhancing ESP profits.RU Analysis. Fig. 9 presents the RU achieved by three DRL-based methods under various allocated edge frequencies. The RU first increases and then decreases as the allocated frequency grows. This is due to when the allocated frequency is low, the required bandwidth to complete the offloaded tasks becomes high. In such a situation, the allocated bandwidth may not be sufficient to process all offloaded tasks, resulting in some tasks being executed locally, leading to a decreasing RU. When the allocated frequency is high, the limited MEC computational capability may reduce the number of offloaded tasks that can be processed, also causing a decrease in RU. The results reveal that properly enhancing the allocated edge frequency can improve the RU. The SliceOff outperforms the SliceDDPG and Off. The reason for this is that Off adopts fixed slice resources, lacking the ability to dynamic MEC resources. The SliceOff can utilize resources better by using two critic-networks, maintaining a higher RU than the SliceDDPG.DVR Analysis. Fig. 10 depicts the DVR achieved by three DRL-based methods under various task maximum tolerable delays. The DVR decreases significantly with the increases of the task maximum tolerable delay that determines the available time to complete tasks. When the task maximum tolerable delay is high (e.g., 1.0 s) , the DVR approaches 0, indicating that the available resources can well meet the demands of processing tasks. Compared to the Off and SliceDDPG, the SliceOff is able to complete more tasks within the maximum tolerable delay and achieve a lower DVR. This is because the SliceOff can adaptively make more rational decisions of computation offloading and resource allocation based on dynamic system states and changeable service demands.C. Testbed ValidationSettup of Real-world Testbed. By using hardware devices, we build a real-world testbed to further evaluate the feasibility and practicality of the proposed SliceOff. the testbed consists of three user devices (i.e., Raspberry 4B equipped with Broadcom BCM2711 SoC@1.5GHz, 4GB RAM, and Raspbian GNU / Linux 11 OS) and three MEC servers (i.e., Jetson TX2 equipped with quad-core Arm CortexA57 MP Core processor, 256-core NVIDIA Pascal GPU, 8GB RAM, and Ubuntu 18.04.6 LTS OS) . All the devices are connected to a 5GHz router, where the communication platform is constructed by using the Flask framework. We regard image classification as a service instance for computation offloading, and users generate image classification tasks with different data sizes and send offloading access requests in different time slots. If the requests are accepted, the images will be uploaded to MEC servers for processing. Otherwise, the task will be processed locally. For slice partitioning, we use the number of MEC servers as the basic partitioning units. For example, if there are insufficient computing resources, all tasks will be offloaded to one MEC server for processing. Otherwise, tasks will be offloaded to different MEC servers for processing. Since the data transmission time may be affected by real world channel environments, we consider the error between the allocated bandwidth and the actual one when calculating the image transmission time. Moreover, we place user devices at different locations in our lab. At each time slot, the data size and computational demand of tasks are proportional to the user traffic used in simulation experiments.Validation Results. Based on the real-world testbed, we first evaluate the task execution time by using different bandwidth allocation strategies, where all offloading requests are accepted. As shown in Fig. 11, when using fixed bandwidth allocation, the task execution time of users grows as their distance to the MEC server increases. This is because the long-distance data transmission in unstable real-world channel environments causes excessive transmission time. When using dynamic bandwidth allocation, the task execution time is comparable in different regions and lower than the fixed one. This demonstrates that the SliceOff can mitigate the performance imbalance caused by diverse user space distribution and thus reduce task execution time. Next, we conduct a comparison of the task execution time for all three users, using the SliceOff, MEC, and Local in different time slots. As shown in Fig. 12, the task execution time achieved by using the Local generally changes along with the user traffic, indicating that the variability in the temporal distribution of user traffic causes a significantly impact on Local. Compared to the Local and MEC, the SliceOff takes less time to process tasks. This is because the SliceOff considers both local and edge computing resources in a time slot, and it can adaptively adjust slice resources according to user traffic, balancing task execution time and system resource costs in different time slots.CONCLUSIONIn this paper, we propose SliceOff, a novel profit-aware offloading framework towards prediction-assisted MEC network slicing. In Sliceoff, we decouple the optimization problem of maximizing long-term ESP profits into the sub-problems of EnS and CoA. For EnS, we design a new prediction-assisted slice partitioning method to theoretically derive the optimal partitioning of network slices based on the accurate prediction of future user traffic. For CoA, we first prove that it is NPhard and then develop an improved DRL-based computation offloading and resource allocation method with twin critic-networks and delay mechanism. By using real-world testbed and datasets of user traffic, we demonstrate the effectiveness of the proposed SliceOff via extensive experiments. Compared to benchmark methods, the SliceOff shows better performance and more stable convergence in terms of ESP profits, RU, and DVR under different scenarios. Further, real-world testbed experiments validate the feasibility and practicality of the SliceOff, which mitigates the performance imbalance caused by diverse spatio-temporal distributions of user requests and thus improves the task execution time.The above are preferred embodiments of the present invention, and any change made in accordance with the technical solution of the present invention shall fall within the protection scope of the present invention if its function and role do not exceed the scope of the technical solution of the present invention.

Claims

1.A Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing: Towards MEC network slicing, formulate the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA) ; For the slicing sub-problem, use a gated recurrent neural network (GRNN) to accurately predict user requests in different regions, and then use the optimal partitioning of network slices with the predicted requests and expected demands; For the offloading sub-problem, incorporating results from slice parti-tioning, use an improved deep reinforcement learning with twin ritic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation.2.The Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing according to claim 1, wherein formulate the optimization problem of maximizing long-term ESP profits and decouple it into the sub-problems of Edge network Slicing (EnS) and Computation offloading Access (CoA) is as follows:When assessing the profits of the ESP, it is essential to take into consideration both the revenues and costs of processing tasks; On the one hand, the ESP receives revenues from users according to the provided services; If a user task can be completed within its maximum tolerable delay Tmax, the ESP will receive revenue Φ; Otherwise, there is no revenue; The revenue received from ui within t defined asAfter completing tasks with different priorities, the ESP receives various revenues. The revenues within h are defined asOn the other hand, the ESP needs to pay for the rented resources; The costs of renting resources within h are defined aswhereandare bandwidth and computing resources rented by the ESP, respectively; ζb and ζf are the unit price of bandwidth and computing resources, respectively;Considering our goal is to maximize the long-term ESP profits, and thus the optimization problem is defined aswhere C1 and C2 represent that the bandwidth and computing resources allocated to the ESP cannot exceed total system resources. C3 represents that the offloading request can only be accepted or rejected by the ESP; C4 and C5 represent that the bandwidth and computing resources allocated to users cannot exceed the available resources of the ESP;Because the allocation of bandwidth and computing resources is a continuous decision-making process and the offloading decision is an integer variable, P1 is a mixed integer nonlinear programming (MINLP) problem; To effectively relieve this issue, we decouple P1 into the sub-problems of Edge network slicing (EnS) and Computation offloading Access (CoA) as follows:· P1.1 (EnS) : Maximize ESP profits in long time slots by partitioning network slices; This sub-problem is defined as:· P1.2 (CoA) : Maximize the ESP revenues in short time slots by conducting computation offloading and resource allocation; This sub-problem is defined as3.The Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing according to claim 2, wherein use a gated recurrent neural network (GRNN) to accurately predict user requests in different regions, and then use the optimal partitioning ofnetwork slices with the predicted requests and expected demands is as follows:As different regions have varying bandwidth demands, based on the predicted user request traffic, adopt the historical average value (HAV) to calculate the expected resource demands, and use Long Short-Term Memory (LSTM) to predict user request traffic in these regions; Finally, theoretically derive the optimal slice partitioning through combining the predicted user request traffic and expected resource demands.4.The Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing according to claim 3, wherein:execute Algorithm 2: After inputting the historical request traffic Z, initialize the learning rate γ, input length Lc, and prediction length Lp, where Lp≥T; For each prediction window, inputs the historical user traffic Zτ into the LSTM cell to predict the traffic for future Lp time slots; Specifically, the LSTM cell controls the information flow into neural networks through the forget, input, and output gates;First, Zτ is used to update the forget gate fτ and the input gate iτ , where fτ determines the information that was forgotten at the previous moment and iτ determines the new informationthat will be stored in the current cell state; Then, the cell candidate stateis defined asNext, cell state Cτ and the output gateare updated, and then the output of the hidden layer Hτ is updated; After analyzing all the data in historical windows, the prediction of user traffic for m regions in Lp short time slots can be obtained, denoted bywhereFor each historical short time slot t, the HAV is used to calculate the expected bandwidth E [bj] and user priority E [ρj] for the completed tasks, and to calculate the bandwidth demand computing demandand expected revenue Rt forthe ESP based onE [bj] , E [ρj] , andf edge; Finally, the optimalandare derived by Lemma 1 and used for slice partitioning;Lemma 1. ESPprofits can be maximized whenand5.The Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing according to claim 4, wherein for the offloading sub-problem, incorporating results from slice partitioning, use an improved deep reinforcement learning with twin critic-networks and delay mechanism, solving the Q-value overestimation and high variance for approximating the optimal offloading and resource allocation is as follows:ESP profits while meeting QoS, regard the problem model of P1.2 as the environment, the DRL agent optimizes policies through continuously interacting with the environment, which can be formulated as a Markov decision process; Specifically, the state space, action space, and reward function are defined as follows:State Space: The state space contains the available resources of the ESP and task attributes in the current time slot; To better capture demand features, convert the data size and computing density of tasks into the demands for uploading rate and computational frequency; Thus, the system state at t is defined asAction Space: design a probability distribution function for the discrete action space of offloading access; Specifically, the action space at t indicates the bandwidth allocation for each user, and it is defined asat=bt,whereNext, the action ofoffloading access is defined asIf the allocated bandwidth is positive, a request for offloading access will be accepted along with the corresponding bandwidth allocation; Otherwise, the request will be rejected and the task will be executed locally;Reward Function: The optimization objective of P1.2 is to maximize the cumulative ESP revenues from users; Therefore, the reward function indicates ESP revenues, and it is defined asBased on the above definitions, execute Algorithm 3: First, we initialize the online networks including two critic-networks Q1 and Q2 and the actor-network μ, and two target critic-networks Q1'and Q2'and the target actor-network μ'; For each training epoch, the environment is first initialized; At each short time slot t, the users send their requests of computation offloading access to the ESP; For these requests, the state st is fed into the actor-network μ, and then the DRL agent explores the action of resource allocation at in the current state according to μ and exploration noise; Next, the actions of bandwidth allocation bt and computation offloading access xt are obtained according to Eqs at=bt andAfter completing bandwidth allocation and computation offloading, the environment provides feedback in the form of immediate reward and the next state; Next, the samples of state transition are stored in the replay buffer, where K training samples are randomly selected for updating network parameters; When updating the critic-network, the actionat st+1 is first obtained by the target actor-network; This process can be described aswhere the network noise ε is a regularization that makes similar actions own comparable rewards;Then, the target Q-value is obtained by using the reward and comparing two critic-networks; This process can be described asNext, the two critic-networks are updated; To reduce the updating frequency of low-quality policies, we design a delay mechanism to update the actor-network and target networks; If t mod 2=0, the actor-network is updated using gradient ascent, and the target networks are updated using soft updates.6.The Profit-aware Offloading Framework towards Prediction-assisted MEC Network Slicing according to claim 5,For each long time slot, first invokes Algorithm 2 to predict the demands ofbandwidth and computing resources for task execution (i.e., and) , and then performs network slicing and calculate resource costs based on the partitioned slice resources; For each short time slot, after users send requests of computation offloading access to the ESP, invokes Algorithm 3 to generate offloading access and bandwidth allocation decisions (i.e., xt and bt) based on the current system state and user demands; Next, users' tasks are processed according to offloading decisions and the ESP receives revenues; Then, calculates ESP profits based on the costs and revenues; Finally, goes to the next long time slot.

Citation Information

Patent Citations

  • Self-adaptive computing unloading method based on server load balancing mechanism in MEC environment

    CN115604274A

  • Computing unloading method for 5G network slice in MEC environment

    CN117202264A

  • Method and system for network slice allocation

    US20180317133A1

  • Edge network computing system with deep reinforcement learning based task scheduling

    US20230153124A1