Traffic-aware lightweight hierarchical offloading framework for adaptive slicing-enabled air-ground integrated networks

By dividing SAGIN into CAP and COP and utilizing sparse probabilistic self-attention and lightweight DRL algorithms, the inefficiency of resource management and computation offloading in SAGIN is solved, efficient resource utilization and service quality are achieved, and the complexity and cost of computation offloading are reduced.

CN118784599BActive Publication Date: 2025-09-26FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410551822.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-09-26
Estimated Expiration
2044-05-07

AI Technical Summary

Technical Problem

The existing Space-Ground Integrated Network (SAGIN) suffers from inefficiency and high complexity in resource management and computation offloading, making it difficult to effectively cope with changes in dynamic network environments and user traffic, resulting in poor service quality and resource utilization efficiency.

Method used

A traffic-aware lightweight hierarchical offloading framework for adaptive slicing is adopted. By dividing SAGIN into the Communication Access Platform (CAP) and the Computing Access Platform (COP), sparse probabilistic self-attention is used to predict user traffic, and adaptive network slicing and lightweight computing offloading algorithms are designed. Combined with DRL and policy distillation technology, resource allocation and computing offloading decisions are optimized.

Benefits of technology

It achieves efficient resource utilization and service quality in a dynamic SAGIN environment, reduces the latency and energy consumption of computational offloading, and improves the benefits and system performance of ESP.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118784599B_ABST
    Figure CN118784599B_ABST
Patent Text Reader

Abstract

The present invention provides a traffic-aware lightweight hierarchical offloading framework for adaptive slicing-enabled air-ground integrated networks. It divides SAGIN into CAP and COP, and uses network slicing to manage resources on each platform. ESP provides computing offloading services while allocating resources. For slice resource allocation, sparse probabilistic self-attention is used to capture dynamic traffic changes, and adaptive network slicing is performed based on predicted traffic and system load. For computing offloading, the communication process and the computing process are separated, sub-channels are allocated on demand according to channel conditions, and virtual machines are assigned to tasks using a lightweight computing offloading algorithm. The converged strategy is refined into a lightweight neural network for online reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of integrated air-ground-space networks, edge computing, computing offloading, etc., and specifically relates to a traffic-aware lightweight layered offloading framework for adaptive slicing-enabled integrated air-ground-space networks. Background Art

[0002] Emerging intelligent applications (such as autonomous driving and video analytics) are computationally intensive and latency-sensitive, yet the limited computing power of end devices severely restricts their further development and adoption. To alleviate this problem, Mobile Edge Computing (MEC) is considered a promising advanced computing paradigm. By deploying computing and storage resources at the edge of the network, MEC can significantly reduce network bandwidth pressure and data transmission latency. However, due to limited coverage and fixed network architecture, existing ground infrastructure, such as base stations (BSs) and roadside units (ROUs), cannot adequately meet the high quality of service requirements of intelligent applications. On the one hand, terrestrial networks cannot provide stable and continuous network access to users worldwide. Over 50% of the world's regions, particularly those with complex terrain such as oceans and isolated islands, still lack effective network coverage. On the other hand, as the core infrastructure of traditional MEC, ground base stations are vulnerable to natural disasters such as earthquakes and floods, leading to interruptions in network communication services. Recent advances in space and aerial communication technologies have led to a shift in the traditional MEC paradigm. Specifically, the aerial network composed of unmanned aerial vehicles (UAVs) and civilian aircraft can provide a temporary communication service for densely populated areas, with the advantages of flexible deployment and low access latency. The satellite network composed of low-Earth orbit (LEO) satellites can provide a communication service with global coverage and universal interconnection by integrating with the ground network. Therefore, through the complementary advantages of these three networks, MEC enabled by Space-Air-Ground Integrated Networks (SAGIN) is expected to provide seamless, full-time global access services for smart applications to better support many application areas that require real-time data sensing and complex computing.

[0003] However, due to the limited resources in SAGIN, when providing services for smart applications, unreasonable supply methods will seriously reduce resource utilization efficiency and service quality. By leveraging software-defined networking and virtualization technologies, infrastructure providers (InPs) can virtualize communication and computing resources into network slices and sell them to edge service providers (ESPs) based on resource pricing. ESPs can deploy different services to appropriate slices based on system status and user needs to provide a resource-customized service. When a user initiates a computation offloading request, the ESP can receive the user's task through SAGIN, execute the task using slice resources, and return the result. Although SAGIN has good characteristics and bright prospects, the following important challenges still exist in designing an efficient computation offloading framework for SAGIN.

[0004] (1) Existing SAGINs lack a comprehensive resource management model, making it difficult to effectively cope with dynamic network environments. To provide services across the entire time domain, ESPs need to simultaneously access and manage multiple communication and computing platforms. However, due to user mobility, user traffic in SAGINs changes over time, resulting in uneven distribution of system load in time and space. Therefore, while ensuring high Quality of Service (QoS) across different platforms, ESPs also need to consider load balancing to reduce resource costs.

[0005] (2) Due to the limited computing power of aerospace nodes such as satellites and drones, existing SAGINs cannot efficiently handle computationally intensive tasks from intelligent applications. Although SAGIN has powerful network access capabilities, satellites and drones in real-world scenarios are typically used to provide communication services and are equipped with limited storage and computing resources, making it difficult to handle complex tasks.

[0006] (3) The high complexity and computational overhead of computational offloading methods severely limit their adaptability and application prospects in SAGIN. Although some computational offloading methods have been proposed for SAGIN, they are difficult to deploy efficiently in SAGIN due to the limited computing resources and low-power architecture design of drones and satellites. As the scale of the network continues to increase, the latency and energy consumption introduced by running these methods are also unacceptable. Summary of the Invention

[0007] In order to overcome the problems existing in the prior art and solve the above challenges, the present invention comprehensively analyzes the advantages and disadvantages of the communication and computing platforms in SAGIN, and explores a new traffic-aware hierarchical computing offloading framework for SAGIN. Specifically, in the coverage area of ​​SAGIN, users can access the services provided by ESP without perception and upload their tasks to the available communication platform. Based on the analysis of user traffic distribution and platform computing power, ESP transfers tasks from the communication platform to the computing platform to perform computing offloading to achieve a balance between QoS and leasing costs. In order to achieve reasonable task offloading, DRL is introduced to interact with dynamic SAGIN, and decisions are made with the goal of maximizing ESP benefits. In particular, in response to the problem of limited computing power of drones, satellites, etc., policy distillation technology is introduced to reduce the scale of deep neural networks while extracting effective strategies in DRL, thereby reducing the delay and energy consumption required for model operation.

[0008] The technical solution specifically adopted by the present invention to solve the technical problem is:

[0009] A traffic-aware lightweight hierarchical offloading framework for adaptive slicing-enabled air-ground integrated networks divides SAGIN into CAP and COP, and uses network slicing to manage resources on each platform. ESP provides computational offloading services while allocating resources. For slice resource allocation, sparse probabilistic self-attention is used to capture dynamic traffic changes, and adaptive network slicing is performed based on predicted traffic and system load. For computational offloading, the communication process and the computation process are separated, sub-channels are allocated on demand based on channel conditions, and virtual machines are assigned to tasks using a lightweight computational offloading algorithm. The converged strategy is refined into a lightweight neural network for online inference.

[0010] Furthermore, at the beginning of the time slot, slice resources are adjusted based on historical traffic, load, and task completion. slice The adaptive network slicing algorithm is executed once per time slot to determine whether to adjust the slice; then, the user sends the task to be offloaded to the CAP closest to it. When both the base station and the drone are unavailable, the user sends the task to the satellite; after receiving the offloading request, the ESP allocates a subchannel for the task and uploads it to the CAP; then, the user task is transmitted from the CAP to the COP through the dedicated link in SAGIN; when the CAP is ground and satellite, the COP is the ground base station; when the CAP is a drone, the COP is the drone or the ground base station; after the task is transmitted to the COP, the lightweight computing offloading algorithm is called to assign the task to the corresponding virtual machine for execution, and the result is transmitted back to the user device after the calculation is completed; at the end of the time slot, the ESP collects task completion status, system load and user traffic information for profit calculation and slice adjustment.

[0011] Furthermore, the adaptive network slicing algorithm specifically includes the following steps:

[0012] Step 1: Predict future user traffic using a Transformer-based traffic prediction model. First, historical traffic is used to construct the encoder and decoder inputs. Next, the encoder output and decoder input are fed into the decoder, which consists of a multi-head sparse probabilistic self-attention layer and a multi-head attention layer. Finally, the encoder output is fed into an MLP, which, after inference, yields a future traffic forecast sequence.

[0013] Step 2: Calculate the expected system load. To accurately measure the slice load, define the system's communication load and computational load separately: the communication load in time slot t is expressed as the ratio of the used channel resources to the total channel resources, and the computational load is expressed as T que With T exe sum and After obtaining the system load, the expected resource requirements are calculated by multiplying the historical load by the ratio of the expected traffic to the historical traffic.

[0014] Step 3: Calculate the expected slice resources and determine whether to adjust the slice. If the difference between the benefit of adjusting the slice and the interruption cost is greater than 0, the slice adjustment is triggered and the communication and computing slice resources are adjusted to B respectively. * and F * , otherwise keep the slice resource unchanged.

[0015] Furthermore, the lightweight computation offloading algorithm is a computation offloading method targeting dynamic task traffic and number of virtual machines to improve resource utilization and ESP benefits in SAGIN, and the interaction between the DRL agent and the SAGIN environment is defined as a Markov decision process.

[0016] Furthermore, in the lightweight computing offloading algorithm:

[0017] First, initialize the network parameters and environment; for each received task, the state s t are all input into the participant network, and then the agent explores the unloading action in the current state according to π;

[0018] After receiving the action, the SAGIN environment executes the action and feeds back the status s of the next task t+1 , Instant Rewards t+1 and time slot state ω t , where ω t Indicates whether the task is the last task of the current time slot and is used to calculate the discounted reward; then, the samples of the state transition function are stored in the buffer;

[0019] When updating the network, in order to reduce the impact of noise caused by task attributes on gradient estimation, the generalized advantage estimate is introduced as the network update target;

[0020] After all rounds of training, the converged strategy executes the offloading decision by interacting with SAGIN;

[0021] To gain experience from the teacher model in SAGIN, the teacher model is used to interact with the environment and the state transition samples are stored in a buffer; next, state transition samples are randomly selected from the buffer for policy distillation.

[0022] Furthermore, the working process of the uninstallation framework includes the following steps:

[0023] (1) Collect historical user traffic of ESP, predict future user traffic, and calculate the resource requirements required by ESP. ESP will perform adaptive network slicing adjustments based on the resource requirements at regular intervals.

[0024] (2) The infrastructure provider responds to the ESP's resource request, divides the network slice for it, and charges the corresponding cost;

[0025] (3) The user accesses and uploads the computation offloading request to ESP through SAGIN;

[0026] (4) Allocate CAP and COP to users based on their locations, and generate subchannel and VM allocation strategies based on task attributes and user priorities;

[0027] (5) During the computation offloading process, THOAS records the user traffic, status, actions taken, rewards obtained, and new status entered in each time slot, and continuously optimizes its own performance based on this information.

[0028] Compared with the existing technology, the present invention and its preferred solution first divide SAGIN into CAP and COP, and use network slicing to manage resources on each platform. Then, a traffic prediction method is designed to capture dynamic traffic changes using probabilistic sparse self-attention, and an adaptive network slicing method is developed based on this. Finally, a lightweight improved DRL offloading method is designed to reduce network complexity while maintaining good performance. In further verification experiments, compared with other existing methods, the solution of the present invention made better slice adjustment and offloading decisions, and showed higher performance in ESP profit, task completion time, RU and DVR. The model complexity can be greatly reduced while retaining the original performance, further proving its practicality in the resource-constrained SAGIN environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0030] Figure 1 This is a diagram of the layered offloading framework for adaptive slicing enabled by SAGIN proposed in an embodiment of the present invention;

[0031] Figure 2 Overview of the traffic-aware lightweight hierarchical offloading framework THOAS designed for an embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the performance of THOAS in traffic prediction and adaptive slicing according to an embodiment of the present invention;

[0033] Figure 4 Schematic diagram of convergence performance in different methods and distillation processes according to embodiments of the present invention;

[0034] Figure 5 Schematic diagram of the performance of THOAS under different distillation network sizes according to an embodiment of the present invention;

[0035] Figure 6 Schematic diagram of the impact of traffic prediction on ESP returns, costs, and benefits according to an embodiment of the present invention;

[0036] Figure 7 This is a schematic diagram comparing task completion times using different methods according to an embodiment of the present invention;

[0037] Figure 8 This is a schematic diagram showing the impact of user traffic on RUs using different methods according to an embodiment of the present invention;

[0038] Figure 9 This is a schematic diagram showing the impact of the maximum tolerable delay of a task on DVR using different methods according to an embodiment of the present invention;

[0039] Figure 10 This is a schematic diagram showing the impact of slice expansion ratio on ESP revenue under different methods according to an embodiment of the present invention;

[0040] Figure 11 This is a schematic diagram of the impact of communication delay ratio on ESP revenue under different methods according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] To make the features and advantages of this patent more clearly understood, the following embodiments are specifically described in detail as follows:

[0042] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which this application belongs.

[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0044] like Figure 1 As shown in FIG, the model example of the layered offloading framework for adaptive slicing enabling SAGIN proposed in the embodiment of the present invention consists of a satellite, a ground base station and several drones. The satellite, base station and drone are all equipped with wireless access channels that can provide network access functions to users, which are called Communication Access Platforms (CAPs). The CAP set is denoted as A = {a1, a2, ..., a R Based on Orthogonal Frequency Division Multiplexing (OFDM) technology, the CAP channel can be divided into several orthogonal sub-channels (SC). j The total number of sub-channels ∈A is recorded as Base stations and drones are equipped with computing units that can provide computing resources for tasks from intelligent applications. They are called Computation Offloading Platforms (COPs). The set of COPs is denoted as O = {o1,o2,...,o S The computing resources of COP are provided in the form of virtual machines (VMs). j ∈O The total number of virtual machines is recorded as

[0045] InP maintains the wireless channels and virtual machines in SAGIN and provides them to ESP in the form of network slices. ESP pays fees to InP to apply for network slices, deploys services to each network slice after obtaining slice resources, and charges service fees by meeting users' computing offload requests. In order to meet service needs in more scenarios and save resource costs, ESP needs to deploy services to multiple network slices in SAGIN, configure appropriate resources for each slice, and continuously monitor and dynamically adjust them. j The number of slice sub-channels configured at is recorded as B j , in o j The number of slice virtual machines configured at is recorded as F j .

[0046] When a user generates a computation offloading request, the user accesses SAGIN through the nearest available CAP and uploads the task to be offloaded. Then, the task is transferred from the CAP to the COP for computation and returned to the user through the CAP after the computation is completed. If the task is completed within the user's maximum tolerable delay, the user will pay the corresponding fee to the ESP. Let all users served by the ESP be denoted as the set U = {u1,u2,...,u N}, where different CAPs are located in different geographical locations and have different communication coverage. j The number of users within the coverage area is recorded as N j , the number of users served by ESP is the sum of all users covered by CAP, that is:

[0047]

[0048] Due to user mobility, user traffic in different areas of SAGIN changes over time, resulting in uneven distribution of user traffic in time and space, and in turn, unbalanced CAP and COP loads. To address this problem, ESP is required to monitor and analyze user traffic and system load in different regions, thereby dynamically adjusting network slices to improve service adaptability and resource utilization efficiency. Specifically, a time slot t∈{1,2,...,T} is defined. At the beginning of time slot t, ESP allocates slice resources to users to meet their offloading requests. At the end of time slot t, ESP collects user access traffic and system load information for each slice. Based on the analysis of traffic and load, ESP can predict future user needs and adjust slices in a timely manner according to the package configuration provided by InP.

[0049] communication model

[0050] From user u i The task ∈U is defined as a six-tuple, denoted as <d i ,η i ,ρ i ,a i ,l i ,o i >, where d i is the amount of data for this task, η i To complete the computational density of this task, ρ i for u i The priority of a i Indicates u i Connected to the CAP, l i for u i and a i The distance between i Indicates execution of ui The COP of the task. Priority reflects the user's service level. The higher the priority, the higher the reward for completing the task.

[0051] Compared to drones and satellites, base stations have more stable communication links and more cost-effective channel costs. For areas outside the coverage of base stations, drones can provide a more flexible expansion of communication and computing capabilities. However, for users in some remote areas (such as the sea and deserts), satellites may be the only available communication method, and users can only access the network via satellite. Therefore, when a user needs to initiate a computing offload request, users within the coverage area of ​​base stations and drones will preferentially access SAGIN through base stations and drones, while other users will access SAGIN via satellite.

[0052] When u i When an uninstall request is initiated, u i Input data needs to be uploaded. Consider the following different situations.

[0053] (1) If the user is within the coverage of the base station or drone, he / she will upload the task through the base station or drone. The signal-to-noise ratio of the upload process is expressed as

[0054]

[0055] Among them, p u is the upload power, σ 2 is the noise power, P loss =10βlog(l i )+C+X G is the average path loss, β is the path loss exponent, C is a constant that depends on the operating frequency and antenna gain, and X G is a Gaussian random variable.

[0056] (2) If the user is not within the coverage of the base station and the drone, the user will upload the task via satellite. The signal-to-noise ratio during data transmission is defined as

[0057]

[0058] Among them G u and G s are the antenna gains of the user and satellite respectively, λ is the wavelength, F rain is the rain fall attenuation, which obeys the Weibull distribution.

[0059] When ESP is assigned to u i The number of sub-channels is According to Shannon's theorem, u i The upload rate is

[0060] ri =b i H log2(1+SNR), (4)

[0061] Where H is the bandwidth of a subchannel. Therefore, u i The time required to upload the task to CAP is

[0062]

[0063] Although satellites can also provide computing services, their energy consumption and resource costs are very expensive compared to ground base stations. Therefore, in real-world scenarios, it is not appropriate to use satellites as a computing node because their costs usually exceed the benefits they generate. Considering that the advantage of satellites is that they can simultaneously connect to users in remote areas and ground base stations with abundant resources, user tasks from remote areas can be transmitted to ground base stations via satellite-to-ground links for offloading, which not only achieves long-distance communication but also saves computing costs. In addition, drones can provide a flexible computing service by carrying small computing units, but due to limitations in computing power and battery energy storage, they may not be able to meet the needs of all tasks. When encountering this situation, it is necessary to consider whether the tasks received by the drone should be appropriately forwarded to the ground base station for execution.

[0064] Therefore, when the user task is transmitted to SAGIN, it is necessary to assign a suitable COP to it according to the CAP accessed by the user. When the user accesses through a ground base station, the task can be executed directly on the base station. When the user accesses a satellite, the task needs to be forwarded to the ground base station through the satellite-ground link for execution. When the user accesses a drone, it will be decided whether to forward it to the ground base station for execution based on the task requirements and network conditions. If forwarding is required, the task will be forwarded to the ground base station for execution through the drone-satellite and satellite-ground links. Otherwise, the task is executed on the drone. Accordingly, u i The forwarding time of a task between different platforms is defined as

[0065]

[0066] Among them, R s2g Indicates the communication rate between the satellite and the base station, R a2s Indicates the communication rate between the drone and the satellite.

[0067] Computational model

[0068] When a task is transferred to a suitable COP, the COP will assign the task to a virtual machine to perform the calculation. The virtual machine may perform multiple tasks. i The queuing time required for a task from arriving at the virtual machine to starting execution is

[0069]

[0070] Among them, Q represents the task queue that exists when a task arrives at the virtual machine, f edge Indicates the computing power of the virtual machine in COP. In addition, u i The actual execution time of the task in the virtual machine is

[0071]

[0072] Finally, the execution results are returned to the user through CAP. Compared to the input data, the output data volume is usually small, so the time to return the results is negligible.

[0073] Return and cost model

[0074] Considering the communication and computing models, processing u i The total time for the task is:

[0075]

[0076] On the one hand, ESP will charge users a certain fee based on the services it provides. max If completed within time slot t, ESP can obtain reward Φ. Otherwise, there is no reward. In time slot t, ESP obtains i The return is defined as

[0077]

[0078] The rewards obtained by completing tasks of users with different priorities are different. Therefore, the total reward of ESP in time slot t is defined as

[0079]

[0080] On the other hand, ESP needs to pay a certain fee to rent the sub-channel and VM resources in SAGIN, which is proportional to the amount of rented resources. Therefore, in time slot t, the total cost required for ESP to rent channel and VM resources is

[0081]

[0082] Among them, b and ζ f Represent the unit price of renting sub-channel and virtual machine resources respectively.

[0083] Based on the model proposed above, the optimization goal is to maximize the long-term benefits of ESP. The optimization problem is formally defined as

[0084]

[0085] Here, B and F represent the communication and computational slicing strategies, respectively, and π represents the computational offloading strategy. Constraints C1 and C2 indicate that the number of subchannels and virtual machines requesting a slice cannot exceed the maximum number of subchannels and virtual machines for the access platform, respectively. It should be noted that when user traffic fluctuates, the offloading strategy may not be able to meet the offloading needs of all users. In this case, the slices need to be adjusted to increase the slice resource capacity, which will also affect the offloading strategy. Therefore, the decision-making between network slicing and computational offloading is mutually coupled. The designed strategy needs to achieve reasonable computational offloading while adapting to changes in user traffic and slice resources in the system.

[0086] Overview of the proposed THOAS technical solution

[0087] In order to solve the optimization problem and maximize the ESP benefit, this paper proposes THOAS, a traffic-aware lightweight hierarchical offloading framework to support SAGIN with adaptive slicing. Figure 2 As shown in Figure 1, ESP provides computational offloading services while performing resource allocation. For slice resource allocation, a new traffic prediction method is designed to analyze and predict future traffic fluctuations. Adaptive network slicing is then performed based on the predicted traffic and system load. For computational offloading, communication and computation are separated, with subchannels first allocated on demand based on channel conditions. Next, an improved DRL is designed to efficiently allocate virtual machines to tasks. The converged strategy is refined into a lightweight neural network for online inference and better adaptability to resource-constrained SAGINs.

[0088] The overview of THOAS is shown in Algorithm 1. At the beginning of a time slot, Algorithm 2 is called to adjust the slice resources based on historical traffic, load, and task completion (lines 2-3). Since frequent adjustments will cause excessive system overhead, consider adjusting the slice resources every T sliceAlgorithm 2 is executed once per time slot to determine whether to adjust the slice. The user then sends the task to be offloaded to the CAP closest to it (e.g., a ground base station or an aerial drone). If neither is available, the user sends the task to the satellite (line 5). After receiving the offload request, the ESP allocates a subchannel for the task and uploads it to the CAP (line 6). The user task is then transferred from the CAP to the COP via a dedicated link in SAGIN (line 7). When the CAP is ground or satellite, the COP is the ground base station. When the CAP is a drone, the COP is either the drone or the ground base station, depending on the drone's load. After the task is transferred to the COP, Algorithm 3 is called to assign the task to the appropriate virtual machine for execution and, after completion, transmits the results back to the user device (lines 8-9). At the end of the time slot, the ESP collects information such as task completion status, system load, and user traffic for subsequent revenue calculation and slice adjustment (line 10).

[0089] In order to better manage the load of different access platforms, the maximum tolerable delay of the task is split into the maximum communication tolerance delay Maximum computational tolerance delay Let ω represent the ratio between the maximum communication tolerance delay and the maximum tolerance delay, that is, Furthermore, an on-demand channel allocation strategy is defined

[0090]

[0091] Among them, SNR depends on u i The CAP to which it is connected.

[0092]

[0093] Adaptive network slicing method

[0094] By analyzing load status and predicting resource requirements, the performance of ESP system resource management can be significantly improved. Typically, system load is affected by user traffic and service demand. Existing advanced prediction models can predict future user request patterns by capturing historical traffic change characteristics, while user service demand depends on the type of service provided by the ESP, which can be determined by analyzing historical load. Therefore, future ESP load conditions are derived by combining user traffic prediction and service demand analysis. When the expected load significantly exceeds or falls below the current ESP slice capacity, slice resources are adjusted to keep the load within a reasonable range. In addition, it is worth noting that frequent slice adjustments may cause service interruptions and affect QoS, so the ESP needs to consider these factors comprehensively to determine the appropriate time for slice adjustment.

[0095] In SAGIN, the traffic pattern of user offload requests includes long-term demand changes (such as the impact of application popularity, etc.) and short-term load fluctuations (such as user mobility, etc.). Among them, long-term demand changes have a more significant impact on SAGIN resource utilization efficiency. Compared with prediction models based on RNN and CNN, Transformer demonstrates outstanding ability in capturing long-term memory dependencies, and it can better learn the global patterns and local trends of user access to services. In addition, the slice windows in the proposed adaptive network slice are dynamic, so the input traffic sequences are usually of unequal length. Thanks to the self-attention mechanism, Transformer can process input data of different time scales without adjusting the model structure, and has better practicality in predicting traffic in different slice windows. Based on user traffic prediction and task demand analysis, an adaptive network slicing method is proposed. Its main steps are shown in Algorithm 2.

[0096] Step 1: Predict future user traffic. A traffic prediction model based on Transformer is designed. First, the input of the encoder and decoder is constructed using historical traffic (line 1). his Indicates historical access traffic, X cur represents the traffic sequence collected in the current slice window, X 0 Represents the time sampling of the traffic sequence to be predicted. When constructing the Encoder, the classic self-attention mechanism needs to calculate the attention weights of all historical time slots, which leads to high computational complexity. To alleviate this problem, sparse probabilistic self-attention is used in the Encoder and self-attention distillation is used between layers to reduce computational overhead (line 2). Specifically, the feature extraction process from layer j to layer j+1 is defined as

[0097]

[0098]

[0099] in,[·] attention represents sparse self-attention, d is Conv1d represents a one-dimensional convolution on the time series, ELU is the activation function, and MaxPool is the maximum pooling operation. Next, the encoder output and decoder input are fed into the decoder, which consists of a multi-head sparse probabilistic self-attention mechanism and a multi-head attention mechanism (line 3). Finally, the encoder output is fed into the MLP, which, after inference, yields a future traffic forecast sequence (line 4).

[0100] Step 2: Calculate the expected system load. Since the system load is positively correlated with user traffic, the load fluctuation can be deduced from the traffic change. In order to accurately measure the slice load, the communication load and computational load of the system are defined separately (line 5). Specifically, the communication load in time slot t is expressed as the ratio of the used channel resources to the total channel resources, and the computational load is expressed as T que With T exe sum and After obtaining the system load, the expected resource requirements can be calculated by multiplying the historical load by the ratio of expected traffic to historical traffic.

[0101] Step 3: Calculate the expected slice resources and determine whether to adjust the slice. Take 1+δ times the load peak as the expected slice resources, where δ represents the proportion of additional resources purchased by ESP (line 6). On the one hand, adjusting the slice may affect the ESP return and bring additional system overhead. On the other hand, the process of adjusting the virtual machine may cause service unavailability, resulting in interruption costs. Therefore, ESP needs to consider these factors comprehensively to decide whether to adjust the slice. The benefits of adjusting the slice can be defined as

[0102] ΔP=ΔR-ΔC, (17)

[0103] Where ΔR represents the difference between the expected ESP return after adjusting the slice and the expected ESP return of keeping the current slice, and ΔC represents the difference between the ESP cost after adjusting the slice and the ESP cost of keeping the current slice (row 7). The interruption cost caused by adjusting the slice can be defined as

[0104]

[0105] Among them, T int Indicates the number of time slots required to adjust the slice (line 8). If the difference between the benefit generated by adjusting the slice and the interruption cost is greater than 0, the slice adjustment is triggered and the communication and computing slice resources are adjusted to B respectively. * and F * , otherwise keep the slice resource unchanged (lines 9-10).

[0106]

[0107] Lightweight computing offloading method

[0108] Based on adaptive network slicing, a computation offloading method was further designed to address dynamic task traffic and the number of virtual machines (VMs) to improve resource utilization and ESP benefits in SAGINs. In recent years, DRL has been widely used to solve complex problems such as computation offloading and task scheduling. Although deep network structures increase DRL's fitting capabilities, they also introduce increased computational costs, limiting its application in resource-constrained SAGINs. In practical decision-making, when faced with a limited search space, the policy generated by a trained DRL agent is primarily influenced by a subset of network parameters. Therefore, compressing the original deep network to obtain a lightweight decision-making model helps improve DRL's training and inference efficiency while maintaining its superior decision-making capabilities. Knowledge distillation, an emerging learning paradigm that transfers knowledge from a deep teacher network to a shallow student network, has been widely used for model compression. For actor-critic DRL, by interacting the trained actor network with the environment, the teacher model's action probability distribution is collected and used as the distillation target. Subsequently, the student model's parameters are trained using a loss function and optimizer to align its output with the teacher model. Based on this idea, a novel lightweight computation offloading method combining DRL with knowledge distillation is proposed. Specifically, the interaction between the DRL agent and the SAGIN environment is defined as a Markov decision process, where the state space, action space, and reward function are defined as follows.

[0109] State space: It contains the VM task queues of the ESP slice, task attributes, and the user priority of sending tasks. In order to better capture the state characteristics, the maximum computational tolerable delay and the CPU cycles required for the task are converted into computational frequency, which can represent the computational resources required for the task. Therefore, the state is defined as

[0110]

[0111] Action Space: When the UAV's VM cannot meet the offloading requirements of all tasks, the tasks can be forwarded to the BS for offloading. Therefore, in order to make the algorithm applicable to UAVs and ground base stations, action spaces that adapt to different scenarios are defined. When the task is executed on the BS, the action space includes the target VM. When the task is executed on the UAV, the action space includes the target VM and forwarding to the BS for execution. Using VM G It means that the task is forwarded to the BS via the SAGIN wireless link, and the task will be reallocated to the virtual machine at the BS. Therefore, the action space is defined as

[0112] a t ∈{VM G ,VM1,VM2,...,VM max}.(20)

[0113] Reward function: Based on the P1 optimization objective, the reward function is defined as the reward obtained by completing the task. When the task is forwarded from the UAV to the BS, the reward is calculated at the BS. A small positive value is used as the reward to distinguish the forwarded task from the failed task. Therefore, r t is defined as

[0114]

[0115] Based on the above definitions, Algorithm 3 outlines the key steps of the proposed computation offloading method. First, the network parameters and environment are initialized (Line 2). For each received task, the state s t are input into the participant network, and then the agent explores the unloading action in the current state according to π (line 4).

[0116] After receiving the action, the SAGIN environment will execute the action and feedback the status of the next task s t+1 , Instant Rewards t+1 and time slot state ω t , where ω t Indicates whether the task is the last task of the current time slot and is used to calculate the discounted reward (line 5). Next, the samples of the state transition function are stored in the buffer.

[0117] When updating the network, in order to reduce the impact of noise caused by task attributes on gradient estimation, the generalized advantage estimate is introduced as the network update target, which is defined as

[0118]

[0119] δ t =R t +γV(s t+1 )-V(s t ), (twenty three)

[0120] where γ is the reward discount rate, λ is the advantage function discount rate, and δ t is the time differential error, R t is the reward, and V is the state value function (lines 6-7). Since the action of each task may affect the queue time and execution time of subsequent tasks, it is necessary to use the rewards of subsequent tasks in the current time slot to calculate the discounted reward, R t Defined as

[0121]

[0122] However, due to the dynamic changes in task traffic and the number of virtual machines, the step size of the policy update in the policy gradient is difficult to determine. To solve this problem, a clipping mechanism is used to ensure that each update is within a certain range (line 8). Therefore, the objective function of the policy update is defined as

[0123]

[0124] Where r(θ) is the ratio of the sample weight under the new strategy to the old strategy, which is defined as

[0125]

[0126] In order to increase the stability of the policy update, the clipping function is first used to avoid the policy update range being too large. At the same time, in order to prevent the fixed confidence interval from causing the update to be too slow, a two-layer dynamic confidence interval is designed, which is defined as follows

[0127]

[0128] where ∈ is the clip ratio, α t is a dynamic confidence factor, which is dynamically adjusted according to the timing differential error, α t Defined as

[0129]

[0130] κ is used to control the update speed of the dynamic confidence factor. Then, by minimizing L critic (φ) (line 9) to optimize the critic network, which is expressed as

[0131] L critic (φ)=E(R t +γV(S t+1 )-V(S t )) 2 . (29)

[0132] After all rounds of training, the converged strategy can be used to execute the unloading decision by interacting with SAGIN. Based on the structure of the neural network, the complexity of the above method is Where P, n p are the number of layers and the number of neural units in each layer. Therefore, in order to better adapt to the limited computing units in SAGIN, it is necessary to reduce P and n p To reduce the algorithm complexity.

[0133] To gain experience from the teacher model (i.e., the converged model) in SAGIN, the teacher model is used to interact with the environment and the state transition samples are stored in a buffer (line 12). Next, state transition samples are randomly selected from the buffer for policy distillation (line 13).

[0134] For each sample, the teacher network’s policy is used as the learning target, and the Adam optimizer is used to train the student network to output a similar action probability distribution as the teacher network (line 14). This process can be regarded as supervised learning. The loss function is constructed using soft labels and KL divergence, which is defined as

[0135]

[0136] in, and are the action probability distributions output by the teacher and student models, respectively, and τ is the temperature. Compared to directly using Q-values ​​as the distillation target, applying a softmax to the action probabilities can reduce the variance of the loss function and facilitate convergence of the student network. Furthermore, in the proposed Markov model, two actions with similar action probabilities may have significantly different rewards. Therefore, sharpening the action probability distribution using a low-temperature softmax allows high-reward actions from the teacher model to be more effectively transferred to the student model.

[0137] Policy distillation effectively reduces the depth and width of the original neural network model. Compared to agents directly using small-scale networks, large networks are more efficient at exploring high-reward actions and fitting complex policies during training. The distillation process also converges quickly, making it worthwhile to balance system performance and overhead by distilling large networks rather than using small ones.

[0138]

[0139]

[0140] According to the above framework design, the workflow can be obtained as follows:

[0141] (1) THOAS collects historical user traffic of ESP, predicts future user traffic, and calculates the resource requirements required by ESP. ESP then performs adaptive network slicing adjustments based on the resource requirements at regular intervals.

[0142] (2) The infrastructure provider responds to the ESP's resource request, divides the network slice for it, and charges the corresponding cost.

[0143] (3) The user accesses through SAGIN and uploads the computation offloading request to ESP.

[0144] (4) THOAS allocates appropriate CAP and COP to users based on their locations, and generates sub-channel and virtual machine allocation strategies based on task attributes (such as task size, computing requirements, maximum tolerable delay, etc.) and user priorities.

[0145] (5) During the computation offloading process, THOAS records the user traffic, status, actions taken, rewards obtained, and new status entered in each time slot, and continuously optimizes its own performance based on the above information.

[0146] Method evaluation

[0147] The proposed system and THOAS were implemented using PyTorch on a workstation equipped with an 8-core Intel(R) Xeon(R) Silver 4208 CPU @ 3.2GHz, an NVIDIA GeForce RTX3090 GPU, and 32GB of RAM. Dynamic user requests were constructed in SAGIN using a real-world dataset of Milan cellular traffic, which includes three types of services: messaging, calling, and internet. User request traffic was recorded over two months at a 10-minute sampling frequency. Specifically, three regions were selected, and the internet service traffic recorded at each sampling time was considered the number of user requests within a time slot. The main parameter settings for the experiment are shown in Table 1, where the different parameters of the satellite, drone, and base station are represented by a list of triples.

[0148] Table 1. Parameter settings

[0149]

[0150] In addition to ESP benefit and task completion time, the following performance metrics are used to further evaluate THOAS.

[0151] (1) Resource Utilization (RU): The ratio between the SC and VM used to execute tasks and the ESP leased resources.

[0152] (2) Deadline Violation Rate (DVR): The ratio of the number of tasks whose delay exceeds the maximum tolerable delay to the total number of tasks.

[0153] THOAS is compared with the following baseline methods to verify its superiority.

[0154] (1) GL-TCN: Temporal convolutional network for flow and adaptive slicing, which uses dilated convolution and residual connections to capture long-term and short-term dependencies in flow sequences;

[0155] (2) PredRNN: Utilizes long short-term memory for traffic and adaptive slicing, which combines recurrent neural networks and gating mechanisms to capture long-term and short-term dependencies in traffic sequences;

[0156] (3) Static: The adaptive network slicing method is not used (i.e., ESP resources remain unchanged), and the offloading algorithm is consistent with THOAS.

[0157] (4) PPO-TO: Proximal Policy Optimization-Penalty Strategy is used to make offloading decisions, which introduces a penalty function to improve the stability of policy updates.

[0158] (5) DDQN-TS: Dual DQN is used for offloading decisions, and the target network is introduced to solve the problem of overestimation of Q value in DQN.

[0159] (6) DQNM: Deep Q-network is used to make offloading decisions, which uses a deep neural network to approximate the Q-network and selects the action with the largest Q value as the candidate action.

[0160] First, the ability of THOAS to capture traffic changes and adaptively adjust resources (including SCs and VMs) was evaluated. Figure 3 As shown in the figure, when user traffic shows a periodic increase, THOAS can accurately predict this growth trend. Accordingly, THOAS will adaptively adjust the slice resources (i.e., increase the number of sub-channels and virtual machines leased by ESP) so that the leased resources can meet the needs of users. When user traffic shows a decreasing trend, THOAS can also adaptively adjust the slice resources again at the appropriate time to reduce resource waste and the cost of ESP leasing resources. It is worth noting that THOAS did not respond to small traffic fluctuations by adjusting slice resources. This is because THOAS performs traffic prediction and slice resource adjustment at regular intervals, which can effectively reduce the overhead of system operation and avoid service interruptions caused by frequent slice adjustments.

[0161] Next, the convergence of different methods and distillation processes was evaluated. Figure 4 As shown, DDQN-TS achieves higher rewards and exhibits better convergence than DQNM. This is because the Q network in DQNM is used for both action evaluation and selection, leading to Q-value overestimation and reducing the correlation between samples. DDQN-TS alleviates this problem by using a target network. Compared to DQNM and DDQN-TS, PPO-TO converges to higher rewards during training because its advantage value function improves the accuracy of action-value estimation. Furthermore, PPO-TO's proximal policy clipping limits the policy update step size, partially alleviating the impact of environmental dynamics such as traffic fluctuations and resource changes on policy updates. However, this fixed interval restriction also reduces the efficiency of policy updates, resulting in inefficient sample utilization. To address this, THOAS employs a dynamic confidence interval clipping strategy, allowing the policy update range to adjust based on the sample advantage value, significantly improving sample utilization and enabling THOAS to converge to better performance.

[0162] Next, the effect of using different network sizes during the distillation process on the performance of THOAS was evaluated. Figure 5 As shown in the figure, when the student model uses the same network size as the teacher model, the performance of the distilled student model is only slightly reduced. When the network size of the student model continues to decrease to 6% of the teacher model, the distilled student model can still maintain 73% of the original performance. This is because when faced with the discrete action space in SAGIN, the policy generated by the trained DRL model is mainly affected by only some network parameters. The proposed policy distillation utilizes the action probability distribution output by the teacher model and passes its core network parameters to a shallow network, which not only maintains the accuracy of the policy generated by the teacher model but also reduces the redundancy of the network structure. Therefore, by utilizing the distillation technology, the proposed THOAS can significantly reduce the network size of the intelligent agent while maintaining superior performance, making it better suitable for SAGIN environments with different levels of computing power.

[0163] Then, the impact of different traffic prediction methods on ESP return, cost and benefit is analyzed, where ESP return is composed of cost and benefit. Figure 6 As shown, when adaptive slicing based on traffic prediction is adopted, the cost is slightly higher than the static method, but the return and benefits are significantly improved. This is because when user traffic increases, the static slicing resources cannot meet user needs. This causes the completion time of some tasks to exceed their maximum tolerable delay, and the ESP is unable to obtain the corresponding return. Similarly, when user traffic decreases, the proposed adaptive network slicing can also reduce resources, avoiding unnecessary resource costs. Compared with GL-TCN and PredRNN, the proposed THOAS can achieve higher traffic prediction accuracy and thus obtain higher returns and benefits. This is because the proposed traffic prediction model based on sparse probabilistic self-attention can better capture the periodic traffic fluctuations in SAGIN, and thus outperforms other methods in prediction performance. The results show that the proposed adaptive slicing based on traffic prediction can effectively improve the benefits of ESP.

[0164] Next, the task completion time of different methods is compared, which is composed of task upload time, transmission time, queuing time and execution time. Since the channel allocation strategy is fixed, the upload time of the four methods is the same. Figure 7As shown in the figure, DQNM incurs longer task queueing times compared to the other three methods. This is because DQNM cannot schedule tasks to appropriate virtual machines, resulting in excessively long task queues on some virtual machines. Compared to PPO-TO and DDQN-TS, THOAS has longer transmission times but shorter execution and queuing times. This is because when traffic in a SAGIN increases to a certain level, user demand within the drone's coverage area may exceed its computing capacity. To address this issue, THOAS forwards tasks to a ground base station via wireless links within the SAGIN for execution. This base station has a higher computational frequency, thus reducing both queueing and execution times. As can be seen from the figure, under this hierarchical offloading framework, the transmission time incurred by task forwarding is less than the saved queueing and execution time, resulting in a shorter total task completion time for THOAS than for the other methods. Furthermore, the network size of the PPO-TO method is larger than that of the DQNM and DDQN-TS methods. This is because PPO-TO uses an actor-critic structure to select and evaluate offloaded actions, which is more complex than the Q network used in DQNM and DDQN-TS. Compared with other methods, THOAS uses distillation technology to effectively compress the network size, which significantly reduces the complexity and training overhead of the model and is more suitable for SAGIN with resource-constrained devices (such as drones).

[0165] Then, the impact of user traffic changes on RU under different methods is compared. Figure 8 As shown in the figure, the RU significantly decreases when the traffic volume increases from 1.0x to 0.5x. This is because as traffic decreases, the chances of ESP-leased resources being utilized decrease. Even in certain time slots where traffic approaches zero, the ESP still needs to maintain basic resources to ensure service availability, resulting in a RU of zero. It is worth noting that sparse traffic can lead to increased prediction errors, resulting in lower resource utilization in the 0.5x traffic scenario compared to the 1.0x traffic scenario. However, when traffic volume increases from 1.0x to 1.5x, the RU increases slightly. This is because as user traffic increases, the probability of additional ESP-leased resources being utilized increases, leaving fewer resources idle, resulting in higher RU. Compared to GL-TCN and PredRNN, THOAS achieves higher traffic prediction accuracy and higher RU, demonstrating that better prediction model performance helps improve resource utilization in SAGIN.

[0166] Then, the impact of the maximum tolerated delay of tasks on DVR under different methods is compared. Figure 9As shown in the figure, as the maximum tolerable delay of a task increases, the DVR decreases significantly. This is because there is more time to complete the task, and therefore the proportion of violations of the task's maximum delay decreases. As the maximum tolerable delay of a task continues to increase, the DVR may approach 0, indicating that the resources leased by ESP can well meet user needs. Compared with PPO-TO and DDQN-TS, THOAS can complete more tasks within the same maximum tolerable delay, thereby achieving a lower DVR. This is because THOAS uses generalized advantage estimation and proximal policy pruning based on dynamic confidence intervals to improve decision-making performance, adaptively making more reasonable decisions in different scenarios and completing more tasks.

[0167] Next, the impact of the slice expansion ratio δ on ESP returns is analyzed. Figure 10 As shown in the figure, as δ increases, the EPS benefits of different methods first increase and then decrease. This is because when δ is low, the additional rented resources may not be able to meet the needs of all tasks, which makes some tasks unable to be completed within their maximum tolerable delay, resulting in lower ESP benefits. As δ increases, the additional rented resources will increase when slicing is adjusted, which can alleviate the above problem and thus improve ESP benefits. When δ continues to increase, ESP benefits begin to decrease. This is because the cost of additional rented resources exceeds the return they bring, resulting in a decrease in ESP benefits. The above results show that based on the expected resources calculated by THOAS, appropriate renting of additional resources can increase ESP benefits while maintaining costs. Under different δ, the performance of THOAS is better than other methods, which shows the superior performance of the proposed accurate traffic prediction and adaptive network slicing in improving ESP benefits.

[0168] Next, the impact of the communication delay ratio ω on ESP benefits is analyzed. Figure 11As shown, as ω increases, the ESP benefits of different methods initially increase and then decrease. This is because when ω is low, the task upload time constraint is more stringent. In this case, ESP needs to allocate more subchannels to tasks, which in turn increases the cost of ESP resource leasing. As ω increases, more task upload time is available, thus alleviating this problem. However, as ω continues to increase, task queuing and execution delays decrease. In this case, the execution frequency of the virtual machines in SAGIN may not meet task requirements, causing some tasks to exceed their maximum tolerable delay, thereby reducing ESP benefits. These results demonstrate that THOAS, by splitting the maximum tolerable delay into communication delay and computation delay, improves its ability to perceive system requirements and, in turn, improves ESP benefits. Compared to PPO-TO, DDQN-TS, and DQNM, the proposed THOAS achieves the highest ESP benefits across all communication scenarios, demonstrating its superiority in handling offloading problems in SAGIN.

[0169] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

[0170] This patent is not limited to the above-mentioned optimal implementation method. Anyone can derive various other forms of traffic-aware lightweight layered offloading frameworks for adaptive slicing-enabled air-space-ground integrated networks under the inspiration of this patent. All equal changes and modifications made within the scope of the patent application of this invention should be covered by this patent.

Claims

1. A traffic-aware lightweight layered offloading framework for adaptive slicing-enabled air-ground integrated networks, characterized by: SAGIN is divided into CAP and COP, and network slicing is used to manage resources on each platform. ESP provides computing offload services while allocating resources. For slice resource allocation, sparse probabilistic self-attention is used to capture dynamic traffic changes and perform adaptive network slicing based on predicted traffic and system load. For computational offloading, the communication and computation processes are separated, sub-channels are allocated on demand based on channel conditions, and virtual machines are assigned to tasks using a lightweight computational offloading algorithm. The converged strategy is refined into a lightweight neural network for online inference. At the beginning of the time slot, slice resources are adjusted based on historical traffic, load, and task completion. The adaptive network slicing algorithm is executed once per time slot to determine whether to adjust the slice. Then, the user sends the task to be offloaded to the CAP closest to it. When both the base station and the drone are unavailable, the user sends the task to the satellite. After receiving the offloading request, the ESP allocates a subchannel for the task and uploads it to the CAP. Then, the user task is transmitted from the CAP to the COP via a dedicated link in SAGIN. When the CAP is ground or satellite, the COP is the ground base station; when the CAP is a drone, the COP is the drone or ground base station. After the task is transmitted to the COP, the lightweight computing offloading algorithm is called to assign the task to the corresponding virtual machine for execution, and the result is transmitted back to the user device after the calculation is completed. At the end of the time slot, the ESP collects task completion status, system load and user traffic information for profit calculation and slice adjustment. The adaptive network slicing algorithm specifically includes the following steps: Step 1: Predict future user traffic through the Transformer-based traffic prediction model; First, use historical traffic to construct the input of the encoder and decoder; Indicates historical access traffic. Indicates the traffic sequence collected by the current slice window, Represents the time sampling of the traffic sequence that needs to be predicted; sparse probabilistic self-attention is used in the Encoder and self-attention distillation is used between layers to reduce computational overhead. Layer to the first The feature extraction process of the layer is defined as: in, represents sparse self-attention, It is Layer encoder input Conv1d represents a one-dimensional convolution on the time series, ELU is the activation function, and MaxPool is the maximum pooling operation. Next, the encoder output and decoder input are fed into the decoder, which consists of a multi-head sparse probabilistic self-attention mechanism and a multi-head attention mechanism. Finally, the encoder output is fed into the MLP, and after inference, the future traffic prediction sequence is obtained. Step 2: Calculate the expected system load. To accurately measure the slice load, define the system's communication load and computing load separately: The communication load under the time slot is expressed as the ratio of the used channel resources to the total channel resources, and the calculation load is expressed as and sum and After obtaining the system load, the expected resource requirements are calculated by multiplying the historical load by the ratio of the expected traffic to the historical traffic. Step 3: Calculate the expected slice resources and determine whether to adjust the slice: times as the desired slice resource, It represents the proportion of additional resources purchased by ESP; the benefit from adjusting the slice is defined as in, It represents the difference between the expected ESP return after adjusting the slice and the expected ESP return of the current slice. It represents the difference between the ESP cost after adjusting the slice and the ESP cost of maintaining the current slice. The interruption cost caused by adjusting the slice is defined as in, Indicates the number of time slots required to adjust the slice; if the difference between the benefit generated by adjusting the slice and the interruption cost is greater than 0, the slice adjustment is triggered and the communication and computing slice resources are adjusted to and , otherwise keep the slice resource unchanged.

2. The traffic-aware lightweight layered offloading framework for adaptive slicing-enabled air-ground integrated networks according to claim 1, characterized in that: The proposed lightweight computation offloading algorithm is a computation offloading method targeting dynamic task traffic and number of virtual machines to improve resource utilization and ESP benefits in SAGIN, and the interaction between the DRL agent and the SAGIN environment is defined as a Markov decision process.

3. The traffic-aware lightweight layered offloading framework for adaptive slicing-enabled air-ground integrated networks according to claim 2, characterized in that: In the Markov decision process, the state space, action space and reward function are defined as follows; State space: Contains the VM task queues of the ESP slice, task attributes, and the user priority of the task being sent; converts the maximum computational tolerable delay and the CPU cycles required for the task into computational frequency to represent the computational resources required for the task; Action space: When the UAV virtual machine cannot meet the offloading requirements of all tasks, the task is forwarded to the BS for offloading; define action spaces that adapt to different scenarios: when the task is executed on the BS, the action space includes the target virtual machine; when the task is executed on the UAV, the action space includes the target virtual machine and forwarding to the BS for execution; use It indicates that the task is forwarded to the BS via the SAGIN wireless link, and the task is reallocated to the virtual machine at the BS; Reward function: Based on the optimization goal, the reward function is defined as the benefits obtained from completing the task; When a task is forwarded from the UAV to the BS, the reward is calculated at the BS; a positive value is used as the reward to distinguish between forwarded tasks and failed tasks.

4. The traffic-aware lightweight layered offloading framework for adaptive slicing-enabled air-ground integrated networks according to claim 2, characterized in that: In the lightweight computation offloading algorithm: First, initialize the network parameters and environment; for each received task, the state are input into the participant network, and then the agent Explore the uninstallation action in the current state; After receiving the action, the SAGIN environment executes the action and reports the status of the next task. , instant rewards and slot status ,in Indicates whether the task is the last task of the current time slot and is used to calculate the discounted reward; then, the samples of the state transition function are stored in the buffer; When updating the network, in order to reduce the impact of noise caused by task attributes on gradient estimation, the generalized advantage estimate is introduced as the network update target, which is defined as in is the reward discount rate, is the advantage function discount rate, is the time differential error, It's a reward. is the state value function; the discounted reward is calculated using the reward of the subsequent task in the current time slot, Defined as A clipping mechanism is used to ensure that each update is within a given range. The objective function of the policy update is defined as in is the ratio of the sample weights under the new strategy to the old strategy, which is defined as A two-level dynamic confidence interval is used, defined as follows in is the editing ratio, is a dynamic confidence factor, which is dynamically adjusted according to the timing differential error. Defined as Used to control the update speed of the dynamic confidence factor; then, by minimizing Optimize the critic network, which is expressed as After all rounds of training, the converged strategy executes the offloading decision by interacting with SAGIN; To gain experience from the teacher model in SAGIN, the teacher model is used to interact with the environment and store state transition samples in a buffer. Next, state transition samples are randomly selected from the buffer for policy distillation. For each sample, the teacher network’s policy is used as the learning target, and the Adam optimizer is used to train the student network to output a similar action probability distribution as the teacher network. This process is considered supervised learning; the loss function is constructed using soft labels and KL divergence, which is defined as in, and are the action probability distributions output by the teacher model and the student model, is the temperature; the action probability is softmaxed as the distillation target.

5. The traffic-aware lightweight layered offloading framework for adaptive slicing-enabled air-ground integrated networks according to claim 1, characterized in that: The working process of the uninstallation framework includes the following steps: (1) Collect ESP historical user traffic, predict future user traffic, and calculate the resource requirements required by ESP. ESP performs adaptive network slicing adjustments based on the resource requirements at regular intervals. (2) The infrastructure provider responds to the ESP's resource request, divides the network slice for it, and charges the corresponding cost; (3) The user accesses and uploads the computation offloading request to ESP through SAGIN; (4) Allocate CAP and COP to users based on their locations, and generate subchannel and VM allocation strategies based on task attributes and user priorities; (5) During the computation offloading process, THOAS records the user traffic, status, actions taken, rewards obtained, and new status entered in each time slot, and continuously optimizes its own performance based on this information.

Citation Information

Patent Citations

  • Slice-based collaborative task unloading method in air-space-ground integrated Internet of Vehicles

    CN116193396A

  • Task unloading method in space-air-ground network based on multi-target depth Q network

    CN116431240A