Slicing and collaborative offloading method for air-ground integrated network
By using Transformer to predict user traffic in an integrated air-space-ground network and combining it with improved DRL for collaborative offloading and resource allocation, the adaptability problem of static slicing schemes is solved, achieving efficient resource utilization and maximizing benefits.
Patent Information
- Application Number
- CN202411345059.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-09-25
AI Technical Summary
In existing integrated air-space-ground networks, static network slicing schemes are difficult to adapt to dynamically changing user traffic, leading to insufficient or excessive resources, affecting QoS and ESP benefits. At the same time, frequent slicing resource allocation operations increase system overhead, and existing DRL-based methods suffer from overestimation of Q-values or excessive variance, resulting in unstable agent convergence or getting trapped in local optima.
We employ a Transformer-based slicing resource partitioning method to predict future user traffic, and combine it with an improved DRL for collaborative offloading and resource allocation. By optimizing resource partitioning and collaborative offloading strategies at the beginning of each slice window, we maximize the cumulative benefits of ESP.
It improves system resource utilization efficiency and task completion rate, reduces the system overhead of frequent resource slicing, and enhances the benefits of ESP and system stability.
Smart Images

Figure CN119071851B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile edge computing technology, specifically to a method for slicing and collaborative offloading in an integrated air-space-ground network. Background Technology
[0002] With the rapid development of artificial intelligence and communications, more and more intelligent applications (such as autonomous driving and virtual reality) are integrating into people's daily lives. These applications typically require significant computing and storage resources to ensure Quality of Service (QoS). However, limited by the computing architecture and energy reserves of mobile devices, the high resource demands of these intelligent applications are difficult to meet. To address this issue, Mobile Edge Computing (MEC) is considered a promising new computing paradigm. In MEC, MEC servers can process tasks offloaded from user mobile devices and return the results to the user upon completion, thus better supporting the complex computing needs of intelligent applications. Compared to traditional cloud computing, MEC significantly reduces the response latency of task requests by deploying computing and storage resources closer to the user at the network edge, thereby effectively improving QoS.
[0003] However, the dispersed location distribution and diverse service needs of mobile users place higher demands on the flexibility of MEC network architecture. Fixed terrestrial access networks have limited service coverage, making it difficult to meet the dynamic and ever-changing user service requirements. As an emerging network architecture, Space-Air-Ground Integrated Networks (SAGIN) builds upon terrestrial networks, collaborating with airborne and spaceborne networks to provide mobile users with a more comprehensive and flexible network access. In SAGIN, terrestrial networks (e.g., base stations) serve nearby users; while airborne and spaceborne networks, typically composed of satellites and drones, can provide services to users in remote areas and alleviate server load pressure in densely populated areas to some extent. However, the heterogeneous node collaboration and diverse resource allocation in SAGIN also increase the difficulty of resource management. Based on network function virtualization and software-defined networking technologies, network slicing technology aims to divide the physical network into multiple logically independent networks according to user needs, representing a potentially feasible solution for efficient heterogeneous resource management. By introducing network slicing, a SAGIN-oriented network architecture can be built. Edge Service Providers (ESPs) can then orchestrate and configure slice resources and deploy different services to appropriate slices to meet the service needs of different users.
[0004] In real-world SAGIN scenarios, user traffic is complex and highly variable, with request volumes and resource demands typically changing dynamically over time. Static network slicing schemes struggle to adapt to this dynamic environment, leading to either insufficient or excessive slice resources, severely impacting QoS and ESP benefits. However, frequent slice allocation operations also incur additional system overhead. By rationally planning the resources and timing of slice allocation, the overhead from frequent allocation operations can be reduced while efficiently utilizing resources. Therefore, designing an efficient collaborative offloading and resource allocation mechanism for dynamically changing network slices is a problem worthy of further in-depth research. Existing research typically employs iterative algorithms or control theory, which are ill-suited to the dynamic SAGIN environment requiring efficient collaboration between heterogeneous platforms. In the field of machine learning, the emerging Deep Reinforcement Learning (DRL) has shown initial application potential in resource management-related problems, providing new perspectives and methods for solving complex issues. By continuously interacting with the dynamic environment, DRL agents are expected to gradually optimize collaborative offloading and resource allocation decisions, thereby improving system resource utilization efficiency and ESP benefits. Although DRL is a promising solution, existing DRL-based works often suffer from overestimation of Q-value or excessive variance, which makes the agent prone to convergence instability or getting trapped in local optima during training. Summary of the Invention
[0005] The purpose of this invention is to provide a method for slicing and collaborative offloading in an integrated air-space-ground network, which is beneficial to improving system resource utilization efficiency and task completion rate.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a method for slicing and collaborative offloading in an integrated air-space-ground network, comprising the following steps:
[0007] (1) Collect historical user traffic of ESP, predict future user traffic, and calculate the resource requirements required by ESP. ESP performs network slicing at the beginning of each slice window based on the resource requirements.
[0008] (2) The infrastructure provider responds to the ESP's resource allocation request and allocates network slices for it;
[0009] (3) Users access the SAGIN and upload a computational offload request to the ESP;
[0010] (4) ESP assigns appropriate edge servers to users based on their access location and request type, and allocates communication and computing resources based on task attributes and user priority;
[0011] (5) During the collaborative unloading process, record the actions taken, rewards obtained, and new states transitioned to in each time slot, and continuously optimize its own performance based on the above information.
[0012] Furthermore, the SAGIN consists of a ground base station (BS), drones, and satellites. Users access the compute offloading service provided by the ESP through the SAGIN. The entire area covered by the SAGIN is divided into three independent regions: the region covered by the BS, the region covered by the drones, and the region not covered by either the BS or the drones. The BS and drones provide services to users in their respective regions, while the satellite provides services to users in regions not covered by either the BS or the drones. The satellite communicates directly with the BS and drones, while the BS and drones communicate indirectly through satellite relay. All users are denoted as the set U = {u1, u2, ..., u...}. N The servers equipped on satellites, BS (Browser / Base Station), and drones are respectively defined as L. sat L uav and L bs Let all servers be denoted as set L = L sat ∪L uav ∪L bs ={L1,L2,...,L M Considering the differences in unloading requests across different regions, the time slot t server L will be used. j The number of uninstallation requests received within the coverage area is denoted as N. j Therefore, the total number of unload requests that ESP needs to handle is equal to the sum of the number of requests from each region, i.e.
[0013] The slicing model of SAGIN is as follows:
[0014] Orchestrate the physical resources of satellites, BS, and drones into resource-customized network slices; for server L j Its bandwidth and computing resources are respectively denoted as and The system time slots are divided into W slice windows, each slice window w containing several time slots t; the bandwidth and computing resources allocated to ESP in slice window w are respectively denoted as... and At the beginning of each slice window, ESP analyzes the historical requests received by each region to predict future demand and then allocates slice resources appropriately.
[0015] The communication model of SAGIN is as follows:
[0016] When user u iWhen initiating a compute offload request to ESP, the process mainly includes four stages: request access, platform collaboration, task execution, and result return; the task attributes of the UI are denoted as the quintuple {a i ,c i ,o i ,ρ i ,l i}, where each element represents the task's data volume, required computing power, type, priority, and u. i The server to which the task is connected; the type of task based on the amount of task data and the required computing power. i Divided into communication-intensive (CMI) and compute-intensive (CPI); task priority ρ i This indicates the level of reward you will receive for completing the task;
[0017] When ESP processes a compute offload request, it needs to consider u i The region where it is located; if u i Within the coverage area of a BS or drone, priority is given to accessing the computing offloading service via the BS or drone; otherwise, u i The computing offloading service can only be accessed via satellite; ESP is allocated to u i The bandwidth is denoted as b i , then u i Task access L j The uplink transmission rate is:
[0018]
[0019] Where, p i is u i The upload power, g0 and σ 2 They are u i With L j The channel gain and Gaussian white noise power between; u i Task request access L j The required time is:
[0020]
[0021] When the task is uploaded to L j Afterwards, ESP, based on u i The access location and task type are used to assign them to the corresponding server L. j' ; when u i When accessing the computing offloading service via BS, both CMI and CPI tasks are executed on the BS; when u i When accessing compute offloading services via drones or satellites, CMI tasks are executed directly on the drones or satellites, while CPI tasks are forwarded to BSs with more abundant computing resources for execution; correspondingly, u iTask data from L j Forward to L j' The required transmission time is:
[0022]
[0023] Among them, R u2s R represents the data transmission rate between the drone and the satellite. s2b This indicates the data transmission rate between the BS and the satellite;
[0024] The calculation model for SAGIN is as follows:
[0025] When a user task is offloaded to the appropriate server, ESP allocates computing resources to it; user u i The CPU computing power allocated to the task is denoted as f. i The time required to execute this task is:
[0026]
[0027] Once the task is completed, the user will receive feedback on the results.
[0028] Complete u i The total time for the task is:
[0029]
[0030] The profit model for SAGIN is as follows:
[0031] If the total time to complete the task is less than its maximum tolerable delay, then ESP receives a unit reward Φ; otherwise, ESP receives no reward. Therefore, at time t, ESP receives a reward from u. i The reward is:
[0032]
[0033] Meanwhile, considering the different priorities of different tasks, the rewards for completing higher priority tasks will be relatively higher; therefore, in time slot t, the total reward obtained by ESP for completing the task is:
[0034]
[0035] Furthermore, the computation offloading service provided by ESP incurs a certain cost, which is related to the actual amount of resources used by the task; therefore, in time slot t, the total resource rental cost of ESP is:
[0036]
[0037] in, and Let R represent the unit rental price of bandwidth and computing resources, respectively; therefore, the revenue obtained by ESP in time slot t is expressed as R. t -C t .
[0038] Furthermore, based on the slicing model, communication model, computation model, and revenue model of SAGIN, it is considered that slice resources are divided at the beginning of each slice window, and corresponding cooperative offloading strategies and communication and computation resource allocation decisions are executed in each time slot to maximize the cumulative total revenue of ESP; therefore, the optimization problem is defined as:
[0039]
[0040] Where, constraints C1 and C2 represent L j The communication and computing resources leased by the server within the slice window w shall not exceed L. j The total resources available; constraints C3 and C4 represent that the communication and computing resources allocated to users in each time slot do not exceed the total resources leased within that slice window; the optimization problem is decomposed into two sub-problems on corresponding time slots: slice partitioning and cooperative offloading.
[0041] P1: Determine a suitable slice resource allocation strategy to maximize the long-term benefit of ESP; define this subproblem as:
[0042]
[0043] By deciding on resource allocation within each slice window, P1 is solved to maximize the long-term benefit of ESP. Considering that in real-world edge environments, ESP may have difficulty accurately obtaining users' actual resource needs, historical user request traffic is analyzed and converted into ESP resource requirements.
[0044] P2: Perform cooperative unloading and resource allocation to maximize the short-term gains of ESP; define this subproblem as:
[0045]
[0046] The goal of P2 is to transfer different types of tasks to the appropriate server through collaborative offloading in each time slot t, and then maximize the benefits of the current time slot ESP through appropriate resource allocation. Since the number of tasks and requests to be processed in each time slot is different, ESP needs to adjust and optimize collaborative offloading and resource allocation decisions to meet the resource requirements of user tasks.
[0047] The proposed slice partitioning and collaborative offloading method integrates a Transformer-based slice resource partitioning method and an improved DRL-based collaborative offloading and resource allocation method. First, it uses Transformer to predict user traffic in future time slots, thereby guiding the ESP to partition slice resources. Then, by addressing the problems of Q-value overestimation and high variance, it optimizes collaborative offloading and resource allocation to maximize the benefits of the ESP.
[0048] Furthermore, the implementation method of the Transformer-based slice resource partitioning method is as follows:
[0049] First, the inputs to the encoder and decoder are constructed using historical user request traffic. Sparse probabilistic self-attention is employed when constructing the encoder, and self-attention distillation is used between layers to reduce computational overhead. The feature extraction process from layer j to layer j+1 is represented as follows:
[0050]
[0051]
[0052] in,[·] attention Denotes sparse self-attention, where d is X l t In the dimension, Conv1d represents one-dimensional convolution, ELU is the activation function, and MaxPool is the max pooling function;
[0053] Next, the encoder output is fed into the decoder, which consists of a multi-head sparse probabilistic self-attention mechanism and a multi-head attention mechanism. Then, the decoder output is fed into the MLP, where inference yields a predicted sequence of future traffic. Next, the maximum value of the predicted traffic in each time slot is taken as the predicted traffic for the next slice window, and the traffic is categorized according to task type and access platform. The communication resource requirement for each platform is the product of the unit communication resource requirement for all types of tasks within its coverage area and the corresponding type of request traffic, i.e. Where b CMI and b CPI These represent the average communication bandwidth required for CMI and CPI tasks, respectively. and These represent the predicted CMI and CPI task flows, respectively; the computing resources required for the CMI task within the BS coverage area are... The CPI target for other regions is Where f CMI and f CPI Let represent the average computing resource requirements for CMI and CPI tasks, respectively; satellites and drones forward CPI tasks within their coverage area to the BS for execution, therefore the computing resource requirements of the BS are... Drones and satellites only need to handle CMI tasks within their coverage area, therefore their computing resource requirements are respectively... and
[0054] Furthermore, after the resource allocation within the slice window is completed, the unloading process between different slice windows is relatively independent; therefore, the long-term optimization problem of maximizing the cumulative ESP reward in P2 is transformed into a short-term optimization problem within a single slice window, and this problem is constructed as a Markov decision process; SAGIN is regarded as the environment, and ESP is abstracted as a DRL agent, which makes cooperative unloading and resource allocation decisions by interacting with the environment, and then updates the policy through the reward signal fed back by the environment; the state space, action space and reward function are defined as follows;
[0055] (1) State Space: Within time slot t, by sensing available communication and computing resources and task attribute information, the agent captures the task requirements and the corresponding resources needed; therefore, the system state in time slot t is represented as:
[0056]
[0057] in,
[0058] (2) Action Space: The action space contains actions that allocate communication and computing resources to users; therefore, the actions in time slot t are represented as:
[0059] a t ={b t ,f t} (15)
[0060] in,
[0061] (3) Reward Function: Problem P2 aims to maximize the short-term return of ESP in each time slot; therefore, the reward function is defined as the ESP return in time slot t, which is expressed as:
[0062] r t =R t -C t (16)
[0063] The implementation method of the improved DRL-based collaborative offloading and resource allocation method is as follows:
[0064] First, initialize the main network Q1, Q2, and μ, and the target network Q'1, Q'2, and μ'. Introduce two Q-value networks with identical structures, and use the minimum value of the two to estimate the value of the next state-action pair when calculating the target value. Next, initialize the experience replay cache RB, the number of slice windows W, and the number of time slots T within each slice window. To reduce the correlation between experience data and improve overall training efficiency, an experience replay mechanism is introduced. At the beginning of the training cycle, initialize the system environment and obtain its initial state. At the beginning of each slice window, predict future time slot user traffic and perform slice resource allocation using a Transformer-based slice resource allocation method. In time slot t, select an action based on the policy network and exploration noise. Then, execute the action and obtain the immediate reward and next state from the environmental feedback. Subsequently, the training samples (s) generated during the interaction with the environment are... t ,a t ,r t ,s t+1 The data is stored in the experience replay cache RB; then, K samples are randomly selected from RB to update the network parameters; based on target policy smoothing regularization, the target action is obtained using the target policy network. This process can be represented as:
[0065]
[0066] The noise follows a normal distribution, and the sampled noise is cropped to reduce the variance of the target estimation; then, the target Q-value y... target The update is based on the minimum of the two Q-value networks, and the process is expressed as follows:
[0067]
[0068] Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the Q-value network parameters. After the Q-value network has been updated a certain number of times, the policy network parameters and the target network parameters are updated again using gradient ascent and soft update methods, respectively.
[0069] Compared with existing technologies, the present invention has the following advantages: The present invention provides a novel method for slice partitioning and collaborative offloading in an integrated air-space-ground network to solve the problems existing in the prior art; within the service range of SAGIN, users can access customized services provided by ESP and transmit data to available edge servers; based on the analysis of historical user traffic trends, ESP rationally partitions slice resources; in order to optimize collaborative offloading and resource allocation strategies, DRL is introduced to interact with dynamic SAGIN, making decisions with the goal of maximizing the long-term cumulative benefits of ESP, thereby improving system resource utilization efficiency and task completion rate. Attached Figure Description
[0070] Figure 1 This is a schematic diagram of an integrated air-space-ground network in an embodiment of the present invention;
[0071] Figure 2 This is a schematic diagram of the network slicing model in an embodiment of the present invention;
[0072] Figure 3 This is a schematic diagram of cooperative offloading and resource allocation based on improved DRL in an embodiment of the present invention;
[0073] Figure 4 This refers to the predictive performance of the method for user traffic in this embodiment of the invention;
[0074] Figure 5 This is a comparison of the convergence of the method in this embodiment of the invention with other methods;
[0075] Figure 6 This is a comparison of the ESP benefits and costs of different methods in the embodiments of the present invention;
[0076] Figure 7 This is a comparison of the task failure rates of different methods in the embodiments of the present invention;
[0077] Figure 8 This is a comparison of resource utilization rates of different methods in the embodiments of the present invention. Detailed Implementation
[0078] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0079] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0080] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0081] This embodiment provides a slice partitioning and collaborative offloading (SPCO) method for integrated air-space-ground networks, including the following steps:
[0082] (1) Collect historical user traffic of ESP, predict future user traffic, and calculate the resource requirements required by ESP. ESP performs network slicing at the beginning of each slice window based on the resource requirements.
[0083] (2) The infrastructure provider responds to the ESP's resource allocation request and allocates network slices for it;
[0084] (3) Users access the SAGIN and upload a computational offload request to the ESP;
[0085] (4) ESP assigns appropriate edge servers to users based on their access location and request type, and allocates communication and computing resources based on task attributes (such as task data size, computing requirements, maximum tolerable latency, etc.) and user priority.
[0086] (5) During the collaborative unloading process, record the actions taken, rewards obtained, and new states transitioned to in each time slot, and continuously optimize its own performance based on the above information.
[0087] 1. System Model and Problem Definition
[0088] like Figure 1 As shown, the SAGIN considered in this method consists of a ground base station (BS), drones, and satellites. Users access the compute offloading service provided by the ESP through the SAGIN. The entire area covered by the SAGIN is divided into three independent regions: the region covered by the BS, the region covered by the drones, and the region not covered by either the BS or the drones. The BS and drones provide services to users in their respective regions, while the satellite provides services to users in the region not covered by either the BS or the drones. The satellite communicates directly with the BS and drones, while the BS and drones communicate indirectly through satellite relay. Specifically, all users are denoted as the set U = {u1, u2, ..., u...}. N The servers equipped on satellites, BS (Browser / Base Station), and drones are respectively defined as L. sat L uav and L bs Let all servers be denoted as set L = L sat ∪L uav ∪L bs ={L1,L2,...,L M Considering the differences in unloading requests across different regions, the time slot t server L will be used. j The number of uninstallation requests received within the coverage area is denoted as N. j Therefore, the total number of unload requests that ESP needs to handle is equal to the sum of the number of requests from each region, i.e. Slice Model
[0089] like Figure 2 As shown, based on the characteristics of each SAGIN platform, the physical resources of satellites, BS, and UAVs are orchestrated into customized network slices. For server L... j Its bandwidth and computing resources are respectively denoted as and The system time slots are divided into W slice windows, and each slice window w contains several time slots t. The bandwidth and computing resources allocated to ESP in slice window w are denoted as follows: and Specifically, at the beginning of each slice window, ESP analyzes the historical requests received by each region to predict future demand and then allocates slice resources appropriately.
[0090] communication model
[0091] When user u i When initiating a compute offload request to ESP, the process mainly includes four stages: request access, platform collaboration, task execution, and result return. The task attributes of the UI are denoted as a quintuple {a}. i ,c i ,o i ,ρ i ,l i}, where each element represents the task's data volume, required computing power, type, priority, and u. i The server to which the task is connected. The task type depends on the amount of data and the required computing power. i Tasks are categorized into Communication-Intensive (CMI) and Computation-Intensive (CPI) tasks. Furthermore, the task priority ρ... i This indicates the level of reward you will receive for completing the task.
[0092] When ESP processes a compute offload request, it needs to consider u i The region in which u is located. i Within the coverage area of a BS or drone, priority is given to accessing the computing offloading service via the BS or drone; otherwise, u i The computational offloading service can only be accessed via satellite. ESP is allocated to u. i The bandwidth is denoted as b i , then u i Task access L j The uplink transmission rate is:
[0093]
[0094] Where, p i is u iThe upload power, g0 and σ 2 They are u i With L j The channel gain and Gaussian white noise power are related. Therefore, u i Task request access L j The required time is:
[0095]
[0096] When the task is uploaded to L j After that, ESP will... i The access location and task type are used to assign them to the corresponding server L. j' Compared to satellites and drones, BS provides communication and computing resources with higher cost-effectiveness and stability. Therefore, when u i When accessing compute offloading services via a BS, both CMI and CPI tasks will be executed on the BS. Compared to a BS, drones and satellites offer wider coverage and more flexible communication capabilities, but due to their limited resource capacity, they cannot meet all compute offloading service requirements. Therefore, when u i When accessing compute offloading services via drones or satellites, CMI tasks will be executed directly on the drone or satellite, while CPI tasks will be forwarded to a BS with more abundant computing resources. Accordingly, u i Task data from L j Forward to L j' The required transmission time is:
[0097]
[0098] Among them, R u2s R represents the data transmission rate between the drone and the satellite. s2b This indicates the data transmission rate between the BS and the satellite.
[0099] Computational model
[0100] When a user task is offloaded to the appropriate server, ESP allocates computing resources to it. User u i The CPU computing power allocated to the task is denoted as f. i The time required to execute this task is:
[0101]
[0102] After the task is completed, the user will receive feedback. Typically, the amount of data returned is small, and its time overhead is far less than other stages of the uninstallation process; therefore, the time overhead of returning the result can be ignored. Therefore, completing u... i The total time for the task is:
[0103]
[0104] Profit Model
[0105] If the total time to complete the task is less than its maximum tolerable delay, ESP will receive a unit reward Φ; otherwise, ESP will not receive a reward. Therefore, at time t, ESP from u i The reward is:
[0106]
[0107] Furthermore, considering the different priorities of various tasks, the rewards for completing higher-priority tasks will be relatively higher. Therefore, the total reward obtained by ESP for completing the task in time slot t is:
[0108]
[0109] Furthermore, ESP incurs a cost for providing computational offloading services, which is related to the actual amount of resources used by the task. Therefore, in time slot t, the total resource rental cost of ESP is:
[0110]
[0111] in, and Let R represent the unit rental price of bandwidth and computing resources, respectively. Therefore, the revenue obtained by ESP in time slot t can be expressed as R. t -C t .
[0112] Based on the above model, this invention considers dividing slice resources at the beginning of each slice window and executing corresponding cooperative offloading strategies and communication and computing resource allocation decisions in each time slot to maximize the cumulative total benefit of ESP. Therefore, the optimization problem can be formally defined as:
[0113]
[0114] Where, constraints C1 and C2 represent L j The communication and computing resources leased by the server within the slice window w shall not exceed L. j The total resources available; constraints C3 and C4 respectively indicate that the communication and computing resources allocated to users in each time slot do not exceed the total resources rented within that slice window. In the original optimization problem, there are decision variables belonging to different time scales, making it difficult to solve them uniformly to obtain the optimal solution. To better solve this problem, this invention decomposes it into two sub-problems on corresponding time slots: slice partitioning and cooperative offloading.
[0115] P1: Determine a suitable slice resource allocation strategy to maximize the long-term benefit of ESP. This subproblem is formally defined as:
[0116]
[0117] This invention solves for P1 by making decisions on resource allocation within each slice window, thereby maximizing the long-term benefits of ESP. Considering that in real-world edge environments, ESP struggles to accurately capture users' actual resource needs, this invention analyzes historical user request traffic and transforms it into ESP resource requirements.
[0118] P2: Perform cooperative unloading and resource allocation to maximize the short-term gains of ESP. This subproblem is formally defined as:
[0119]
[0120] The goal of P2 is to transfer different types of tasks to appropriate servers through collaborative offloading in each time slot t, thereby maximizing the benefits of ESP in the current time slot through appropriate resource allocation. Since the number of tasks and requests to be processed varies in each time slot, ESP needs to adjust and optimize collaborative offloading and resource allocation decisions to meet the resource requirements of user tasks.
[0121] The SPCO framework proposed in this invention
[0122] To address the aforementioned system model and problem definition, this invention proposes a novel SAGIN-oriented Slice Partitioning and Collaborative Offloading (SPCO) method to solve optimization problems P1 and P2. This method effectively integrates a Transformer-based slice resource partitioning method and an improved DRL-based collaborative offloading and resource allocation method. First, the Transformer is used to predict user traffic in future time slots, guiding the ESP in partitioning slice resources. Then, by addressing the issues of Q-value overestimation and high variance, the collaborative offloading and resource allocation are optimized to maximize the ESP's benefits.
[0123] Transformer-based slice resource partitioning
[0124] In SAGIN, user traffic typically changes dynamically over time. If resources are allocated only under a fixed slice allocation, slice resources will not be able to adapt well to traffic fluctuations. Based on historical user traffic data, the resource requirements of future time slots are predicted to rationally utilize slice resources to handle user computation offloading requests. Therefore, this invention solves optimization problem P1 by predicting future resource requirements for rational slice resource allocation. However, the actual resource requirements of user tasks cannot be accurately measured. To address this problem, this invention uses Transformer to effectively perceive future user traffic trends and combines them with historical resource usage to predict the demand for communication and computing resources in future time slots, thereby allocating slice resources accordingly. The key steps of the Transformer-based slice resource allocation method proposed in this invention are shown in Algorithm 1.
[0125]
[0126]
[0127] First, construct the inputs to the encoder and decoder using the user's historical request traffic (line 1), where X his Represents historical traffic, X cur Represents the flow sequence of the current slice window, X 0 This represents the time sampling of the traffic sequence to be predicted. Classical self-attention mechanisms require calculating attention weights for all historical time slots, leading to high computational complexity. To alleviate this problem, this invention employs sparse probabilistic self-attention in the encoder construction and utilizes self-attention distillation between layers to reduce computational overhead (line 2). Specifically, the feature extraction process from layer j to layer j+1 is represented as follows:
[0128]
[0129]
[0130] in,[·] attention Indicates sparse self-attention, d is The dimensions are: Conv1d represents one-dimensional convolution, ELU is the activation function, and MaxPool is the max pooling function.
[0131] Next, the encoder output is fed into the decoder (line 3), which consists of a multi-head sparse probabilistic self-attention mechanism and a multi-head attention mechanism. Then, the decoder output is fed into the MLP, and through inference, a predicted sequence of future traffic is obtained (line 4). Next, the maximum value of the predicted traffic in each time slot is taken as the predicted traffic for the next slice window, and then categorized according to task type and access platform (line 5). The communication resource requirement for each platform is the product of the unit communication resource requirement for all types of tasks within its coverage area and the corresponding type of request traffic, i.e. Where b CMI and b CPI These represent the average communication bandwidth required for CMI and CPI tasks, respectively. and These represent the predicted CMI and CPI task traffic, respectively. Since the computational load varies across different platforms, computational resources need to be configured accordingly. Specifically, the computational resources required for the CMI task within the BS coverage area are... The CPI target for other regions is Where f CMI and f CPI These represent the average computing resource requirements for CMI and CPI tasks, respectively. Satellites and drones will forward CPI tasks within their coverage area to the BS for execution; therefore, the computing resource requirements for the BS are... (Line 6). Drones and satellites only need to handle CMI tasks within their coverage area, therefore their computational resource requirements are respectively... and (Lines 7-8)
[0132] Collaborative offloading and resource allocation based on improved DRL
[0133] After the resource allocation within a slice window is completed, the unloading process between different slice windows is relatively independent. Therefore, the long-run optimization problem of maximizing the cumulative ESP reward in P2 can be transformed into a short-run optimization problem within a single slice window, and this problem can be constructed as a Markov decision process. Specifically, the SAGIN under consideration is regarded as the environment, and the ESP is abstracted as a DRL agent, which makes cooperative unloading and resource allocation decisions by interacting with the environment, and then updates its policy through reward signals fed back from the environment. Figure 3 The proposed collaborative offloading and resource allocation method based on improved DRL is illustrated. Accordingly, the state space, action space, and reward function are defined as follows.
[0134] (1) State Space: Within time slot t, by sensing available communication and computing resources and task attributes, the agent can capture the task requirements and the corresponding resources needed. Therefore, the system state in time slot t can be represented as:
[0135]
[0136] in,
[0137] (2) Action Space: The action space contains actions that allocate communication and computing resources to users. Therefore, the action in time slot t can be represented as:
[0138] a t ={b t ,f t} (15)
[0139] in,
[0140] (3) Reward Function: Problem P2 aims to maximize the short-term return of ESP in each time slot. Therefore, the reward function is defined as the ESP return in time slot t, which can be expressed as:
[0141] r t =R t -C t (16)
[0142]
[0143] The key steps of the proposed cooperative offloading and resource allocation method based on improved DRL are shown in Algorithm 2. First, the main networks Q1, Q2, and μ, and the target networks Q'1, Q'2, and μ' are initialized (lines 1-2). Classical DRL employs a strategy of maximizing Q-values, which leads to a non-uniform overestimation of Q-values. To address this, this method introduces two Q-value networks with identical structures, using the minimum value among them to estimate the value of the next state-action pair when calculating the target value. Next, the experience replay buffer RB, the number of slice windows W, and the number of time slots T within each slice window are initialized (line 3). To reduce the correlation between experience data and improve overall training efficiency, an experience replay mechanism is introduced. At the beginning of the training cycle, the system environment is initialized and its initial state is obtained (line 5). At the beginning of each slice window, Algorithm 1 is called to predict future time slot user traffic and perform slice resource allocation (line 7). In time slot t, actions are selected based on the policy network and exploration noise (line 9). Next, the action is performed and an immediate reward and next state are obtained from environmental feedback (line 10). Then, the training samples (s) generated during the interaction with the environment are... t ,a t ,r t ,s t+1Store the data in the experience replay cache RB (line 11). Next, randomly select K samples from RB to update the network parameters (line 12). Based on target policy smoothing regularization, obtain the target action using the target policy network. (Line 13), this process can be represented as:
[0144]
[0145] The noise follows a normal distribution, and the sampled noise is cropped to reduce the variance of the target estimation. Next, the target Q-value y... target The update is based on the minimum of the two Q-value networks (line 14), and this process can be represented as:
[0146]
[0147] Finally, the gradient descent algorithm is used to minimize the error between the evaluation and target values, thereby updating the Q-value network parameters (line 15). Furthermore, a delayed update method is employed to improve the training stability of the agent. After the Q-value network has been updated a certain number of times, the policy network parameters and target network parameters are updated again using gradient ascent and soft update methods respectively (lines 16-18).
[0148] Performance evaluation
[0149] Based on a workstation equipped with an 8-core CPU @ 3.2GHz and 32GB RAM, this embodiment constructs a SAGIN environment and implements the proposed SPCO framework using PyTorch. The Milan, Italy cellular network traffic dataset is used to simulate user traffic trends. This dataset records user communication traffic changes in the city over two months at a sampling frequency of 10 minutes, covering three types of real network traffic data: calls, SMS, and internet. In the experiment, the sampled internet communication traffic in the dataset is considered as the number of user requests per time slot; calls and SMS traffic are considered CMI-type tasks; and internet traffic is considered CPI-type tasks. One training epoch contains 24 slice windows, each containing 6 scheduling time slots. The constructed SAGIN environment includes one satellite, one UAV, and one BS. The main parameter settings are shown in Table 1, where the parameters of the satellite, UAV, and BS are represented using triples, and the parameters of the CMI and CPI tasks are represented using binary tuples.
[0150] To evaluate the superiority of the SPCO framework proposed in this invention, this embodiment comprehensively compares it with the following benchmark methods:
[0151] (1) TD3_RA: Resource allocation is performed using the Twin Delayed DeepDeterministic policy gradient algorithm (TD3), but adaptive slice partitioning is not considered.
[0152] (2) DDPG_RA: Resource allocation is performed using the Deep Deterministic Policy Gradient (DDPG) algorithm, but adaptive slice partitioning is not considered.
[0153] (3) HPFS: Tasks are executed according to the High Priority First Service (HPFS) policy, and communication and computing resources are allocated according to the average demand of the tasks.
[0154] (4) FCFS: Tasks are executed according to the First Come First Service (FCFS) policy, and the allocation of communication and computing resources is the same as HPFS.
[0155] Table 1 Main Parameter Settings
[0156]
[0157] First, the performance of this method for predicting user traffic was evaluated across multiple time slices. In the experiments, user traffic in the next time slice was predicted based on the user traffic from the first six time slices, and the actual and predicted user traffic values were compared on different servers. Figure 4 As shown, user request traffic exhibits a cyclical trend, and this method can effectively capture this trend and accurately predict future user traffic. Next, the user traffic prediction results will serve as the basis for resource allocation, and this method will further combine task attributes and environmental conditions to perform reasonable resource allocation to improve ESP (Effective Service Points) benefits.
[0158] Next, the convergence of different methods was analyzed. For example... Figure 5As shown, the FCFS and HPFS methods are limited by their single, agentless decision-making framework, which causes their reward curves to remain unchanged during training. Compared with other DRL-based methods, FCFS and HPFS perform poorly. This is because the strategies used by FCFS and HPFS in allocating resources to tasks are too rigid, lacking consideration for system dynamics and task-specific needs. As a result, the resource requirements of many tasks cannot be met, preventing ESP from obtaining the expected rewards for completing these tasks. The HPFS method prioritizes allocating resources to high-priority tasks, thus achieving higher rewards. The SPCO and TD3_RA methods demonstrate superior performance in training stability and convergence rate compared to the DDPG_RA method. This is because SPCO and TD3_RA methods use a dual-policy network to estimate the Q-value, which effectively alleviates the problem of overestimation of the Q-value, thereby improving the stability and performance of the algorithm. Meanwhile, both SPCO and TD3_RA methods employ delayed learning to update the policy network and introduce target policy smoothing. This avoids frequent network updates and effectively alleviates the high variance problem in the DDPG_RA method, enabling more stable convergence of model training. Compared to the TD3_RA method, the SPCO proposed in this invention achieves further performance improvements. This is because, compared to the TD3_RA method which uses static slice resource partitioning, the proposed SPCO can accurately predict future time-slot user request traffic and then dynamically partition slice resources, thereby better meeting user service resource requests.
[0159] Next, the benefits of ESP using different methods and the costs of resource rental were compared. For example... Figure 6 As shown, by classifying task types, the proposed SPCO can better identify the communication and computing resource requirements of tasks in SAGIN. Compared with other methods, SPCO can more rationally lease communication and computing resources while improving ESP benefits. This is because SPCO's strategy of adjusting slice resources based on predicted user traffic can effectively adapt to the resource needs of user tasks and better support edge nodes in providing services to users. The results show that the proposed traffic prediction-based slice partitioning strategy can more rationally and efficiently utilize resources to improve ESP benefits.
[0160] Next, the task completion performance of different methods under different maximum tolerable delays was compared. For example... Figure 7As shown, the task failure rate decreases with increasing maximum tolerable latency. This is because ESP has more time to process tasks, reducing the number of tasks that fail due to exceeding the maximum tolerable latency. Compared to DRL-based methods, FCFS and HPFS methods exhibit higher task failure rates under different maximum tolerable latency conditions. This is because FCFS and HPFS methods allocate resources only based on the average resource requirements of tasks, resulting in unreasonable resource allocation, thus some tasks cannot obtain sufficient resources. Compared to other methods, the proposed SPCO can complete more tasks and achieve the lowest task failure rate under different maximum tolerable latency conditions. This is because the dual Q-value network architecture and delayed update mechanism in SPCO improve decision-making performance. Therefore, SPCO can make more reasonable resource allocation decisions in different scenarios, thereby completing more tasks.
[0161] Finally, the utilization of slice resources by different methods under different user traffic multipliers was compared. For example... Figure 8 As shown, when user traffic is 0.5 times the initial setting, resource utilization decreases significantly. This is because as user requests decrease, the proportion of resources allocated to tasks also decreases, resulting in many resources being idle. When user traffic increases to 1.5 times the initial setting, resource utilization slightly improves. This is because the increase in user requests leads to more resources being allocated and used, reducing idle resources and thus exhibiting higher resource utilization. Compared to other methods, the proposed SPCO achieves higher resource utilization, indicating that adjusting the slicing strategy based on traffic prediction helps to utilize limited slice resources more rationally and efficiently.
[0162] This invention proposes a novel collaborative offloading method for slicing in integrated air-space-ground networks. First, a traffic prediction method is designed, which uses a probabilistic sparse self-attention mechanism to perceive the changing trends of user traffic, and based on this, a network slicing method is developed. Next, an improved DRL collaborative offloading and resource allocation method is designed, addressing problems such as Q-value overestimation and convergence difficulties caused by high variance, achieving reasonable allocation of communication and computing resources and efficient collaborative offloading across heterogeneous platforms. Experiments based on real traffic datasets verify the effectiveness of this method. Compared with other methods, this method effectively increases ESP (Effective Strategies for Network Expectations) and exhibits higher performance in terms of task completion rate and resource utilization.
[0163] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0167] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for slicing and collaborative offloading in an integrated air-space-ground network, characterized in that, Includes the following steps: (1) Collect historical user traffic of ESP, predict future user traffic, and calculate the resource requirements required by ESP. ESP performs network slicing at the beginning of each slice window based on the resource requirements. (2) The infrastructure provider responds to the ESP's resource allocation request and allocates network slices for it; (3) Users access the SAGIN and upload a computational offload request to the ESP; (4) ESP assigns appropriate edge servers to users based on their access location and request type, and allocates communication and computing resources based on task attributes and user priority; (5) During the collaborative unloading process, record the actions taken, rewards obtained, and new states transitioned to in each time slot, and continuously optimize its own performance based on the information recorded in each time slot. The slice partitioning and collaborative offloading method integrates a Transformer-based slice resource partitioning method and an improved DRL-based collaborative offloading and resource allocation method. First, the Transformer is used to predict user traffic in future time slots, thereby guiding the ESP to partition slice resources. Then, by solving the problems of Q-value overestimation and high variance, the collaborative offloading and resource allocation are optimized to maximize the benefits of the ESP. The implementation method of the Transformer-based slice resource partitioning method is as follows: First, the inputs to the encoder and decoder are constructed using historical user request traffic. Sparse probabilistic self-attention is employed when constructing the encoder, and self-attention distillation is used between layers to reduce computational overhead. The feature extraction process from layer j to layer j+1 is represented as follows: in,[·] attention Describing sparse self-attention, d is In the dimension, Conv1d represents one-dimensional convolution, ELU is the activation function, and MaxPool is the max pooling function; Next, the encoder output is fed into the decoder, which consists of a multi-head sparse probabilistic self-attention mechanism and a multi-head attention mechanism. Then, the decoder output is fed into the MLP, where inference yields a predicted sequence of future traffic. Next, the maximum value of the predicted traffic in each time slot is taken as the predicted traffic for the next slice window, and the traffic is categorized according to task type and access platform. The communication resource requirement for each platform is the product of the unit communication resource requirement for all types of tasks within its coverage area and the corresponding type of request traffic, i.e. Where b CMI and b CPI These represent the average communication bandwidth required for CMI and CPI tasks, respectively. and These represent the predicted CMI and CPI task flows, respectively; the computing resources required for the CMI task within the BS coverage area are... The CPI target for other regions is Where f CMI and f CPI Let represent the average computing resource requirements for CMI and CPI tasks, respectively; satellites and drones forward CPI tasks within their coverage area to the BS for execution, therefore the computing resource requirements of the BS are... Drones and satellites only need to handle CMI tasks within their coverage area, therefore their computing resource requirements are respectively... and 2. The method for slicing and collaborative offloading in an integrated air-space-ground network according to claim 1, characterized in that, The SAGIN consists of a ground base station (BS), drones, and satellites. Users access the compute offloading service provided by the ESP through the SAGIN. The entire area covered by the SAGIN is divided into three independent regions: the region covered by the BS, the region covered by the drones, and the region not covered by either the BS or the drones. The BS and drones provide services to users in their respective regions, while the satellite provides services to users in the region not covered by either the BS or the drones. The satellite communicates directly with the BS and drones, while the BS and drones communicate indirectly through satellite relay. All users are denoted as the set U = {u1, u2, ..., u...}. N The servers equipped on satellites, BS (Browser / Base Station), and drones are respectively defined as L. sat L uav and L bs Let all servers be denoted as set L = L sat ∪L uav ∪L bs ={L1,L2,...,L M Considering the differences in unloading requests across different regions, the time slot t server L will be used. j The number of uninstallation requests received within the coverage area is denoted as N. j Therefore, the total number of unload requests that ESP needs to handle is equal to the sum of the number of requests from each region, i.e. The slicing model of SAGIN is as follows: Orchestrate the physical resources of satellites, BS, and drones into resource-customized network slices; for server L j Its bandwidth and computing resources are respectively denoted as and The system time slots are divided into W slice windows, each slice window w containing several time slots t; the bandwidth and computing resources allocated to ESP in slice window w are respectively denoted as... and At the beginning of each slice window, ESP analyzes the historical requests received by each region to predict future demand and then allocates slice resources appropriately. The communication model of SAGIN is as follows: When user u i When initiating a compute offload request to ESP, the process mainly includes four stages: request access, platform collaboration, task execution, and result return; i The task attribute is denoted as a quintuple {a i ,c i ,o i ,ρ i ,l i }, where each element represents the task's data volume, required computing power, type, priority, and u. i The server to which the task is connected; the type of task based on the amount of task data and the required computing power. i Divided into communication-intensive (CMI) and computation-intensive (CPI); task priority ρ i This indicates the level of reward you will receive for completing the task; When ESP processes a compute offload request, it needs to consider u i The region where it is located; if u i Within the coverage area of a BS or drone, priority is given to accessing the computing offloading service via the BS or drone; otherwise, u i The computing offloading service can only be accessed via satellite; ESP is allocated to u i The bandwidth is denoted as b i , then u i Task access L j The uplink transmission rate is: Where, p i is u i The upload power, g0 and σ 2 They are u i With L j The channel gain and Gaussian white noise power are interdependent; will u i Task request access L j The required time is: When the task is uploaded to L j Afterwards, ESP, based on u i The access location and task type are used to assign them to the corresponding server L. j' ; when u i When accessing the computing offloading service via BS, both CMI and CPI tasks are executed on the BS; when u i When accessing compute offloading services via drones or satellites, CMI tasks are executed directly on the drones or satellites, while CPI tasks are forwarded to BSs with more abundant computing resources for execution; correspondingly, u i Task data from L j Forward to L j' The required transmission time is: Among them, R u2s R represents the data transmission rate between the drone and the satellite. s2b This indicates the data transmission rate between the BS and the satellite; The calculation model for SAGIN is as follows: When a user task is offloaded to the appropriate server, ESP allocates computing resources to it; user u i The CPU computing power allocated to the task is denoted as f. i The time required to execute this task is: Once the task is completed, the user will receive feedback on the results. Complete u i The total time for the task is: The profit model for SAGIN is as follows: If the total time to complete the task is less than its maximum tolerable delay, then ESP receives a unit reward Φ; otherwise, ESP receives no reward. Therefore, at time t, ESP receives a reward from u. i The reward is: Meanwhile, considering the different priorities of different tasks, the rewards for completing higher priority tasks will be relatively higher; therefore, in time slot t, the total reward obtained by ESP for completing the task is: Furthermore, the computation offloading service provided by ESP incurs a certain cost, which is related to the actual amount of resources used by the task; therefore, in time slot t, the total resource rental cost of ESP is: in, and Let R represent the unit rental price of bandwidth and computing resources, respectively; therefore, the revenue obtained by ESP in time slot t is expressed as R. t -C t .
3. The method for slicing and collaborative offloading in an integrated air-space-ground network according to claim 2, characterized in that, Based on the slicing model, communication model, computation model, and revenue model of SAGIN, this paper considers allocating slice resources at the beginning of each slice window and executing corresponding cooperative offloading strategies and communication and computation resource allocation decisions in each time slot to maximize the cumulative total revenue of ESP; therefore, the optimization problem is defined as: Where, constraints C1 and C2 represent L j The communication and computing resources leased by the server within the slice window w shall not exceed L. j The total resources available; constraints C3 and C4 represent that the communication and computing resources allocated to users in each time slot do not exceed the total resources leased within that slice window; the optimization problem is decomposed into two sub-problems on corresponding time slots: slice partitioning and cooperative offloading. P1: Determine a suitable slice resource allocation strategy to maximize the long-term benefit of ESP; define this subproblem as: By deciding on resource allocation within each slice window, P1 is solved to maximize the long-term benefit of ESP. Considering that in real-world edge environments, ESP may have difficulty accurately obtaining users' actual resource needs, historical user request traffic is analyzed and converted into ESP resource requirements. P2: Perform cooperative unloading and resource allocation to maximize the short-term gains of ESP; define this subproblem as: The goal of P2 is to transfer different types of tasks to the appropriate server through collaborative offloading in each time slot t, and then maximize the benefits of the current time slot ESP through appropriate resource allocation. Since the number of tasks and requests to be processed in each time slot is different, ESP needs to adjust and optimize collaborative offloading and resource allocation decisions to meet the resource requirements of user tasks.
4. The method for slicing and collaborative offloading in an integrated air-space-ground network according to claim 3, characterized in that, After the slice resources within the slice window are divided, the unloading process between different slice windows is relatively independent; Therefore, the long-run optimization problem of maximizing the cumulative ESP reward in P2 is transformed into a short-run optimization problem within a single slice window, and this problem is constructed as a Markov decision process; SAGIN is regarded as the environment, and ESP is abstracted as a DRL agent, which makes collaborative unloading and resource allocation decisions by interacting with the environment, and then updates the policy through the reward signal fed back by the environment; the state space, action space and reward function are defined as follows; (1) State Space: Within time slot t, by sensing available communication and computing resources and task attribute information, the agent captures the task requirements and the corresponding resources needed; therefore, the system state in time slot t is represented as: in, (2) Action Space: The action space contains actions that allocate communication and computing resources to users; therefore, the actions in time slot t are represented as: a t ={b t ,f t } (15) in, (3) Reward Function: Problem P2 aims to maximize the short-term return of ESP in each time slot; therefore, the reward function is defined as the ESP return in time slot t, which is expressed as: r t =R t -C t (16) The implementation method of the improved DRL-based collaborative offloading and resource allocation method is as follows: First, initialize the main network Q1, Q2, and μ, and the target network Q1', Q'2, and μ'. Introduce two Q-value networks with identical structures, and use the minimum value of the two to estimate the value of the next state-action pair when calculating the target value. Next, initialize the experience replay cache RB, the number of slice windows W, and the number of time slots T within each slice window. To reduce the correlation between experience data and improve overall training efficiency, an experience replay mechanism is introduced. At the beginning of the training cycle, initialize the system environment and obtain its initial state. At the beginning of each slice window, predict future time slot user traffic and perform slice resource allocation using a Transformer-based slice resource allocation method. In time slot t, select an action based on the policy network and exploration noise. Then, execute the action and obtain the immediate reward and next state from the environmental feedback. Subsequently, the training samples (s) generated during the interaction with the environment are... t ,a t ,r t ,s t+1 The data is stored in the experience replay cache RB; then, K samples are randomly selected from RB to update the network parameters; based on target policy smoothing regularization, the target action is obtained using the target policy network. This process can be represented as: The noise follows a normal distribution, and the sampled noise is cropped to reduce the variance of the target estimation; then, the target Q-value y... target The update is based on the minimum of the two Q-value networks, and the process is expressed as follows: Finally, the gradient descent algorithm is used to minimize the error between the evaluation value and the target value, thereby updating the Q-value network parameters. After the Q-value network has been updated a certain number of times, the policy network parameters and the target network parameters are updated again using gradient ascent and soft update methods, respectively.
Citation Information
Patent Citations
Deep learning-based intelligent identification method for color steel tile building along railway
CN114239755A
Computing unloading method for 5G network slice in MEC environment
CN117202264A