Space-air-ground integrated internet of vehicles based on slice cooperative task offloading method
By adopting the M/M/1 queuing model and DDQN method in the integrated air-space-ground vehicle network, the RAN slicing framework realizes adaptive slice window duration and heterogeneous base station cooperation, solves the problems of dynamic adjustment of slice window length and multi-dimensional resource orchestration, and improves task completion rate and system adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2026-03-31
AI Technical Summary
In the integrated air-space-ground vehicle network, existing technologies struggle to effectively handle dynamic adjustment of slice window length, multi-dimensional resource orchestration, and collaboration between heterogeneous base stations, resulting in low task offloading efficiency and difficulty in providing differentiated QoS guarantees, especially in high-speed mobile environments.
A RAN slicing framework based on the M/M/1 queuing model is proposed. Combined with the Dual Deep Q Learning (DDQN) method, it realizes adaptive slice window duration, spectrum and computing resource orchestration, and cooperative task offloading among heterogeneous base stations. Joint decision-making is carried out through the MEC controller to maximize the number of tasks completed.
It improves task completion rate, reduces task failure rate, enhances system adaptability and resource utilization, and is superior to traditional methods.
Smart Images

Figure CN116193396B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle networking technology, specifically a slice-based collaborative task offloading mechanism in Space-Air-Ground integrated Vehicular Networks (SAGVNs). Background Technology
[0002] The high bandwidth, millisecond latency and ultra-high density of 5G networks provide a prerequisite for the development of vehicle-to-everything (V2X) technology. V2X technology connects vehicles, base stations and service providers into an organic whole, enabling real-time acquisition of information from all parties [1]. Vehicle-mounted equipment has limited computing and storage capabilities, making it difficult to meet the needs of complex, data-intensive and latency-sensitive applications. One feasible solution is MEC-assisted task offloading. The tasks issued by the vehicle are offloaded to the MEC server on the base station (BS) or roadside unit (RSU), and the processed results are sent back to the vehicle to achieve low latency and high agility vehicle services. However, terrestrial RANs are difficult to cover all road networks and have problems such as rigid network structure and slow service response [2]. The high mobility of vehicles, the complexity of urban road networks and the differences in task requirements exacerbate the difficulty of task offloading and resource allocation.
[0003] Space-Air-Ground integrated Vehicular Networks (SAGVNs) can provide seamless information services for vehicle users and meet the service needs of all time and all domains [3]. This network is based on a ground-based network and supplemented and extended by space-based and air-based networks, providing comprehensive information protection for vehicles in a wide area of space [5]. The ground-based network consists of BS and RSU, providing services to areas with dense pedestrian and vehicle traffic; the air-based network consists of UAV base stations, which have advantages such as mobile deployment and line-of-sight (LoS) communication; the space-based network includes a space-based access network composed of low-orbit satellites, which is a necessary facility to achieve full coverage and ubiquitous connectivity. Both UAVs and low-orbit satellites can serve as airborne MEC platforms [5], providing network access and task offloading opportunities for vehicles in areas where the ground-based network coverage is at the edge or lacks infrastructure.
[0004] With the development of intelligent transportation and autonomous driving, new vehicle applications are emerging, which can be roughly divided into latency-sensitive (such as autonomous driving, path planning, collision warning) and latency-tolerant (such as map download, video distribution [6]). Network slicing [7] technology can divide the shared physical RAN into multiple isolated virtual networks (i.e., RAN slices) to provide customized services for different types of applications. RAN slicing is a key enabling technology for providing differentiated QoS guarantees for vehicle network task offloading. The controller configures computing and communication resources for RAN slices based on information such as task traffic. In sliced vehicle networks, the offloading strategy determines where to offload the task based on information such as task attributes, base station load, vehicle speed and direction. In the evolution of SAGVN, a natural step is to "extend" RAN slices from ground-based networks to air-based and space-based networks to support ubiquitous and diverse vehicle network applications.
[0005] SAGVN is a dynamic environment characterized by the convergence of multiple networks and high-speed vehicle movement, which presents many challenges for RAN slicing and task offloading:
[0006] (1) Dynamic adjustment of slice window length.
[0007] This is a fundamental issue and the key to balancing overhead and QoS. Due to the dynamic nature of the network and the time-varying nature of task traffic, the service provision capability of slices will gradually weaken over time. The MEC controller must periodically reallocate resources for RAN slices. If the slice window duration is shortened, resource reallocation will be triggered frequently, resulting in huge control and computation costs; if the slice window is extended, task traffic fluctuations may cause the slice performance isolation to be destroyed. Reference [9] studies a dynamic RAN slicing framework for vehicle networking, which divides time into multiple slice windows of equal length and calculates the optimal resource allocation strategy for each window. Reference
[11] proposes a hierarchical soft slicing framework that supports differentiated QoS, and performs resource slicing at the network layer and base station layer at large and small time scales, respectively. These schemes allocate resources under a fixed slice window duration.
[0008] (2) Multi-dimensional resource orchestration for multi-tier networks.
[0009] Traffic flow in the road network is unevenly distributed in the time and spatial domains. There are significant differences in the deployment, coverage and resources of different types of base stations. Resource allocation is affected by both service changes and vehicle movement. Resource coupling in heterogeneous networks exacerbates the complexity of decision-making. Most existing works consider terrestrial networks or single-type resource slicing. Ye et al. proposed a downlink spectrum resource slicing framework for heterogeneous wireless networks to provide differentiated QoS guarantees for machine-type devices and user equipment
[12] . Reference
[13] further incorporated a transmit power adjustment mechanism and designed a spectrum slicing strategy based on multi-access edge computing. A spectrum and computing resource slicing framework for RAN slicing was proposed in
[14] to meet the task offloading of differentiated QoS requirements in vehicle networks. Reference
[21] combined deep deterministic policy gradient (DDPG) and hierarchical learning to make decisions on the joint allocation of multi-dimensional resources in vehicular networks.
[0010] (3) Cooperation between heterogeneous base stations.
[0011] The interaction between vehicles moving at high speed and BS is very short. Vehicle speed, direction and road shape will affect the task offloading effect. The cooperation between air-to-ground, air-to-ground and air-to-air base stations is the key to reducing latency and improving task completion rate. Traditional model optimization and heuristic methods
[15]
[16]
[17] are difficult to handle the real-time task offloading problem in dynamic scenarios. Deep reinforcement learning (DRL)
[18] integrates the decision advantage of reinforcement learning RL and the perception advantage of deep learning DL, enabling individuals to perceive the environment and establish actions that match it, so that it can handle higher-dimensional state-action spaces. Most existing edge collaboration for vehicle networking considers the ground network environment. Kai et al. proposed an offloading scheme based on pipeline
[19] . Mobile devices and edge nodes can offload tasks to edge nodes or the cloud according to their own computing and communication capabilities. Li et al. proposed a task partitioning and scheduling algorithm for vehicle networking
[20] , which maintains service continuity by pre-selecting edge servers and reduces computing latency through edge-side collaboration. Summary of the Invention
[0012] To address the aforementioned problems in existing technologies, this invention proposes a slice-based collaborative task offloading method for Space-Air-Ground integrated Vehicular Networks (SAGVNs), which provides differentiated QoS guarantees for task offloading of high-speed vehicles while maximizing the number of tasks completed.
[0013] The method of this invention also proposes:
[0014] A service-oriented RAN slicing framework is proposed, supporting adaptive slice window duration, spectrum and computational resource orchestration, and collaboration among heterogeneous base stations. Based on the M / M / 1 queuing model, the joint decision-making process for Radio Access Network (RAN) slicing and task offloading is modeled as a problem of maximizing the number of long-term task completions. This problem is decoupled into three sub-problems: slice window duration allocation, resource allocation, and cooperative workflow scheduling.
[0015] The solution is solved alternately by a multi-access edge computing (MEC) controller, forming a closed loop with a slice window as the cycle. Whenever a new slice window arrives, the controller determines the window duration through a task traffic-aware strategy and allocates resources to the slice using an optimization method.
[0016] The workflow scheduling within a slice window at small timescales is determined by a method based on Double Deep Q-Learning Network (DDQN).
[0017] Simulation results show that, compared with the benchmark method, this method demonstrates superiority in terms of adaptability, task completion rate, and control overhead. Attached Figure Description
[0018] Figure 1 This represents a SAGVNs scenario;
[0019] Figure 2 Indicates the RAN slice frame;
[0020] Figures 3(a) and 3(b) illustrate collaborative workflow scheduling cases, where Figure 3(a) represents delay-sensitive task scheduling (Case-1) and Figure 3(b) represents delay-tolerant task scheduling (Case-2).
[0021] Figure 4 This represents the state machine of the MEC controller;
[0022] Figures 5(a) and 5(b) show the fitting of task traffic growth with the optimal slice window length, where Figure 5(a) shows the case where task traffic is continuously decreasing, and Figure 5(b) shows the case where task traffic is continuously increasing.
[0023] Figure 6 This represents a cooperative worker path scheduling decision based on DDQN;
[0024] Figure 7 Indicates the receipt of rewards;
[0025] Figures 8(a) to 8(c)The figures show the impact of increasing the number of training rounds on system performance, where: Figure 8(a) shows the system task completion reward, Figure 8(b) shows the number of tasks completed, and Figure 8(c) shows the task failure rate;
[0026] Figures 9(a) and 9(b) show a comparison between static and dynamic windows, where: Figure 9(a) represents the task failure rate, and Figure 9(b) represents the number of slice windows;
[0027] Figures 10(a) and 10(b) show the impact of different types of resource quantities on task failure rate, where: Figure 10(a) represents an increase in spectrum resources, and Figure 10(b) represents an increase in computing resources;
[0028] Figures 11(a) and 11(b) illustrate the impact of load changes on performance. Figure 11(a) shows the task failure rate and the number of tasks, while Figure 11(b) shows the task completion rate and the proportion of latency-sensitive tasks. Detailed Implementation
[0029] 1 Overview
[0030] To address the problems existing in the prior art, this invention proposes a slice-based cooperative task offloading method for SAGVNs, providing differentiated QoS guarantees for vehicle task offloading and maximizing the number of tasks completed. It mainly includes three technical contributions:
[0031] Design a service-oriented heterogeneous RAN slicing framework that supports dynamic allocation of slice window durations, orchestration of spectrum and computing resources, and cooperative task offloading. Based on the M / M / 1 queuing model, the joint decision-making for RAN slicing and task offloading is modeled as an optimization problem that maximizes the number of long-term tasks completed under coupling and resource constraints.
[0032] To balance QoS and signaling overhead, an adaptive strategy / method for RAN slice window duration is designed. During peak traffic periods, the slice window duration is shortened to facilitate resource reallocation. During off-peak periods, the window duration is lengthened to reduce overhead. For each slice window, an optimization method is used to solve the constrained slice resource allocation problem.
[0033] A collaborative workflow scheduling method based on Double Deep Q-Learning Network (DDQN) is designed to allocate tasks at small time scales within a decision slice window, balancing the overall network load. This method comprehensively considers factors such as vehicle speed and direction, association patterns, base station resources, and task type. Simulation results demonstrate that the proposed scheme outperforms benchmark methods in terms of adaptability, resource utilization, and task completion rate.
[0034] 2 System Model
[0035] This section introduces the RAN slicing framework, communication model, and collaborative workflow scheduling framework.
[0036] like Figure 1 As shown, consider a SAGVN scenario comprising a constellation of low-Earth orbit satellites, ground base stations, and drones. Ground base stations and drones have limited coverage, while satellites can seamlessly cover the entire road network. Vehicles are equipped with three types of transceivers that can connect to satellites, ground base stations, or drones, but can only connect to one base station at a time slot. Satellites connect to the core network via ground workstations; ground (drone) base stations connect to the core network via wired (wireless) connections. A MEC-assisted controller connects to various base stations through the core network, responsible for allocating and scheduling resources and tasks on the RAN side.
[0037] 2.1 RAN Slicing Framework
[0038] Consider a service-oriented RAN slicing framework. The physical resources of each satellite / ground / UAV base station are arranged into two RAN slices, named slice 1 and slice 2, which are used to handle latency-sensitive and latency-tolerant tasks, respectively. Task type o=1 (o=2) represents latency-sensitive (latency-tolerant) tasks. The former includes applications such as autonomous vehicle formation control
[14] , with a latency constraint of 100-150ms; the latter corresponds to autonomous vehicle high-definition map download
[24] , with a more relaxed latency requirement. The set of satellite, ground base station and UAV base station is represented as and base station The amount of spectrum resources and computing resources held are respectively denoted as c. j and s j The amount of spectrum and computing resources allocated by base station j to slice o∈{1,2} is represented by c. j,o and s j,o .
[0039] The duration of the slice window can be adaptively adjusted according to the network situation. For example... Figure 2 As shown, time is divided into a series of slice windows of unequal length. Each slice window contains multiple scheduling slots. The set of scheduling slots contained in slice window w is represented as... The duration of the slice window w is represented as f. (w) The controller collects workflow scheduling decisions within slice window w to determine the resource allocation strategy for slice window w+1. At the start of slice window w, the spectrum and computing resources of each base station are allocated according to the workflow scheduling decisions within window w-1. RAN slicing decisions continue until the end of slice window w. During scheduling time slots... Initially, the controller distributes the collected tasks to different base stations for processing. The base stations allocate resources to the tasks and transmit the processed results back to the original vehicle. At the end of each slice window, the controller collects the workflow scheduling decisions for use in the next resource allocation.
[0040] 2.2 Communication model
[0041] Since the satellite and the vehicle are far apart, the impact of changes in the vehicle's position on the vehicle-to-satellite channel gain is negligible within a local area. The average channel gain of vehicle i within the coverage area of base station j is expressed as g. i,j The method in
[25] is used for quantification.
[0042] The transmit power of vehicle i and base station j is represented as p i and p j During interaction with base station j, vehicle i will experience interference from other base stations. Spectrum resources in the slice are allocated to vehicles in an orthogonal manner. If base station j allocates a bandwidth of c to task m generated by vehicle i... i,j,m , σ 2 The uplink transmission rate when vehicle i submits task m to base station j, representing the average background noise, is calculated as follows:
[0043]
[0044] Where σ 2 This represents the average background noise. The downlink transmission rate from base station j back to task m and then to vehicle i is...
[0045]
[0046] 2.3 Workflow Scheduling Framework
[0047] To address the high-speed mobility of vehicles, this invention designs a collaborative workflow scheduling method. Task scheduling no longer relies on a single base station, but instead allows task offloading and processing to be performed on different base stations. Each base station contains two queues, named processing queue 1 and 2, to buffer latency-sensitive and latency-tolerant tasks. The MEC controller also contains two corresponding queues, named offloading queue 1 and 2, to buffer the two types of tasks transferred from the data acquisition base station. Integrating multi-source information, tasks in the offloading queues are transferred to different base stations for collaborative processing. An example is provided below for understanding:
[0048] 1) Delay-Sensitive Task Scheduling: As shown in Figure 3(a), when a vehicle generates a task, it is within the coverage area of both the satellite and ground base station b1. Following the proximity principle, the task is collected by base station b1 and then transferred to the offloading queue 1 of the MEC controller. Based on the vehicle's direction of travel and speed, the satellite, UAV, and base station b2 are selected as candidate cooperating base stations. Due to the low latency requirement, the controller selects base station b2, which has a lower load, to process the task. Base station b2 allocates resources for the task according to the first-come, first-served (FCFS) rule and transmits the processed results back to the vehicle.
[0049] 2) Delay-Tolerant Task Scheduling: As shown in Figure 3(b), when a vehicle generates a task, it is within the coverage area of a satellite, a drone, and base station b2. The drone is selected as the receiving base station based on proximity. The drone transfers the received task to the controller's offloading queue 2. Based on vehicle speed and direction of travel, either the satellite or base station b1 is listed as a candidate cooperating base station. The controller selects the satellite with the lower load to process the task.
[0050] As can be seen from the above cases, collaborative workflow scheduling needs to comprehensively consider factors such as vehicle location, speed, driving direction and base station load. This invention quantifies task queuing delay based on queuing theory
[26] , providing a basis for workflow scheduling. The data volume (bits), required computing resources and delay constraints of task m are respectively represented as ε m ,τ m ,d m .
[0051] 2.3.1 Unloading Delay Calculation
[0052] Unloading latency refers to the time it takes for a task to be published and submitted by the base station to the controller's unloading queue.
[0053] Within the slice window w, the set of vehicles is represented as The set of tasks of type O collected by base station j and the cardinality of that set are represented as follows: and Where o = 1 (o = 2) represents a delay-sensitive (delay-tolerant) task. Let α i,m =1 means task m is loaded by vehicle i, otherwise α i,m The value is 0. According to equation (1), the average time for a task of type o to be transmitted from the vehicle to the base station is calculated as follows:
[0054]
[0055] The arrival of tasks for individual vehicles and base stations is modeled as a Poisson process. Let the binary variables... This means that vehicle i has established a task upload connection with base station j (equivalent to step ① in Figures 3(a) and 3(b)); otherwise... The arrival rate of tasks in the controller's unloading queue o is 0.
[0056]
[0057] in This represents the arrival rate of tasks of type O generated by vehicle i in window w.
[0058] The unloading queue processes only one task at a time. The task unloading process is modeled as an M / M / 1 queue model. The service intensity of the unloading queue o is defined as...
[0059]
[0060] Enqueueing is determined by the task arrival rate, while dequeueing is determined by the task allocation rate. When the enqueueing rate exceeds the dequeueing rate, the continuously accumulating tasks may cause queue overflow. To maintain queue stability (prevent overflow), equation (5) needs to satisfy...
[0061]
[0062] After task m arrives in the unloading queue, the set of indices of tasks preceding task m is denoted as Ω(m). The delay at which task m generated by vehicle i is forwarded by base station j to the controller is denoted as ζ. i,j,m The unloading delay for this task is calculated as follows:
[0063]
[0064] 2.3.2 Processing Delay Calculation
[0065] Processing latency refers to the time it takes for a task to be processed from the time it enters the processing queue of the cooperating base station.
[0066] The base station allocates computing resources to each task as needed. Assume the maximum CPU cycles for base station j′ are... Hz (per second). Within the slice window w, the average processing time for tasks in processing queue o at this base station is calculated as...
[0067]
[0068] Tasks in the controller's offload queue are distributed to the processing queues of different base stations. The arrival of tasks in the processing queues also follows a Poisson process. The proportion of spectrum resources allocated by base station j′ to slice o out of the total spectrum resources of all slices of the same type is...
[0069]
[0070] The task arrival rate in processing queue o of cooperative base station j′ is The task processing process is modeled as an M / M / 1 queue model. Based on (4), (8), and (9), the service intensity of the processing queue o in the cooperative base station j′ is defined as
[0071]
[0072] To maintain the stability of the processing queue o, equation (10) needs to satisfy...
[0073]
[0074] In the processing queue of base station j′, the set of task indices preceding task m is denoted as Ψ. j′ (m). The processing delay of task m is calculated as...
[0075]
[0076] 2.3.3 Calculation of handover delay
[0077] After each task is computed in the processing queue of the cooperating base station j′, the base station transmits the result back to the vehicle (corresponding to step ⑤ in Figures 3(a) and 3(b)). The volume of the result data of task m generated by vehicle i after processing is denoted as θ. i,m Based on equation (2), the delay at which base station j′ transfers the processing result of task m to vehicle i is expressed as:
[0078]
[0079] The service delay D of task m generated by vehicle i i,m It is the sum of equations (7), (12), and (13), that is
[0080]
[0081] Let the binary variable The controller will transfer the task generated by vehicle i to the cooperating base station j′ for processing (corresponding to step ⑤ in Figures 3(a) and 3(b)), otherwise the value will be 0. Assume the travel distance from when vehicle i issues task m to when it leaves the coverage area of base station j′ is ω. i,j′,m The driving speed is The time it takes for vehicle i to move out of the coverage area of base station j′ after submitting task m is:
[0082]
[0083] Considering both the duration and the inherent delay constraints of the task itself, the delay d of task m generated by vehicle i is... i,m The demand was reformulated as
[0084]
[0085] This is one of the conditions that must be met for successful return of results, but it is not the only condition. Changes in the vehicle's speed and direction may cause the vehicle to fail to "meet as scheduled" with the cooperating base station j′. In this case, even if the task is completed by the cooperating base station j′ within the time (as shown in equation (15)), the result cannot be returned to the vehicle i.
[0086] 3. Problem Modeling
[0087] This section models the joint optimization of RAN slicing and collaborative workflow scheduling as a constrained stochastic optimization problem.
[0088] Combining (14) and (15), the following binary variables are defined.
[0089]
[0090] e i,j′,m =1 if and only if the cooperating base station j′ transmits the processing result of task m back to vehicle i within the specified time.
[0091] Definition 1: Within a slice window w, the average reward the system receives for completing a task is defined as...
[0092]
[0093] Where u j′,o ∈(0,1) represents the reward factor for a task of type o successfully completed on the cooperative base station j′.
[0094] Definition 2: Within a slice window w, the average loss caused by incomplete tasks is defined as...
[0095]
[0096] Where h j′,o ∈(0,1) represents the loss factor for a task of type o that failed to be completed on the cooperative base station j′.
[0097] In the proposed framework, a challenging problem is the joint optimization of RAN slice resource orchestration and cooperative workflow scheduling. Within the slice window w, the sets of spectrum and computational resource allocation strategies are represented as follows:
[0098]
[0099] and
[0100]
[0101] The set of collaborative workflow scheduling strategies is represented as
[0102]
[0103] in Representing time slots The set of collaborative workflow scheduling strategies within the scope. The set of slice window indices and the cardinality of the set are represented as... And W. The problem of maximizing the number of tasks completed over a long cumulative time, P1, is modeled as
[0104]
[0105]
[0106]
[0107]
[0108]
[0109] (6)and(11) (19e)
[0110] P1 essentially allocates spectrum and computing resources to each slice through online decision-making, balancing the load on each base station and maximizing the average number of tasks completed over a long period. Constraint (19a) ensures that each base station holds a certain amount of spectrum resources for allocation. The amount of spectrum and computing resources allocated to each vehicle by each base station should not exceed its total resources, corresponding to constraints (19b) and (19c). Constraint (19d) means that each vehicle can only connect to a single base station. Constraint (19e) is a condition for maintaining queue stability. Resource allocation and workflow scheduling decisions both affect queue stability.
[0111] The objective of problem P1 is a long-running, non-smooth localization function. Constraint (19d) contains two binary integer variables, and the variables in constraint (19e) are also coupled. Therefore, it is difficult to obtain an exact optimal solution for problem P1 using traditional optimization methods.
[0112] 4 solutions
[0113] For ease of processing, P1 is decoupled into 3 subproblems:
[0114] 1) Slice window duration division;
[0115] 2) Resource allocation (large time scale);
[0116] 3) Collaborative workflow scheduling (small time scale).
[0117] These subproblems are solved alternately by the MEC controller, forming a continuously running closed loop. The behavior of the MEC controller is abstracted as a state machine containing three states. Each state corresponds to a subproblem-solving module. Whenever the system reaches a state, the corresponding functional module is activated:
[0118] • Adaptive slice window duration (State 1): When slice window w-1 ends, the controller determines the slice window length f based on task traffic fluctuations. (w) (Section 4.1 provides specific details).
[0119] • Resource allocation (State 2): After the duration of slice window w is determined, the workflow scheduling decision within slice window w-1 is made. Resource allocation decisions for a given window w and Given the conditions (discussed in detail in Section 4.2), the resource allocation decision for RAN slices is determined by the controller at the beginning of window w and remains unchanged until the slice window w ends.
[0120] • Collaborative workflow scheduling (State 3): The slice window w is divided into multiple equal-length scheduling slots. At the beginning of each scheduling slot, The DDQN algorithm is input into the design to determine workflow scheduling decisions (implementation details are described in Section 4.3). At the end of the last scheduling slot, all scheduling decisions within the slice window w are saved as... It was subsequently used for calculation and
[0121] 4.1 Slice Window Duration Division
[0122] The slice window duration partitioning subproblem aims to maximize the number of tasks completed by dynamically dividing the slice window duration.
[0123]
[0124] Considering the relatively long interval between RAN segmentations and the dynamic nature of the network, historical task traffic information obtained at the end of slice window w-1 is used to determine the duration of window w. P1.1 is simplified to
[0125]
[0126] In reality, vehicle request issuance is time-varying and uncertain. If RAN slice resource allocation is performed within a fixed time window, RAN slice resource scheduling will be unable to cope with fluctuations in request arrival. During peak traffic periods, the proportion of various tasks will fluctuate continuously and significantly. In this case, shortening the window duration can promote resource reallocation and adapt to fluctuations in task traffic. During idle periods, the proportion of various tasks is relatively stable.
[27] At this point, the window duration can be appropriately extended to reduce unnecessary expenses.
[0127] This study explores the optimal match between task traffic fluctuations and slice window duration using experimental methods. The minimum granularity for window length adjustment is 10 minutes. A relatively short window duration is preset in the system, and the window length is then increased tentatively. Multiple initial time points are selected, and numerical pairs of task traffic fluctuations and the optimal slice window duration are collected. The following function is constructed to fit these numerical pairs.
[0128] y = αlog2x + β (20)
[0129] The fitting process involves finding parameters α and β that minimize the sum of squared residuals. Two fitting curves are generated. Figures 5(a) and 5(b) correspond to the case where the task flow continuously decreases (increases), where the slice window duration gradually increases (decreases) as the flow decreases (increases). It is evident that the more drastic the fluctuation in task flow, the smaller the optimal slice window length. This pattern is consistent with expectations.
[0130] At the end of slice window w-1, the ARIMA-ANN model
[27] Used to predict the task flow at the beginning of the next window w, the predicted value is represented as ARIMA and ANN models are suitable for handling historical data with linear and non-linear characteristics, respectively; their synergy can improve prediction accuracy. Based on (20) and The duration of the slice window w is determined as follows:
[0131]
[0132] Where γ is a constant representing the minimum unit length of the slice window. and It represents rounding up and down.
[0133] 4.2 Resource Allocation
[0134] The resource allocation subproblem maximizes the number of tasks completed by allocating spectrum and computational resources across RAN slices, described as follows:
[0135]
[0136] st(19a),(19b),(19c)
[0137] According to equations (17) and (18), the decision for each slice window is independent, and resources are allocated independently to each task within the window. In reality, traffic flow does not fluctuate continuously and drastically, and the traffic flow in adjacent slice windows is similar. Based on the workflow scheduling decision within the upper slice window w-1, the controller can calculate the amount of communication and computing resources required for each slice within window w. Accordingly, P1.2 is transformed into a one-shot optimization problem that maximizes the number of tasks completed within each window, i.e.
[0138]
[0139] st(19a),(19b)and(19c)
[0140] P1.2a is a multivariate optimization problem with multiple constraints. The Lagrange Multiplier is used to solve the problem, transforming a multivariate optimization problem with multiple constraints into a multivariate unconstrained optimization problem. Let... and As parameters of this extremum problem, problem P1.2a is transformed into
[0141]
[0142] The optimal resource allocation scheme for P1.2b can be obtained using the gradient descent method.
[0143] 4.3 Collaborative Workflow Scheduling
[0144] The collaborative workflow scheduling subproblem maximizes the number of tasks completed under latency constraints by selecting appropriate cooperative base stations for the collected tasks.
[0145]
[0146] st(19d),(6),(11)
[0147] As described in Section 4.2, the resource orchestration operations for each slice window in P1.1 are independent of each other. Given a fixed resource allocation, the collaborative workflow scheduling operations within each slice window are also independent. Therefore, the long-run optimization problem in P1.3 can be decomposed into a short-run optimization problem for each individual slice window, which is a Markov decision problem with a finite horizon.
[0148] The collaborative workflow scheduling subproblem within a single slice window is constructed as a Markov decision process (MDP). The MEC controller is abstracted as an agent. Training rounds The environmental state is represented as Controller according to Perform workflow scheduling actions The rewards given by the environment are represented as The controller is based on the state transition probability. Update the environment status to The states, actions, and rewards are expressed as follows:
[0149] • State space S: Workflow scheduling needs to consider task parameters, vehicle information, resources and load of each base station, etc. The number of tasks in processing queue o in base station j is... Vehicle i is at position l i Training rounds The state is represented as
[0150]
[0151] • Action Space A: The system in the training round The workflow scheduling action performed is represented as
[0152]
[0153] in Represents training rounds The workflow scheduling decision set within the controller is the set of tasks that the controller assigns to different cooperating base stations. In (19d), the decision variable for each action is either 0 or 1, determined by the current state.
[0154] • Reward R: The reward reflects the quality of the action performed in a given state. The system objective shifts from maximizing the number of tasks completed to maximizing the reward obtained. Based on (20) and (21), the reward is represented as...
[0155]
[0156] in, This is represented by the training round. The total rewards obtained from completing the task due to internal factors. Represents the training round Inside
[0157] The total loss from failed tasks. Workflow scheduling actions determine the base station's task processing. If a task is completed, the environment provides a reward to acknowledge the action. Simultaneously, the system introduces a penalty mechanism to prevent decisions that could lead to high base station load or disrupt processing queue stability.
[0158] In MDP, workflow scheduling refers to the controller maximizing rewards by allocating tasks from the offload queue to different cooperating base stations.
[0159]
[0160] Where Π is the set of all possible allocation strategies. In epoch The discount factor. Due to the unpredictability of request publication, state transitions are difficult to determine. Problem P1.2a Traditional model-based methods (such as value iteration and policy iteration) cannot be used.
[28] The solution involves a model-free approach that does not rely on state transition probabilities. However, due to the complexity of collaborative workflow scheduling, traditional model-free RL algorithms struggle to handle complex action and state spaces. Deep Q-Learning Network (DQN), as an improvement on Q-learning, does not rely on prior knowledge and can adapt to large action and state spaces. DDQN further separates the prediction and target networks in DQN training, avoiding overestimation caused by bootstrapping. Therefore, this section designs a DDQN-based method to address the collaborative workflow scheduling subproblem.
[0161] The core of Q-learning lies in constructing a Q-table. In the state space, the reward for each action is estimated and stored in the Q-table. The action-value function is expressed as... The maximum reward for each state in the Q-table represents the maximum possible future return. By consulting the Q-table, the action with the maximum reward in each state is determined.
[0162]
[0163] Applying the Bellman Equation to equation (25), the values in the Q table can be obtained. The calculation process is as follows:
[0164]
[0165] In the above formula, φ represents the learning rate, and υ represents the greedy probability.
[0166] like Figure 6 As shown, the DDQN-based workflow scheduling scheme uses two identical neural networks (prediction network and target network) for training. Q-learning and mean squared error are used to construct the loss function. The DDQN-based collaborative workflow scheduling is described as Algorithm 1. Compared to DQN, this scheme adds an experience replay pool and a target network. The experience replay mechanism constructs a data pool. A set of data is randomly drawn from the data pool during each training session, improving data utilization and reducing training correlation. DDQN parameter updates depend on... and The target network and the evaluation network work together to update the parameters, thus avoiding overestimation.
[0167] Algorithm 1: DDQN-based workflow scheduling algorithm
[0168]
[0169]
[0170] 5. Performance Evaluation
[0171] The effectiveness of the proposed scheme was verified using simulation methods. A 1000-meter-long four-lane highway was considered. The origin of the coordinate system was the starting point of the road. The scenario included two macro base stations, three drones, and one satellite base station. The satellite covered the entire road, and each of the two macro base stations covered approximately 500 meters of the road. The drone hovered above the road, with an effective coverage radius of 80 meters. The transmission powers of the satellite, macro base stations, and drones were 27W, 40dBm, and 0.1W, respectively. Traffic flow data was selected from the OpenITS open data platform①. The vehicle density on the road was set to 0.4 (vehicles / m²). Unmanned vehicle platooning control and high-definition map download for autonomous driving were used to simulate latency-sensitive and latency-tolerant tasks. Other simulation parameters are shown in Table 1.
[0172] Four benchmark methods were selected. Consistent with the proposed scheme, each benchmark method includes three functional modules: slice window adjustment, resource allocation, and workflow scheduling. These modules are integrated into, for example, Figure 4 Within the framework shown, Table 2 provides the implementation details of the different methods.
[0173] 5.1 Convergence Analysis
[0174] In DRL, learning speed and training performance are affected by the update cycle and learning rate. The agent's preference for long-term and short-term rewards is influenced by the discount rate in the cumulative discounted reward. Our simulations observe the reward acquisition and convergence of the proposed method when the initial learning rate is set to 0.1, 0.005, and 0.001, respectively.
[0175] The reward value is directly proportional to the number of tasks completed. For example... Figure 7 As shown, when the learning rate is high at 0.1, the maximum reward converges to around 1500. The reward value fluctuates significantly throughout the process. When the learning rate is relatively small at 0.001, the maximum reward converges to approximately 2700. At this point, the system is trapped in a local optimum, and even increasing the number of training epochs makes it difficult to improve the training effect. A compromise learning rate of 0.005 outperforms the previous two settings. It not only obtains the maximum reward but also steadily improves the training effect with each epoch. When training reaches 100 epochs, the reward value rises to 3157.
[0176] 5.2 Impact of Training Rounds on Performance
[0177] Next, we observe the impact of the number of training epochs on performance. The DRL-based approach uses the same settings and processes the same 4000 data points. Baseline-4 only considers link quality and does not involve model training; its results are only used as a reference baseline.
[0178] As shown in Figure 8(a), the rewards obtained in the first 20 training rounds show a rapid growth trend, after which the growth slows down. Since the parameters are randomly selected, the agent cannot adapt to the environment initially; only through learning with a large amount of data can it capture data correlations and update parameters until convergence. Figure 8(b) shows the number of tasks completed by different schemes. After 5 training rounds, the number of tasks completed by the proposed scheme and Baseline-3 begins to exceed that of Baseline-4, and then continues to rise steadily. The number of tasks completed by the proposed scheme is consistently higher than that of Baseline-3. In Figure 8(c), the task failure rate of Baseline-4 remains at 29%. As a variant of DQN, DDQN reduces data correlation, resulting in better learning and convergence performance. After 100 training rounds, the task failure rates of the proposed method and Baseline-1 are approximately 21% and 25%, respectively. The former consistently outperforms the latter.
[0179] 5.3 Task Completion Effect Analysis
[0180] This simulation aims to verify the performance improvement effect of the proposed adaptive slicing window strategy. In Figure 9(a), as the proportion of latency-sensitive tasks increases, the task failure rate of the static window schemes (baseline-1 and baseline-2) shows an upward trend, while the task failure rate of the dynamic window schemes (the proposed scheme and baseline-3) remains relatively stable, indicating that the dynamic window scheme is more adaptable to load fluctuations. Figure 9(b) shows the number of slicing windows generated by different methods within 2 hours. Whenever a new window arrives, the controller triggers the resource reallocation of RAN slices, generating significant signaling overhead. Combining Figures 9(a) and 9(b), it can be seen that the proposed method has a lower number of windows and a lower task failure rate than baseline-1, meaning that the proposed method can provide higher quality service with lower overhead, proving the effectiveness of the dynamic window scheme.
[0181] Figure 10(a) illustrates the impact of increasing spectrum resources on task failure rate when the number of computing resources is fixed at 15. The task failure rates of each scheme continuously decrease, and the differences gradually narrow, eventually stabilizing at around 10%. Sufficient spectrum resources provide the controller with greater decision space, which is an important condition for performance improvement, but not the only one. Next, we examine the performance improvement effect of increasing computing resources when the number of sub-channels is fixed at 20. As shown in Figure 10(b), the task failure rate decreases rapidly in the initial stage, but when the number of computing resources reaches 15, further increasing computing resources no longer helps to improve performance; at this point, the performance bottleneck lies in spectrum resources.
[0182] Now, we simulate a scenario where the number of task releases continuously increases within one hour. As shown in Figure 11(a), due to resource constraints, the overall task failure rate shows an upward trend. Due to the lack of flexibility in the Max-SINR scheme, the task failure rate of Baseline-4 increases from 35% to 52%. The failure rates of DQN-based baselines-2 and-3 increase from 20% to 31%. Thanks to the collaboration of heterogeneous base stations, the task failure rate of the proposed method remains lower than that of other schemes as it increases from 18% to 28%. The increase in the proportion of latency-sensitive tasks also leads to a decrease in task completion rate. As shown in Figure 11(b), when the proportion of such tasks is 0.2, the task completion rate of baseline-4 is 52%, while the task completion rates of baseline-3 and the proposed scheme are 78% and 85%, respectively; when the proportion is 0.8, the task completion rates of the proposed scheme, baseline-3, and baseline-4 are 58%, 53%, and 26%, respectively. The workflow scheduling strategy generated by the proposed scheme is more reasonable compared to other benchmark methods.
[0183] References
[0184] [1] Zhuang, W., Ye, Q., Lyu, F., Cheng, N., & Ren, J. SDN / NFV empowered futureIoV with enhanced communication, computing, and caching. Proceedings of the IEEE, 2020, 108(2): 274-291.
[0185] [2] Zhang, W., Zhang, Z., & Chao, HCCooperative fog computing for dealing with big data in the internet of vehicles: Architecture and hierarchical resource management. IEEE Communications Magazine, 2017, 55(12): 60-67.
[0186] [3] Liu, J., Shi, Y., Fadlullah, ZM, & Kato, N. Space-air-ground integrated network: A survey. IEEE Communications Surveys&Tutorials, 2018, 20(4): 2714-2741.
[0187] [4]Zeng Y, Zhang R, Lim T J. Wireless communications with unmannedaerial vehicles: Opportunities and challenges [J]. IEEE Communications Magazine, 2016, 54(5): 36-42.
[0188] [5] Dong Chao, Tao Ting, Feng Simeng, et al. A review of media access control protocols for UAV ad hoc networks and vehicle-to-everything (V2X) networks [J]. Journal of Electronics and Information Technology, 2022, 44: 1-13.
[0189] [6]Ning, Z., Hu,
[0190] [7] Sexton, C., Marchetti, N., & DaSilva, LACustomization and tradeoffs in 5G RAN slicing. IEEE Communications Magazine, 2019, 57(4): 116-122.
[0191] [8] Dong Chao, Tao Ting, Feng Simeng, et al. A review of media access control protocols for UAV ad hoc networks and vehicle-to-everything (V2X) networks [J]. Journal of Electronics and Information Technology, 2022, 44: 1-13.
[0192] [9] Zhang N, Zhang S, Yang P.Software defined space-air-groundintegrated vehicular networks: Challenges and solutions. IEEE Communications Magazine, 2017, 55(7): 101-109.
[0193]
[10] Lyu F, Yang P, Wu H, et al. Service-oriented dynamic resource slicing and optimization for space-air-ground integrated vehicular networks [J]. IEEE Transactions on Intelligent Transportation Systems, 2021.
[0194]
[11] Li J, Shi W, Yang PA hierarchical soft RAN slicing framework for differentiated service provisioning[J]. IEEE Wireless Communications, 2020, 27(6):90-97.
[0195]
[12] Ye Q,Zhuang W,Zhang S,Jin AL,Shen X,Li X.Dynamic radio resourceslicing for a two-tier heterogeneous wireless network.IEEE Transactions onVehicular Technology,2018,67(10):9896-9910.
[0196]
[13] Peng,H.,Ye,Q.,&Shen,X.Spectrum management for multi-access edgecomputing in autonomous vehicular networks.IEEE Transactions on IntelligentTransportation Systems,2019,21(7):3001-3012.
[0197]
[14] Wu,W.,Chen,N.,Zhou,C.,Li,M.,Shen,X.,Zhuang,W.,&Li,X.Dynamic RANslicing for service-oriented vehicular networks via constrained learning.IEEEJournal on Selected Areas in Communications,2020,39(7),2076-2089.
[0198]
[15] Chen,M.,Hao,Y.,Hu,L.,Huang,K.,&Lau,V.K.Green and mobility-awarecaching in 5G networks.IEEE Transactions on Wireless Communications,2017,16(12),8347-8361.
[0199]
[16] Ji,J.,Zhu,K.,Niyato,D.,&Wang,R.Joint cache placement,flighttrajectory,and transmission power optimization for multi-UAV assistedwireless networks.IEEE Transactions on wireless communications,2020,19(8),5389-5403.
[0200]
[17] Sun X,Ansari N.Jointly optimizing drone-mounted base stationplacement and user association in heterogeneous networks.IEEE InternationalConference on Communications 2018:1-6.
[0201]
[18] Lim W Y B,Luong N C,Hoang D T.Federated learning in mobile edgenetworks:A comprehensive survey[J].IEEE Communications Surveys&Tutorials,2020,22(3):2031-2063.
[0202]
[19] C.Kai,H.Zhou,Y.Yi and W.Huang,Collaborative Cloud-Edge-End TaskOffloading in Mobile-Edge Computing Networks With Limited CommunicationCapability,IEEE Transactions on Cognitive Communications and Networking,2021,7(2):624-634.
[0203]
[20] Mushu Li, Jie Gao, Lian Zhao, Xuemin Shen. Deep Reinforcement Learning for Collaborative Edge Computing in Vehicular Networks, IEEE Transactions on Cognitive Communications and Networking, 2020, 6(4): 1122–1135.
[0204]
[21] Peng H, Shen X. Deep reinforcement learning based resource management for multi-access edge computing in vehicular networks. IEEE Transactions on Network Science and Engineering, 2020, 7(4): 2416-2428.
[0205]
[22] Xu Xiaolong, Fang Zijie, Qi Lianyong, et al. A distributed service offloading method based on deep reinforcement learning in the edge computing environment of vehicle networking [J]. Chinese Journal of Computers, 2021.
[0206]
[23] Zhu Zhengze, Zhou Haiying, Fu Yongzhi, et al. Cooperative control of connected autonomous vehicles based on delay compensation [J]. Journal of System Simulation, 2019, 31(7):1448.
[0207]
[24] Javanmardi, E., Gu, Y., Javanmardi, M., & Kamijo, S. Autonomous vehicle self-localization based on abstract map and multi-channel LiDAR in urban area. IATSS research, 2019, 43(1): 1-13.
[0208]
[25] Erceg,V.,Greenstein,L.J.,Tjandra,S.Y.,Parkoff,S.R.,Gupta,A.,Kulic,B.,...&Bianchi,R.An empirically based path loss model for wirelesschannels in suburban environments.IEEE Journal on selected areas incommunications,1999,17(7):1205-1211.
[0209]
[26] Xue,J.,Wang,Z.,Zhang,Y.,&Wang,L.Task allocation optimizationscheme based on queuing theory for mobile edge computing in 5G heterogeneousnetworks.Mobile Information Systems,2019,62(2):1-3.
[0210]
[27] Zeng D,Xu J,Gu J.Short term traffic flow prediction using hybridARIMA and ANN models.IEEE Workshop on Power Electronics and IntelligentTransportation System.,2008,621-625.
[0211]
[28] MNIH,Volodymyr,et al.Playing atari with deep reinforcementlearning.arXiv preprint arXiv:1312.5602,2013.
[0212]
[29] GOLCHI,Mahya Mohammadi;SARAEIAN,Shideh;HEYDARI,Mehrnoosh.A hybridof firefly and improved particle swarm optimization algorithms for loadbalancing in cloud environments:Performance evaluation.Computer Networks,2019,162:106860.
[0213]
[30] Ye,Q.,Shi,W.,Qu,K.,He,H.,Zhuang,W.,&Shen,X.Joint RAN slicing andcomputation offloading for autonomous vehicular networks:A learning-assistedhierarchical approach.IEEE Open Journal of Vehicular Technology,2021,2:272-288.
[0214]
[31] PENG,Haixia;YE,Qiang;SHEN,Xuemin.Spectrum management for multi-access edge computing in autonomous vehicular networks.IEEE Transactions onIntelligent Transportation Systems,2019,21.7:3001-3012.
[0215]
[32] ZHANG,Chuanting.Dual attention-based federated learning forwireless traffic prediction.In:IEEE INFOCOM 2021-IEEE conference on computercommunications.IEEE,2021.p.1-10.
[0216]
[33] Ye,Q.,Zhuang,W.,Zhang,S.,Jin,A.L.,Shen,X.,&Li,X.Dynamic radioresource slicing for a two-tier heterogeneous wireless network.IEEETransactions on Vehicular Technology,2018,67(10):9896-9910.
Claims
1. A slice-based cooperative task offloading method in space-air-ground integrated vehicle networking, wherein the space-air-ground integrated vehicle networking (SAGVNs) scenario includes a low earth orbit satellite group, a ground base station, and a UAV base station, and a satellite serves as a satellite base station, hereinafter referred to as a satellite. Unmanned aerial base station, hereinafter referred to as UAV; Satellites seamlessly cover the entire road network; The signal transceivers equipped on vehicles are connected to satellites, ground base stations and UAVs respectively, and only one base station is connected in the same time slot; the satellites are connected to the core network through the ground station; the ground base stations and UAVs are also connected to the core network; The MEC controller connects various base stations through the core network and is responsible for allocating and scheduling resources and tasks on the RAN side; the resources include spectrum resources and computing resources; The design steps of the cooperative task offloading method include: First, a service-oriented RAN slicing framework is designed, which supports adaptive slicing window length, spectrum and computing resource orchestration, and cooperation between heterogeneous base stations; In this RAN slicing framework, based on the M / M / 1 queuing model, the RAN slicing and task offloading joint decision is modeled as a problem of maximizing the long-term task completion number; Then, the problem of maximizing the long-term task completion number is decoupled into three sub-problems: slicing window length division problem, resource allocation problem and cooperative workflow scheduling problem, which are solved alternately by the MEC controller to form a closed loop with the slicing window as the period; The behavior of the MEC controller is abstracted as a state machine containing 3 states, each state corresponding to a sub-problem solving module; when a state comes, the corresponding solving module is activated: The solution method of the slicing window length division problem is to determine the window length through a task flow perception strategy; The solution method of the resource allocation problem is to allocate resources to slices through an optimization method; The solution method of the cooperative workflow scheduling problem is to use the DDQN method to decide the task scheduling within the slicing window and determine to offload tasks to the corresponding base station for processing; The MEC controller collects the workflow scheduling decisions within the current slicing window to determine the resource allocation strategy for the next slicing window; at the beginning of the current slicing window, resources are allocated to each base station according to the workflow scheduling decisions in the previous slicing window; at the beginning of each scheduling time slot within the slicing window, the controller transfers the collected tasks to different base stations for processing; the base station allocates resources for the task and returns the processed results to the original vehicle; at the end of each slicing window, the controller collects the workflow scheduling decisions within the window for use in the next resource allocation.
2. The method of claim 1, wherein the method further comprises: determining a slice type of the task; and determining a slice type of the target computing resource. In the RAN slicing framework, the physical resources of each satellite, ground base station and UAV base station are arranged into 2 RAN slices for processing delay-sensitive tasks o=1 and delay-tolerant tasks o=2; o represents the task type; The set of satellites, ground base stations and drone stations are denoted as and Base stations The number of spectrum resources and computing resources held are denoted as c j and s j ; The number of frequency spectrum resources and computing resources allocated by the base station j to the slice o e {1,2} is denoted as c j,o and s j,o ; The slicing window length is adaptively adjusted according to the network situation; time is divided into a series of slicing windows of unequal length, each slicing window contains multiple scheduling time slots; The set of scheduled time slots contained by the slice window w is denoted as The duration of the slice window w is denoted as f (w) ; During task scheduling, the offloading and processing of tasks are allowed to be performed at different base stations; each base station contains two processing queues to buffer the collected delay-sensitive and delay-tolerant tasks, and the MEC controller contains two corresponding offloading queues to buffer the delay-sensitive and delay-tolerant tasks transferred from the collection base station; based on comprehensive multi-source information, the tasks in the offloading queue are transferred to different base stations for cooperative processing. 3.The slice-based cooperative task offloading method in space-air-ground integrated vehicular networking according to claim 2, characterized in that RAN slicing and task offloading joint decision is modeled as a long-term task completion maximization problem P1: Define binary variable e i,j′,m = 1 indicates if and only if the cooperating base station j' returns the processing result of the task m to the vehicle i within a specified time; Definition 1 The average reward obtained from completing tasks within a slicing window w is defined as where u j′,o ∈ (0,1) represents a reward factor for successful completion of task of type o by cooperating base station j' on link u; Definition 2 The average loss caused by not completing tasks within a slicing window w is defined as where h j′,o ∈(0,1) represents a loss factor for a task of type o that failed to complete on the cooperating base station j'. Joint optimization of RAN slicing resource orchestration and collaborative workflow scheduling: The set of spectrum and computing resource allocation strategies within slicing window w are denoted as and The set of collaborative workflow scheduling strategies is denoted as wherein representing a time slot cooperative workflow scheduling policy set within a slice window; a set of slice window indices and a cardinality of the set denoted as and W; P1 is modeled as Service intensity of the offload queue o service intensity of the processing queue o in the cooperating base station j' Constraint (a) guarantees that each base station holds a certain amount of spectrum resources for allocation; Constraints (b) and (c) guarantee that the amount of spectrum and computing resources allocated to vehicles by each base station should not exceed the total amount of resources it holds; Constraint (d) guarantees that each vehicle can only connect to a unique base station; Constraints (e) and (f) guarantee to maintain queue stability.
4. The method of claim 3, wherein the method further comprises: determining a slice type of the task; and determining a slice type of the target computing resource. Slicing window length partition sub-problem: The slicing window length partition sub-problem aims to maximize the number of tasks completed by dynamically partitioning the slicing window length, i.e. Due to the long interval of RAN slicing windows and network dynamics, the historical task flow information obtained at the end of the lower slicing window w-1 is used to determine the length of window w, then P1.1 is simplified as Collect the numerical pairs of task flow fluctuations and optimal slicing window length, construct the function y = a log2x + β to fit these numerical pairs, and find the parameters a and β that minimize the sum of squared residuals; At the end of the previous slice window w-1, the ARIMA-ANN model is used to predict the task traffic at the beginning of the window w, and the predicted value of the task traffic is denoted as The task traffic value of the previous slice window w-1 is denoted as Duration f of the slice window w (w) For wherein where γ is a constant representing the minimum unit of the slice window length, and respectively represent the upward and downward rounding, and α1 and β1, α2 and β2 represent the parameters of the minimized residual square sum of the corresponding functions, respectively.
5. The method of claim 3, wherein the method further comprises: determining a slice type of the task; and determining a slice type of the target node based on the slice type of the task. Resource allocation sub-problem: The resource allocation sub-problem maximizes the number of tasks completed by allocating spectrum and computing resources to each RAN slice, denoted as The decision for each slicing window is independent, and each task within the window is allocated resources independently; due to the fact that vehicle flow will not experience continuous drastic fluctuations, and there is similarity in vehicle flow within adjacent slicing windows; based on the workflow scheduling decisions within the upper slicing window w-1, the required resource amount for each slice within window w is calculated, then P1.2 is transformed into a one-shot optimization problem that maximizes the number of tasks completed within each window, i.e. Lagrange Multiplier is used to solve the problem, which transforms a multi-variable and multi-constraint optimization problem into a multi-variable and unconstrained extreme value problem; let and become the parameters of the extreme value problem, and the problem P1.2a is transformed into The optimal resource allocation scheme for P1.2b is obtained by gradient descent method.
6. The method of claim 3, wherein the method further comprises: Collaborative workflow scheduling sub-problem: The collaborative workflow scheduling sub-problem maximizes the number of tasks completed under delay constraints by selecting appropriate collaborative base stations for the collected tasks, i.e. The resource orchestration operation for each slicing window is independent of each other, and under the condition that resource allocation is determined, the collaborative workflow scheduling operation within each slicing window is also independent of each other, then the long-term optimization problem in P1.3 is decomposed into a short-term optimization problem for a single slicing window, which belongs to a finite horizon Markov decision problem; The collaborative workflow scheduling subproblem within a single slice window is constructed as a Markov decision process (MDP), and the MEC controller is abstracted as an agency; training rounds The environmental state is represented as Controller according to Perform workflow scheduling actions The rewards given by the environment are represented as The controller is based on the state transition probability. Update the environment status to The states, actions, and rewards are then expressed as follows: • State space S: The number of tasks in the processing queue o in base station j is The vehicle i position is l i , training round The state is represented as • Action space A: In training round The workflow scheduling actions made are represented as wherein representing a training round a set of workflow scheduling decisions within the controller, i.e. the controller assigns a set of tasks to different cooperating base stations; In the constraint Each action corresponds to a decision variable of 0 or 1, determined by the current state. · Reward R: The reward reflects the quality of the action taken in a certain state; the goal is transformed from maximizing the number of tasks completed to maximizing the reward obtained; based on task flow fluctuations and optimal slicing window length and the length of slicing window w, the reward is denoted as wherein, represents the total reward obtained by the base station for completing tasks within the training round; represents the total loss of the base station for failing to complete tasks within the training round; The action decision of the workflow scheduling determines the task processing of the base station; if a task is processed, a reward is provided to affirm the action; meanwhile, a punishment mechanism is introduced to prevent the decision that may cause the base station to face high load or destroy the stability of the processing queue; In MDP, workflow scheduling refers to the controller allocating tasks in the offloading queue to different collaborative base stations to obtain the most rewards, i.e. wherein Π is the set of all possible allocation policies, is the discount factor at time t. Due to unpredictability of the release of vehicle requests, the DDQN-based method is adopted to deal with the cooperative workflow scheduling sub-problem; the input of the DDQN-based cooperative workflow scheduling method includes the information of tasks, the physical resources held by base stations and the information of queues, and the output is the optimal workflow scheduling scheme; In the state space, the reward obtained by each action is estimated and stored in the Q table; the action value function is represented as The maximum reward of each state in the Q table represents the maximum return that can be obtained in the future; by querying the Q table, the action a * is determined as The value in the Q table is obtained by using Bellman Equation, and the calculation process is as follows: In the formula, φ represents the learning rate, and υ represents the greediness probability. The DDQN-based workflow scheduling method uses the same prediction network and target network for training; Q-learning and mean square error method are used to construct the loss function.
Citation Information
Patent Citations
Prediction-based virtual network function scheduling method for 5G network slices
CN108965024A
Methods and apparatuses for network slice minimum and maximum resource quotas
US20210377814A1