Intelligent wireless access network resource allocation method based on delay SLA
By combining deep reinforcement learning and hybrid slicing methods in the OFDMA system, the resource allocation of URLLC and eMBB is optimized, solving the problems of insufficient URLLC resource utilization and eMBB rate loss, and achieving efficient resource allocation in the industrial Internet scenario.
Patent Information
- Application Number
- CN202410871348.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-07-01
AI Technical Summary
Existing orthogonal slicing methods cannot fully utilize URLLC slicing resources, non-orthogonal slicing methods lead to eMBB rate loss and excessive control signal overhead, and traditional methods have difficulty in effectively adjusting resource allocation when channel status changes.
Combining deep reinforcement learning with hybrid slicing methods, by building an OFDMA system model, we optimize deterministic traffic delay jitter and eMBB throughput, and use heuristic algorithms and deep reinforcement learning models to dynamically adjust resource allocation at large and small time scales.
While ensuring URLLC latency and reliability requirements, it reduces latency jitter, improves eMBB service throughput, and achieves efficient and adaptive configuration of slice resources.
Smart Images

Figure CN118843206B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communications, and in particular to a delay SLA-oriented intelligent wireless access network resource allocation method. Background Art
[0002] The fifth-generation new radio system supports three major application scenarios: enhanced mobile broadband (eMBB), ultra-reliable low-latency communications (URLLC), and massive machine-type communications (mMTC). Traditional communication networks are primarily designed to serve single mobile broadband services and are unable to adapt to the diverse business scenarios of 5G. To meet the heterogeneous service level agreement (SLA) requirements of different services in terms of speed, latency, and reliability, 5G adopts the concept of network slicing to achieve end-to-end guaranteed service provision and traffic isolation. An end-to-end network slice typically consists of a radio access network sub-slice, a transport network sub-slice, and a core network sub-slice. To implement slicing on the radio access network, appropriate resource scheduling strategies must be developed to enable efficient coexistence of different traffic types.
[0003] Slice resource allocation schemes for scenarios where eMBB and URLLC services coexist fall into two main areas of research. The first research direction is orthogonal slicing. Service providers reserve a portion of bandwidth for eMBB and URLLC users, respectively. This approach ensures service isolation between the two slices. However, due to the dynamic changes in URLLC traffic and channels, the resources allocated to URLLC slices may not be fully utilized. The second research direction is non-orthogonal slicing. To facilitate support for eMBB and URLLC network slicing, 5G New Radio has standardized mini-slot-based broadcasting and preemptive scheduling techniques for service multiplexing in the radio access network. Non-orthogonal slicing involves preempting the arrival of URLLC packets at the mini-slot scale to transmit them. This approach effectively meets the latency requirements of URLLC services, but preemptive scheduling results in rate loss for eMBB users and control signal overhead for preemption instructions.
[0004] In industrial internet scenarios, traffic primarily consists of eMBB services and URLLC services, with URLLC traffic constituting the majority of traffic. eMBB services have non-real-time, high-bandwidth requirements, while URLLC services require low latency and high reliability. Because orthogonal slicing methods fail to fully utilize the resources allocated to URLLC slices, non-orthogonal slicing methods significantly reduce eMBB rates when URLLC traffic dominates, while also requiring significant control signaling overhead. Therefore, a combination of orthogonal and non-orthogonal slicing methods is being considered to reduce resource preemption and improve eMBB service throughput while ensuring URLLC latency and reliability requirements.
[0005] Radio access network resource allocation in scenarios where eMBB and URLLC services coexist has been extensively studied in academia. However, traditional model-based optimization resource allocation methods assume that user traffic demand and channel gain distribution are known a priori. This is difficult to implement in industrial manufacturing scenarios where channel state information is constantly changing. Model-free deep reinforcement learning methods can learn optimal wireless resource allocation strategies by interacting with the environment without manual guidance. They can also dynamically adjust resource allocation strategies based on different wireless environments to adapt to different network conditions and user needs. The present invention combines deep reinforcement learning with a hybrid slicing method to determine the appropriate slice resource allocation strategy.
[0006] 5G can promote the digital transformation of the manufacturing industry. By integrating 5G networks with industrial networks, they can support deterministic communications with bounded latency requirements. URLLC traffic, a common type of deterministic traffic in industrial scenarios, is periodic, and ensuring deterministic latency for this type of traffic is of great value and significance. Due to the impact of small-scale fading on channel gain, the amount of resources required for deterministic traffic upon arrival is constantly changing. The computational complexity of determining the latency for each user using an exhaustive method scales exponentially with the number of users, resulting in excessive computational time when there are many URLLC users. Therefore, a URLLC user resource allocation method with low computational complexity is needed. Summary of the Invention
[0007] The present invention provides a latency SLA-oriented intelligent wireless access network resource allocation method to address the shortcomings in industrial Internet scenarios, such as the inability to fully utilize the resources allocated to URLLC slices by the existing orthogonal slicing method and the excessive control signal overhead and eMBB rate loss generated by the non-orthogonal slicing method.
[0008] An embodiment of the present invention provides a method for allocating intelligent wireless access network resources based on latency SLA, comprising the following steps:
[0009] Step 1: Build a single-base station downlink transmission OFDMA (Orthogonal Frequency Division Multiple Access) system model and formulate an optimization problem for minimizing the sum of deterministic traffic delay and jitter for intra-slice resource allocation. Also, establish an optimization problem for maximizing eMBB service throughput for inter-slice resource allocation, given a defined intra-slice resource allocation strategy.
[0010] Step 2: Convert the optimization problem of minimizing the sum of the deterministic traffic delay and jitter into the problem of minimizing the expected variance of the URLLC resource requirements for each mini-timeslot under deterministic delay and solve the optimization problem using a heuristic method;
[0011] Step 3: Set the location distribution and traffic model of URLLC users and eMBB users, and determine the resource allocation plan for the two services at a small time scale.
[0012] Step 4: To optimize the eMBB service throughput, within the framework of a deep reinforcement learning network, the cell downlink communication system is considered the environment, the base station is considered the agent, and the large-scale inter-slice resource allocation process is modeled as a Markov decision process. The reinforcement learning state, action, and reward functions are designed to build a deep reinforcement learning model.
[0013] Step 5: Based on the deep reinforcement learning model built in step 4, allocate resources based on the model output on a large time scale and on a small time scale based on the resource allocation scheme in step 3. Through the interaction between the base station and the wireless environment, learn the optimal strategy for slice resource allocation.
[0014] Optionally, in one embodiment of the present invention, step 1 further includes:
[0015] Step 1.1: A single-base station downlink transmission scenario includes two types of users: eMBB users and URLLC users. eMBB users have non-real-time, high-bandwidth requirements, while URLLC users require high reliability and low latency. Network slicing technology is used for resource allocation. Wireless spectrum resources are divided into shared slices and reserved slices for URLLC users.
[0016] In step 1.2, the eMBB user adopts slot-based transmission. The length of a slot is 1ms. At the beginning of each slot, the bandwidth of the shared slice is allocated to the eMBB user.
[0017] In step 1.3, URLLC users use mini-slot-based transmission. One time slot consists of seven mini-slots. At the beginning of each mini-slot, resources in the URLLC reserved slice are allocated based on the URLLC user scheduling priority. If resources in the reserved slice are insufficient, transmission is performed by preempting resources allocated to eMBB users.
[0018] Step 1.4: The delay of URLLC in the wireless access network includes transmission delay, queuing delay, and propagation delay. A short URLLC packet is transmitted within a mini-slot, so the transmission delay is the length of a mini-slot. The queuing delay of the mth data packet of the nth URLLC user is d n,m , subtract the transmission delay and propagation delay from the total delay requirement of the URLLC data packet to obtain the maximum queuing delay, and record the maximum allowed queuing delay of the nth URLLC user as D n ,The arrival of URLLC traffic includes random arrival and periodic arrival;
[0019] In step 1.5, URLLC users preempting resources allocated to eMBB users in a shared slice will cause rate loss for the eMBB users. The rate loss ratio of the eMBB users caused by URLLC preemption in time slot t is recorded as loss(t).
[0020] Step 1.6: Establish an optimization problem for minimizing the sum of delay and jitter of deterministic traffic. The details are as follows:
[0021]
[0022] Among them, b n,τ is the number of physical resource blocks required for URLLC user n to send data packets in mini-time slot τ, i n,τ Indicates whether URLLC user n sends a data packet in mini-timeslot τ. 1 indicates sending, and 0 indicates not sending. B is the total bandwidth. The constraints guarantee the latency and reliability of the URLLC user respectively.
[0023] Step 1.7: Establish the optimization problem of maximizing eMBB throughput as follows:
[0024]
[0025] stω e ≤B
[0026] Among them, R e,n (t) is the eMBB user e n The rate when the URLLC user does not preempt resources in time slot t. loss(t) is the eMBB rate loss caused by URLLC preemption in time slot t. ω e For the bandwidth allocated to eMBB users, the constraints ensure that the resources allocated to eMBB users do not exceed the total resources.
[0027] Optionally, in one embodiment of the present invention, step 2 further includes:
[0028] Step 2.1: Convert the deterministic traffic delay jitter minimization optimization problem into the problem of minimizing the expected variance of URLLC resource requirements for each mini-timeslot under deterministic delay. The details are as follows:
[0029]
[0030] Among them, l n is the deterministic delay of user n. When resources are sufficient, all URLLC data packets are transmitted with this deterministic delay. τThe expected value of resources required by URLLC users for transmission with deterministic delay in mini-slot τ is obtained by combining the packet size and statistical channel state information.
[0031] Step 2.2, in, is a constant, so the optimization problem is equivalent to:
[0032]
[0033] Step 2.3, the steps for solving the optimization problem using the heuristic method are as follows:
[0034] Using the packet size of URLLC periodic traffic and statistical channel status information, the expected value k of the number of RBs required to send each user data packet is calculated. n , and sort them in descending order of expected value;
[0035] Initialize all mini-slot RB usage to 0 and the user to 1;
[0036] repeat:
[0037] Within the user k delay requirement range, calculate the ∑ corresponding to each propagation delay τ b τ 2 Growth value, and at the same time determine the mini-slot occupied resources b corresponding to each delay τ Whether it exceeds the set threshold. If so, the delay is excluded.
[0038] Select ∑ τ b τ 2 The delay corresponding to the minimum growth value is updated τ and k; if all delays do not meet the requirements, the user is set as the first priority user, k = 1;
[0039] Until: the delay of all users is determined.
[0040] Optionally, in one embodiment of the present invention, step 3 further includes:
[0041] In step 3.1, URLLC users and eMBB users are randomly distributed within the cell. Each user adopts an appropriate traffic model and updates the user's location information and statistical channel state information at the beginning of each large time scale.
[0042] Step 3.2: At the beginning of each time slot, resources are allocated to eMBB users. The resources in the shared slice are allocated to specific users using a proportional fair resource allocation algorithm. At the beginning of each mini-time slot, resources are allocated to URLLC users. For random traffic packets in the buffer area, the D n -d n,m Schedule in ascending order of priority, give priority to allocating resources in URLLC reserved slices, if the reserved slice resources are allocated and there are still D n -d n,m Data packets shorter than a mini-slot will seize resources in the shared slice for transmission;
[0043] In step 3.3, for resource allocation of periodic URLLC traffic, the heuristic algorithm in step 2 is used to calculate the deterministic delay of each URLLC user. For packets of periodic traffic in the buffer area, they are preferentially transmitted according to the pre-configured deterministic delay. If resources are insufficient, they wait in the buffer area.
[0044] Optionally, in one embodiment of the present invention, step 4 further includes:
[0045] Step 4.1: Model the large-scale inter-slice resource allocation process as a partially observable Markov decision process.
[0046] The state, action, and reward functions designed based on the optimization problem of maximizing the eMBB service throughput are as follows:
[0047] Status: Statistical channel status information of all users;
[0048] Action: The ratio of resources allocated to shared slices and URLLC reserved slices;
[0049] Reward: The weighted sum of the eMBB rate, URLLC rate, and the proportion of eMBB preempted resources.
[0050] Optionally, in one embodiment of the present invention, step 5 further includes:
[0051] Step 5.1: Initialize the deep reinforcement learning network weights and algorithm hyperparameters, apply the algorithm to the agent, and make it interact with the wireless communication environment in step 1 for several rounds;
[0052] Step 5.2: Initialize the base station user position and channel state information at the beginning of each interaction round, and design the time step in each interaction round;
[0053] In step 5.3, at each time step, the agent collects user statistical channel state information and inputs it into the deep reinforcement learning network. It then adjusts the slice resource allocation ratio based on the output of the deep reinforcement learning network.
[0054] Step 5.4: After the slice resource allocation ratio is determined, perform intra-slice resource allocation according to step 3. Measure and record the eMBB rate, URLLC rate, and preempted resource ratio. Collect channel state information and calculate the reward function. Cache the original state, action, reward, and next state in the experience buffer. When the buffer has sufficient data, randomly extract data in batches to train the neural network.
[0055] In step 5.5, repeat the above interactive process until the deep reinforcement learning network converges, and save the deep reinforcement learning model and the trained parameter configuration.
[0056] The delay SLA-oriented intelligent wireless access network resource allocation method according to the embodiment of the present invention has the following beneficial effects:
[0057] 1. The present invention transforms the optimization problem of minimizing the sum of deterministic traffic delay and jitter into the problem of minimizing the expected variance of mini-slot URLLC resource requirements, and solves it using a low-complexity heuristic algorithm. The algorithm can effectively reduce delay and jitter while meeting the deterministic traffic delay SLA requirements.
[0058] 2. This invention combines orthogonal and non-orthogonal slicing, leveraging a deep reinforcement learning framework to achieve adaptive adjustment of slice resource configuration. This intelligently and efficiently finds the optimal slice resource allocation strategy in the current network, achieving high eMBB service throughput and slice isolation while ensuring URLLC service latency SLA requirements.
[0059] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0061] Figure 1 A flowchart of a method for allocating resources in an intelligent wireless access network oriented to latency SLA according to an embodiment of the present invention;
[0062] Figure 2 Schematic diagram of a hybrid slicing model according to an embodiment of the present invention;
[0063] Figure 3Performance comparison chart of the heuristic URLLC deterministic delay scheduling method designed for an embodiment of the present invention and a baseline solution;
[0064] Figure 4 Performance comparison chart of the hybrid slicing resource allocation method based on deep reinforcement learning designed for an embodiment of the present invention and traditional orthogonal and non-orthogonal slicing methods. DETAILED DESCRIPTION
[0065] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0066] Figure 1 The present invention provides a flowchart of a method for allocating intelligent wireless access network resources based on latency SLA according to an embodiment of the present invention.
[0067] like Figure 1 As shown, the delay SLA-oriented intelligent wireless access network resource allocation method includes the following steps:
[0068] Step 1: Construct a single-base station downlink transmission OFDMA system model and establish an optimization problem for minimizing the sum of deterministic traffic delay and jitter for intra-slice resource allocation. Furthermore, under the premise of determining the intra-slice resource allocation strategy, establish an optimization problem for maximizing the eMBB service throughput for inter-slice resource allocation.
[0069] In an embodiment of the present invention, step 1 further comprises:
[0070] In step 1.1, for a single-base station downlink transmission scenario, there are two types of users: eMBB users and URLLC users. eMBB users have non-real-time, high-bandwidth requirements, while URLLC users require high reliability and low latency. Network slicing technology is used for resource allocation. Wireless spectrum resources are divided into shared slices and reserved slices for URLLC users.
[0071] Step 1.2: eMBB users use slot-based transmission, with a slot length of 1ms. At the beginning of each slot, the bandwidth of the shared slice is allocated to the eMBB user. e,n (t) indicates the time slot t allocated to eMBB user e n The bandwidth, p e,n (t) is the power allocated to the user by the base station, g e,n (t) is the channel gain from the base station to the user, σ 2 is the noise variance. The noise is additive Gaussian white noise. Then user en The signal-to-noise ratio at time slot t is:
[0072]
[0073] User e n The rate when the resource is not preempted by URLLC users in time slot t is obtained by Shannon's formula:
[0074] R e,n (t) = ω e,n (t)log2(1+Γ e,n (t))
[0075] In step 1.3, URLLC users use mini-slot-based transmission. One slot contains seven mini-slots. At the beginning of each mini-slot, resources in the URLLC reserved slice are allocated according to the URLLC user scheduling priority. When the resources in the reserved slice are insufficient, transmission is performed by preempting the resources allocated to the eMBB user. Since URLLC data in industrial control signals is usually transmitted in short packets, the URLLC user u n The rate uses the short packet transmission rate formula based on finite block length coding:
[0076]
[0077] Among them, ω u,n (τ) is the bandwidth allocated to the user in the mini-time slot τ, Γ u,n (τ) is the signal-to-noise ratio of the mini-slot τ. -1 (·) is the inverse function of Gaussian Q function. ∈ is the decoding error probability. V(Γ u,n (τ)) is the channel dispersion. M is the block length. The channel dispersion is calculated as:
[0078]
[0079] Step 1.4, the delay of URLLC in the wireless access network includes transmission delay, queuing delay and propagation delay. The propagation delay depends on the physical characteristics of the transmission medium, such as transmission speed and distance. The queuing delay depends on the time interval between the arrival of the data packet and the allocated resources. The transmission delay depends on the size of the data packet and the transmission rate. In the present invention, a short packet of URLLC is transmitted within a mini-time slot. Therefore, the transmission delay is the time length of a mini-time slot. Since the propagation delay is short and cannot be optimized through resource allocation, the present invention focuses on the queuing delay of the URLLC data packet. The queuing delay of the mth data packet of the nth URLLC user is denoted as d n,m The maximum queuing delay can be obtained by subtracting the transmission delay and propagation delay from the total delay requirement of the URLLC data packet. The maximum allowed queuing delay of the nth URLLC user is denoted as Dn The arrival of URLLC traffic includes random arrival and periodic arrival.
[0080] In step 1.5, URLLC users preempting resources allocated to eMBB users in a shared slice results in a rate loss for the eMBB users. The proportion of the eMBB user rate loss due to preemption in time slot t is denoted as loss(t).
[0081] Step 1.6: Establish an optimization problem for minimizing the sum of delay and jitter of deterministic traffic. The details are as follows:
[0082]
[0083] Among them, b n,τ is the number of physical resource blocks (RBs) required for URLLC user n to send data packets in mini-time slot τ, i n,τ Indicates whether user n sends a data packet in mini-slot τ. 1 indicates sending, and 0 indicates not sending. B is the total bandwidth. The constraints guarantee latency and reliability for URLLC users, respectively.
[0084] Step 1.7: Establish the optimization problem of maximizing eMBB throughput as follows:
[0085]
[0086] stω e ≤B
[0087] where loss(t) is the eMBB rate loss caused by URLLC preemption in time slot t. The constraint ensures that the resources allocated to eMBB users do not exceed the total number of resources.
[0088] Specifically, in step 1, a single-base station downlink transmission OFDMA system model is first established. This embodiment of the present invention assumes a cell radius of 500 meters, with the base station located at the cell center. There are 40 URLLC users and 2 eMBB users within the base station range, with users randomly distributed within the cell. The shortest distance from a user to the base station is 100 meters. Assume a communication bandwidth of 10 MHz and a subcarrier spacing of 15 kHz. The receiver noise power is -174 dBm / Hz. The base station transmit power is 1 W, with equal power allocated to each time-frequency resource block.
[0089] The channel is assumed to be flat fading. The large time scale is 1 second, and the user's large-scale fading coefficient remains constant over the large time scale. The large-scale fading coefficient includes path loss and shadow fading. The path loss model is PL(d) = 128.1 + 37.6lg(d), expressed in dB. d is the distance from the user to the base station, expressed in kilometers. The shadow fading model follows a lognormal distribution with a standard deviation of 8 dB. Small-scale fading is assumed to follow a Rayleigh fading distribution. The small time scale is 1 millisecond, and the user's small-scale fading coefficient remains constant over the small time scale.
[0090] URLLC users have latency and reliability requirements. Substituting the decoding error probability and signal-to-interference-and-noise ratio into the short packet transmission rate formula for finite block length coding, the URLLC user rate per physical resource block can be obtained. The calculation is as follows:
[0091]
[0092] In this embodiment, the duration of a mini-slot is two OFDM symbols, and a physical resource block contains 12 subcarriers. Therefore, the block length M is 24. The subcarrier spacing is 15 kHz, and the unit RB bandwidth is 180 kHz. The decoding error probability is 0.001 and 0.00001. The packet size s is n The number of RBs required for transmission can be obtained by combining the unit RB transmission rate with the following calculation:
[0093]
[0094] The data packet size is 32 bytes and 48 bytes. τ It is the duration of a mini-slot (1 / 7ms).
[0095] Slice transmission model such as Figure 2 As shown in Figure 2, at the beginning of each large time scale, the resource reservation ratio of shared slices and URLLC slices is determined. Within the small time scale range, resources are allocated to eMBB and URLLC users respectively.
[0096] For periodic URLLC traffic, an optimization problem is established for small-time-scale URLLC user scheduling. The optimization objective is to minimize the sum of the delay jitter of deterministic traffic. User delay jitter is the difference between the maximum delay and the minimum delay. The details are as follows:
[0097]
[0098] The queuing delay of the mth data packet of the nth URLLC user is denoted as d n,m The maximum allowed queuing delay of the nth URLLC user is denoted as D nIn this embodiment, the mini-slot durations are 0, 1, and 3. n,τ is the number of physical resource blocks required for URLLC user n to send data packets in mini-time slot τ, i n,τ Indicates whether user n sends a data packet in mini-slot τ. 1 indicates sending, and 0 indicates not sending. B is the total bandwidth.
[0099] Under the hybrid slice transmission model, an optimization problem is established for inter-slice resource allocation, with the goal of maximizing the throughput of eMBB services while ensuring URLLC service requirements. The details are as follows:
[0100]
[0101] stω e ≤B
[0102] Where loss(t) is the eMBB rate loss caused by URLLC preemption in time slot t. The constraint ensures that the resources allocated to the eMBB user do not exceed the total number of resources. In this embodiment, a convex rate loss function is used. The details are as follows:
[0103]
[0104] Where μ is the ratio of URLLC preempted resources to the total resources allocated by eMBB. The probability of each RB in the shared slice bandwidth being preempted is equal.
[0105] Step 2: The optimization problem of minimizing the sum of the deterministic traffic delay jitter is converted into the problem of minimizing the expected variance of the URLLC resource requirements of each mini-timeslot under deterministic delay, and the heuristic method is used to solve the optimization problem.
[0106] In an embodiment of the present invention, step 2 further comprises:
[0107] In step 2.1, due to the presence of small-scale fading, the resources required for URLLC user transmission will change and be unpredictable. To reduce the risk of insufficient resources to support deterministic delay propagation due to channel variations, the deterministic traffic delay jitter minimization optimization problem proposed in step 1 is transformed into the problem of minimizing the expected variance of URLLC resource requirements for each mini-timeslot under deterministic delay. The details are as follows:
[0108]
[0109] Among them, l n is the deterministic delay of user n. When resources are sufficient, all URLLC data packets are transmitted with this delay. τThe expected value of resources required by each URLLC user in a mini-slot τ for transmission with deterministic delay. The expected value of resources required by each user is obtained from the packet size and statistical channel state information.
[0110] Step 2.2, in, is a constant, so the optimization problem is equivalent to:
[0111]
[0112] Step 2.3, the heuristic algorithm steps designed to solve the above optimization problem are as follows:
[0113] Using the packet size of URLLC periodic traffic and statistical channel status information, the expected value k of the number of RBs required to send each user data packet is calculated. n . And sort them in descending order of expected value.
[0114] Initialize all mini-slot RB usage to 0 and user to 1
[0115] repeat:
[0116] Within the user k delay requirement range, calculate the ∑ corresponding to each propagation delay τ b τ 2 Growth value, and at the same time determine the mini-slot occupied resources b corresponding to each delay τ Whether it exceeds the set threshold. If so, exclude this delay.
[0117] Select ∑ τ b τ 2 The delay corresponding to the minimum growth value is updated τ If all delays do not meet the requirements, the user is set as the first priority user, k = 1.
[0118] Until: the delay of all users is determined.
[0119] Step 3: Set the location distribution and traffic model of URLLC users and eMBB users, and determine the resource allocation plan for the two services at a small time scale.
[0120] Among them, the two services refer to eMBB service and URLLC service.
[0121] In an embodiment of the present invention, step 3 further comprises:
[0122] In step 3.1, eMBB and URLLC users are randomly distributed within the cell, and each adopts an appropriate traffic model. At the beginning of each large time scale, the user's location information and statistical channel state information are updated.
[0123] Step 3.2: At the beginning of each time slot, allocate resources to eMBB users. Allocate resources in the shared slice to specific users using the proportional fair resource allocation algorithm. At the beginning of each mini-time slot, allocate resources to URLLC users. For random traffic packets in the buffer area, follow the D n -d n,m Prioritize the allocation of resources in the URLLC reserved slice. If the reserved slice resources are allocated and there are still D n -d n,m Data packets shorter than a mini-slot will seize resources in the shared slice for transmission.
[0124] In step 3.3, for resource allocation of periodic URLLC traffic, the deterministic delay for each URLLC user is calculated using the low-complexity heuristic algorithm proposed in step 2. For packets in the buffer for periodic traffic, they are prioritized for transmission according to the pre-configured deterministic delay. If resources are insufficient, they are left waiting in the buffer.
[0125] Specifically, eMBB users are uniformly and randomly distributed within a range of 100 to 500 meters from the base station. URLLC users are randomly drawn with equal probability from (0, 0, 1), (0, 0.25, 0.75), (0, 0.5, 0.5), (0.3, 0.4, 0.3), (0, 1, 0), (0.25, 0.75, 0), and (0, 0.75, 0.25). URLLC users are uniformly distributed within the distance intervals of [300, 500], [200, 400], and [100, 300]. At the beginning of each large time scale, the user position is updated. The distance between the user and the base station is added to a uniformly distributed random value in the range [-5, 5], which does not exceed the range [100, 500].
[0126] eMBB user traffic uses a fully buffered model, meaning there is always data in the buffer. URLLC user traffic uses both random and periodic traffic. Random traffic uses a Poisson arrival model, with a 0.1 probability of arrival of each user's data packet per mini-slot. Periodic traffic has six period types: 7, 8, 9, 10, 12, and 14 mini-slots.
[0127] At the beginning of each time slot, resources are allocated to eMBB users. Resources in the shared slice are allocated to specific users using a proportional fair resource allocation algorithm. The specific operations are as follows:
[0128]
[0129] Among them, r i(t) is the rate of the RB allocated to eMBB user i, calculated by the Shannon formula. is the average rate, and the average rate is updated before allocating RBs. After allocating RB to user i, the updated average rate is Until all RBs in the shared slice are allocated.
[0130] At the beginning of each mini-time slot, resources are allocated to URLLC users. For random traffic packets in the buffer area, according to D n -d n,m Schedule the resources in the reserved slice of URLLC first. If the reserved slice resources are allocated and there are still D n -d n,m Data packets shorter than a mini-slot will seize resources in the shared slice for transmission.
[0131] For resource allocation of periodic URLLC traffic, a designed heuristic algorithm is used to calculate the deterministic delay for each URLLC user. For packets in the cache for periodic traffic, they are prioritized for transmission based on the pre-configured deterministic delay. If resources are insufficient, they are left waiting in the cache.
[0132] Step 4: To optimize the eMBB service throughput, within the framework of a deep reinforcement learning network, the cell downlink communication system is considered the environment, the base station is considered the agent, and the large-scale inter-slice resource allocation process is established as a Markov decision process. The state, action, and reward functions of the reinforcement learning are designed to build a deep reinforcement learning model.
[0133] In an embodiment of the present invention, step 4 further includes:
[0134] Step 4.1: At the beginning of each large time scale, the resource ratio allocated to the shared slice and the URLLC reserved slice needs to be determined. The reserved resource ratio not only affects the throughput of the current eMBB user, but also affects the resource allocation strategy and throughput of subsequent eMBB. Therefore, the embodiment models the interaction process as a partially observable Markov decision process (S, A, R, P, γ). Where S and A are the state and action spaces, R(S, A) is the reward function, P is the state transition probability, and γ is the discount factor. The state action value function is:
[0135]
[0136] The Bellman equation can be expressed as:
[0137] Q π (s,a)=E π,p (R(s,a)+γQπ (s′,a′)
[0138] Among them, s′ and a′ can be obtained from P and π respectively.
[0139] The goal of deep reinforcement learning is to find the optimal strategy to maximize the value of states and actions. This example uses DoubleDQN to address the problem of overestimation of Q values in the traditional DQN algorithm. The specific design of the state, action, and reward function is as follows:
[0140] Status: Statistical channel status information of all users, which can be obtained by the base station. In the embodiment of the present invention, it is calculated by power and large-scale fading coefficient:
[0141] s i =p-PL(t)-sf
[0142] Wherein, p is the power per unit RB, which is 10*lg(20)dBm in the embodiment of the present invention, PL(t) is the path loss, and sf is the shadow fading.
[0143] Action: The ratio of resources allocated to shared slices and URLLC reserved slices. In this embodiment of the present invention, the action space size is 20, representing the number of RBs allocated to reserved slices from 0 to 19.
[0144] Reward: The weighted sum of the eMBB rate, URLLC rate, and the proportion of eMBB preempted resources. In this embodiment, the weighting factors for the eMBB rate and URLLC rate are 1, and the weighting factors for the preempted resource proportion are 0, 5, or 25.
[0145] Step 5: Based on the deep reinforcement learning model built in step 4, allocate resources based on the model output on a large time scale and on a small time scale based on the resource allocation scheme in step 3. Through the interaction between the base station and the wireless environment, learn the optimal strategy for slice resource allocation.
[0146] In an embodiment of the present invention, step 5 further includes:
[0147] In step 5.1, the weights of the deep reinforcement learning network and its algorithm hyperparameters are initialized, and the algorithm is applied to the agent, so that it interacts with the wireless communication environment described in step 1 for several rounds.
[0148] Step 5.2: Initialize the base station user position and channel state information at the beginning of each interaction round, and design the time step in each interaction round.
[0149] In step 5.3, at each time step, the agent collects user statistical channel state information and inputs it into the deep reinforcement learning network, and then adjusts the slice resource allocation ratio based on the output of the deep reinforcement learning network.
[0150] In step 5.4, after the slice resource allocation ratio is determined, intra-slice resource allocation is performed according to step 3. The eMBB rate, URLLC rate, and preempted resource ratio are measured and recorded, channel state information is collected, and the reward function is calculated. The original state, action, reward, and next state are cached in the experience buffer. When the buffer has sufficient data, data is randomly batched and used for neural network training.
[0151] In step 5.5, repeat the above interactive process until the deep reinforcement learning algorithm converges. Save the neural network model and the trained parameter configuration.
[0152] This embodiment of the present invention uses a deep reinforcement learning algorithm based on Double DQN for inter-slice resource allocation. The network structure is Linear(42,128), ReLU, Linear(128,64), ReLU, Linear(64,20). The batch size is 128, the experience pool size is 10,000, the number of training interaction rounds is 4,000, and the number of time steps per interaction round is 200. The discount factor is 0.1, the learning rate is 0.0002, and the optimizer is Adam.
[0153] The present invention evaluates the proposed method of intelligent wireless access network resource allocation for delay SLA through simulation experiments. The heuristic algorithm proposed in the present invention is used with the greedy method to allocate URLLC resources, and the actual RB demand variance curve of each mini-time slot is obtained, as shown in Figure 2. Figure 3 As shown. It can be found that this heuristic method can effectively reduce the variance of resource requirements for each mini-time slot. The hybrid slicing resource allocation method based on deep reinforcement learning proposed in this invention is combined with the URLLC resource allocation method for latency SLA, as well as the orthogonal slicing and non-orthogonal slicing methods to allocate resources, and the eMBB throughput curve is obtained, as shown in Figure 4 It can be found that when the URLLC traffic demand is high, the invented method can improve the throughput of eMBB users compared with the traditional method.
[0154] The intelligent wireless access network resource allocation method for latency SLAs, according to an embodiment of the present invention, targets hybrid industrial Internet URLLC and eMBB service scenarios. It provides an inter-slice resource allocation method based on deep reinforcement learning and an intra-slice URLLC resource allocation method based on a heuristic algorithm. These two methods are applied to large and small time scales, respectively. Within the large time scale, the large-scale fading coefficient of the channel remains constant. Within the small time scale, the small-scale fading coefficient remains constant. The present invention first establishes an optimization problem for minimizing the sum of deterministic traffic delay jitter through intra-slice resource allocation and, based on this, maximizing eMBB service throughput through inter-slice resource allocation. Next, the deterministic traffic delay jitter minimization problem is transformed into the problem of minimizing the expected variance of URLLC resource requirements for each mini-slot under deterministic delay, and a heuristic method is designed to solve this problem. Finally, a deep reinforcement learning optimization algorithm is designed to adaptively adjust inter-slice resource allocation based on statistical channel state information to improve eMBB throughput. This method features low complexity, strong adaptability, and excellent performance.
[0155] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0156] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0157] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or N executable instructions for implementing a custom logical function or step of a process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
Claims
1. A method for allocating intelligent wireless access network resources based on latency SLA, characterized in that: The following steps are involved: Step 1: Build a single-base station downlink transmission OFDMA system model and formulate an optimization problem for minimizing the sum of deterministic traffic delay and jitter for intra-slice resource allocation. Also, establish an optimization problem for maximizing eMBB service throughput for inter-slice resource allocation, given a defined intra-slice resource allocation strategy. Step 2: Convert the optimization problem of minimizing the sum of the deterministic traffic delay and jitter into the problem of minimizing the expected variance of the URLLC resource requirements for each mini-timeslot under deterministic delay and solve the optimization problem using a heuristic method; Step 3: Set the location distribution and traffic model of URLLC users and eMBB users, and determine the resource allocation plan for the two services at a small time scale. Step 4: To optimize the eMBB service throughput, within the framework of a deep reinforcement learning network, the cell downlink communication system is considered the environment, the base station is considered the agent, and the large-scale inter-slice resource allocation process is modeled as a Markov decision process. The reinforcement learning state, action, and reward functions are designed to build a deep reinforcement learning model. Step 5: Based on the deep reinforcement learning model built in step 4, allocate resources based on the model output on a large time scale and on a small time scale based on the resource allocation scheme in step 3. Through the interaction between the base station and the wireless environment, learn the optimal strategy for slice resource allocation.
2. The method according to claim 1, characterized in that The step 1 further comprises: Step 1.1: A single-base station downlink transmission scenario includes two types of users: eMBB users and URLLC users. eMBB users have non-real-time, high-bandwidth requirements, while URLLC users require high reliability and low latency. Network slicing technology is used for resource allocation. Wireless spectrum resources are divided into shared slices and reserved slices for URLLC users. In step 1.2, the eMBB user adopts slot-based transmission. The length of a slot is 1ms. At the beginning of each slot, the bandwidth of the shared slice is allocated to the eMBB user. In step 1.3, URLLC users use mini-slot-based transmission. One time slot consists of seven mini-slots. At the beginning of each mini-slot, resources in the URLLC reserved slice are allocated based on the URLLC user scheduling priority. If resources in the reserved slice are insufficient, transmission is performed by preempting resources allocated to eMBB users. Step 1.4: The delay of URLLC in the wireless access network includes transmission delay, queuing delay, and propagation delay. A short URLLC packet is transmitted within a mini-slot, so the transmission delay is the length of a mini-slot. The queuing delay of the mth data packet of the nth URLLC user is d n,m , subtract the transmission delay and propagation delay from the total delay requirement of the URLLC data packet to obtain the maximum queuing delay, and record the maximum allowed queuing delay of the nth URLLC user as D n ,The arrival of URLLC traffic includes random arrival and periodic arrival; In step 1.5, URLLC users preempting resources allocated to eMBB users in a shared slice will cause rate loss for the eMBB users. The rate loss ratio of the eMBB users caused by URLLC preemption in time slot t is recorded as loss(t). Step 1.6: Establish an optimization problem for minimizing the sum of delay and jitter of deterministic traffic. The details are as follows: Among them, b n,τ is the number of physical resource blocks required for URLLC user n to send data packets in mini-time slot τ, i n,τ Indicates whether URLLC user n sends a data packet in mini-timeslot τ. 1 indicates sending, and 0 indicates not sending. B is the total bandwidth. The constraints guarantee the latency and reliability of the URLLC user respectively. Step 1.7: Establish the optimization problem of maximizing eMBB throughput as follows: Among them, R e,n (t) is the eMBB user e n The rate when the URLLC user does not preempt resources in time slot t. loss(t) is the eMBB rate loss caused by URLLC preemption in time slot t. ω e For the bandwidth allocated to eMBB users, the constraints ensure that the resources allocated to eMBB users do not exceed the total resources.
3. The method according to claim 2, characterized in that The step 2 further comprises: Step 2.1: Convert the deterministic traffic delay jitter minimization optimization problem into the problem of minimizing the expected variance of URLLC resource requirements for each mini-timeslot under deterministic delay. The details are as follows: Among them, l n is the deterministic delay of user n. When resources are sufficient, all URLLC data packets are transmitted with this deterministic delay. τ The expected resource value required by URLLC users in mini-slot τ for transmission with deterministic delay. The expected resource value required by each user is obtained from the packet size and statistical channel state information. Step 2.2, in, is a constant, so the optimization problem is equivalent to: Step 2.3, the steps for solving the optimization problem using the heuristic method are as follows: Using the packet size of URLLC periodic traffic and statistical channel status information, the expected value k of the number of RBs required to send each user data packet is calculated. n , and sort them in descending order of expected value; Initialize all mini-slot RB usage to 0 and the user to 1; repeat: Within the user k delay requirement range, calculate the ∑ corresponding to each propagation delay τ b τ 2 Growth value, and at the same time determine the expected value b of resources required for each delay corresponding to the mini-time slot τ τ Whether it exceeds the set threshold. If so, the delay is excluded. Select ∑ τ b τ 2 The delay corresponding to the minimum growth value is updated τ and k; if all delays do not meet the requirements, the user is set as the first priority user, k = 1; Until: the delay of all users is determined.
4. The method according to claim 3, characterized in that The step 3 further comprises: In step 3.1, URLLC users and eMBB users are randomly distributed within the cell. Each user adopts an appropriate traffic model and updates the user's location information and statistical channel state information at the beginning of each large time scale. Step 3.2: At the beginning of each time slot, resources are allocated to eMBB users. The resources in the shared slice are allocated to specific users using a proportional fair resource allocation algorithm. At the beginning of each mini-time slot, resources are allocated to URLLC users. For random traffic packets in the buffer area, the D n -d n,m Schedule in ascending order of priority, give priority to allocating resources in URLLC reserved slices, if the reserved slice resources are allocated and there are still D n -d n,m Data packets shorter than a mini-slot will seize resources in the shared slice for transmission; In step 3.3, for resource allocation of periodic URLLC traffic, the heuristic algorithm in step 2 is used to calculate the deterministic delay of each URLLC user. For packets of periodic traffic in the buffer area, they are preferentially transmitted according to the pre-configured deterministic delay. If resources are insufficient, they wait in the buffer area.
5. The method according to claim 4, characterized in that The step 4 further comprises: Step 4.1, model the large-scale inter-slice resource allocation process as a partially observable Markov decision process; The state, action, and reward functions designed based on the optimization problem of maximizing the eMBB service throughput are as follows: Status: Statistical channel status information of all users; Action: The ratio of resources allocated to shared slices and URLLC reserved slices; Reward: The weighted sum of the eMBB rate, URLLC rate, and the proportion of eMBB preempted resources.
6. The method according to claim 3, characterized in that The step 5 further comprises: Step 5.1: Initialize the deep reinforcement learning network weights and algorithm hyperparameters, apply the algorithm to the agent, and make it interact with the wireless communication environment in step 1 for several rounds; Step 5.2: Initialize the base station user position and channel state information at the beginning of each interaction round, and design the time step in each interaction round; In step 5.3, at each time step, the agent collects user statistical channel state information and inputs it into the deep reinforcement learning network. It then adjusts the slice resource allocation ratio based on the output of the deep reinforcement learning network. Step 5.4: After the slice resource allocation ratio is determined, perform intra-slice resource allocation according to step 3. Measure and record the eMBB rate, URLLC rate, and preempted resource ratio. Collect channel state information and calculate the reward function. Cache the original state, action, reward, and next state in the experience buffer. When the buffer has sufficient data, randomly extract data in batches to train the neural network. In step 5.5, repeat the above interactive process until the deep reinforcement learning network converges, and save the deep reinforcement learning model and the trained parameter configuration.
Citation Information
Patent Citations
Heterogeneous network resource slicing method with eMBB and URLLC mixed service
CN114340017A
Mixed service intelligent resource scheduling method based on reinforcement learning algorithm
CN116234047A