A resource allocation method for service-oriented wireless access network

By using the DDQN algorithm to construct a resource allocation solution in a service-oriented radio access network, the problems of computational complexity and resource cost are solved, efficient bandwidth utilization and low-energy resource allocation are achieved, and diversified service needs can be adapted.

CN116017736BActive Publication Date: 2025-09-23XIDIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211475526.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-09-23
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively solve the resource allocation problem in service-oriented wireless access networks, resulting in high computational complexity, huge computational resource costs, and an inability to meet diverse service requirements.

Method used

A dual deep Q network (DDQN) algorithm is used to build a wireless resource allocation solution. By obtaining slice performance indicators and base station computing resource costs, a system cumulative utility function is constructed. The optimization objective is solved in the DDQN algorithm network, and bandwidth resources are dynamically allocated to meet the QoS requirements of different slices.

Benefits of technology

It achieves the goal of improving bandwidth utilization and reducing computing energy consumption while ensuring the QoS of different services, thus adapting to diversified service needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116017736B_ABST
    Figure CN116017736B_ABST
Patent Text Reader

Abstract

The present invention provides a resource allocation method for a service-based wireless access network. The method obtains performance indicators and slice states of different slices at the current moment; constructs a cumulative utility function for system data transmission based on the combined performance indicators of the different slices and the computing resource costs of the base station; imposes a maximization constraint on the system cumulative utility function to obtain an optimization target; inputs the slice states into a trained DDQN algorithm network so that the trained DDQN algorithm network solves the optimization target, obtains the action with the maximum reward, and the bandwidth allocation strategy corresponding to the action; and allocates bandwidth resources to users within eMBB slices and URLLC slices according to the bandwidth allocation strategy. The method ensures the QoS of different services while also considering computing resource costs. Therefore, the method is highly feasible and can achieve higher bandwidth utilization and lower computing energy consumption while meeting user QoS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network resource allocation, and in particular relates to a resource allocation method for a service-oriented wireless access network. Background Art

[0002] Faced with the rapid growth of data traffic, the increasing diversity of service types, and stringent QoS requirements, research on flexible dynamic resource allocation solutions has become crucial. Therefore, current research on radio access networks (RANs) focuses primarily on dynamic resource allocation. RAN slicing technology is being introduced in service-based RAN systems, and wireless resource allocation solutions are being studied specifically for this scenario to support flexible resource configuration in service-based RANs. Traditional optimization algorithms suffer from the complexity of calculating long-term optimization objectives and the difficulty of solving them.

[0003] Currently, most resource allocation solutions use mathematical methods to find optimal solutions based on optimization objectives and constraints, and are suitable for solving static or transient problems. However, analysis of the optimization objectives reveals that the resource allocation problem in wireless access networks is modeled as a dynamic decision-making problem, which makes the optimization function overly complex and difficult to obtain an optimal solution.

[0004] A conventional approach to analyzing the latency and throughput trade-offs in radio access network (RAN) slices based on dynamic intelligent resource allocation has been proposed to meet the diverse service requirements in smart grid scenarios. This approach builds a model based on the throughput requirements of eMBB slices and the user latency constraints of URLLC slices. This approach maximizes the throughput of eMBB slices while ensuring user latency for URLLC slices. While this approach takes into account the diverse requirements of slices, it also incurs significant computing resource deployment costs.

[0005] Existing technologies propose another efficient RAN slicing solution based on offline reinforcement learning and low-complexity heuristic algorithms. This solution allocates wireless resources to different slices, aiming to maximize resource utilization while meeting the traffic requirements of each RAN slice. The idea behind this solution is to define QoS based on rate. However, in real-world environments, services are diverse, and this solution, which only meets the specific requirements of a single service, is not applicable. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a resource allocation method for a service-oriented radio access network. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0007] The present invention provides a resource allocation method for a service-oriented radio access network, comprising:

[0008] Step 1: Obtain the performance indicators and slice status of different slices at the current moment;

[0009] The slices include eMBB slices and URLLC slices, and each slice has multiple users;

[0010] Step 2: Construct a system cumulative utility function for system data transmission based on the performance indicators of different slices and the computing resource cost of the base station;

[0011] Step 3: Maximize the cumulative utility function of the system to obtain the optimization target;

[0012] Step 4: Input the slice state into the trained DDQN algorithm network, so that the trained DDQN algorithm network solves the optimization objective and obtains the action with the maximum reward and the bandwidth allocation strategy corresponding to the action;

[0013] Step 5: Allocate bandwidth resources to users in the eMBB slice and URLLC slice according to the bandwidth allocation policy.

[0014] The present invention provides a resource allocation method for a service-based wireless access network. The method obtains performance indicators and slice states of different slices at the current moment; constructs a cumulative utility function for system data transmission based on the combined performance indicators of the different slices and the computing resource costs of the base station; imposes a maximization constraint on the system cumulative utility function to obtain an optimization objective; inputs the slice states into a trained DDQN algorithm network so that the trained DDQN algorithm network solves the optimization objective, obtains the action with the maximum reward, and the bandwidth allocation strategy corresponding to the action; and allocates bandwidth resources to users within eMBB slices and URLLC slices according to the bandwidth allocation strategy. The method ensures the QoS of different services while also considering computing resource costs. Therefore, the method is highly feasible and can achieve higher bandwidth utilization and lower computing energy consumption while meeting user QoS.

[0015] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of a resource allocation method for a service-based wireless access network provided by the present invention;

[0017] Figure 2 Schematic diagram of the wireless network slicing model provided by the present invention;

[0018] Figure 3 This is a model diagram of the service-oriented RAN wireless resource allocation solution based on the DDQN algorithm provided by the present invention;

[0019] Figure 4This is a bandwidth utilization comparison chart provided by the present invention;

[0020] Figure 5 This is a comparison chart of computing energy consumption provided by the present invention;

[0021] Figure 6 This is a comparison chart of eMBB slice QoS satisfaction provided by the present invention;

[0022] Figure 7 This is a comparison chart of URLLC slice QoS satisfaction provided by the present invention. DETAILED DESCRIPTION

[0023] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0024] In order to ensure the QoS of each service while minimizing the cost of deploying computing resources, the present invention proposes a wireless resource allocation method for service-oriented RAN. The idea is to define the weighted sum of the rate of eMBB (enhanced Mobile Broadband) services, the delay benefit of URLLC (Ultra-Reliable Low-Latency Communications) services, and the computing resource cost of the base station as a utility function, and construct a network slice utility function maximization problem. The constraints are defined as follows: the wireless resources allocated to users do not exceed the total bandwidth resources, the URLLC slice rate is not less than the rate lower limit, and the eMBB slice processing delay is not greater than the processing delay upper limit. The wireless resource allocation decision process is modeled as a Markov Decision Process (MDP) and solved using a double deep Q network (DDQN) algorithm. Combined with the constraints, the system utility function is transformed as a reward, and limited time-frequency resources are dynamically allocated on a physical network to maximize the cumulative reward.

[0025] The specific process of a resource allocation method for a service-oriented radio access network provided by the present invention is described in detail below.

[0026] like Figure 1 As shown, the present invention provides a resource allocation method for a service-oriented radio access network, including:

[0027] Step 1: Obtain the performance indicators and slice status of different slices at the current moment;

[0028] The slices include eMBB slices and URLLC slices, and each slice has multiple users; the performance indicators include: a variable indicating whether sub-bandwidth i is allocated to slice k at time t in the base station; the base station's transmit power on sub-bandwidth i; the channel gain coefficient of slice k; the size of the first data packet in the queue buffer for storing data packets to be sent in slice k at time t is δ k (); When a data packet arrives and is not transmitted immediately, it waits in the buffer. The time the data packet waits in the buffer Packet encoding time of the same slice

[0029] based on Figure 2 In the system model shown, it is assumed that a user only requests one type of service. The optimization goal is to obtain higher QoS benefits for the slice with the least possible computing resources. Different slices have different QoS requirements. The eMBB slice has higher throughput requirements, with throughput as the QoS indicator; the URLLC slice has higher latency requirements, with processing latency as the QoS indicator. It is assumed that there is bandwidth resource isolation between slices, that is, the bandwidth resources of two slices do not overlap, so there is no interference between slices. Each user only accesses one slice. There are M users in slice k. At the current time t, the bandwidths obtained by the eMBB slice and URLLC slice are b1() and b2() respectively. The bandwidth resources of the slice are evenly distributed to the M users in the slice, so the user bandwidth is:

[0030]

[0031] in, is the bandwidth of user d in slice k at time t.

[0032] Step 2: Construct a cumulative utility function for system data transmission based on the performance indicators of different slices and the computing resource cost of the base station;

[0033] Step 2 includes:

[0034] Step 21: Calculate the data transmission rate θ of slice k at time t based on the performance index k (t);

[0035] Step 22: Based on the first packet size δ k (), calculate the transmission time of the data packet

[0036] Step 23: According to the transmission time The time the data packet waits in the buffer And the encoding time of the data packets of the same slice Calculate the packet processing time τ k();

[0037] Step 24: Obtain the computing resources c(t) of the base station;

[0038] Step 25: According to the data transmission rate θ of slice k at time t k (t), packet processing time τ k () and the computing resources c(t) of the base station, and construct the cumulative utility function of the system transmission data.

[0039] According to Shannon's formula, the data transmission rate θ of slice k at time t k (t) are as follows:

[0040]

[0041] Among them, a i,k () is a variable indicating whether sub-bandwidth i is allocated to slice k at time t in the base station, a i,k (t) = {0, 1}; p i () is the transmission power of the base station in sub-bandwidth i at time t; h k () is the channel gain coefficient of slice k at time t; σ 2 is the power spectral density of additive white Gaussian noise;

[0042] The base station's transmission power is constant after system deployment, with a total power of P, which is evenly distributed to each sub-bandwidth, i.e., p i (t) = / N, substituting into (2), the data transmission rate θ of slice k at time t is k (t) is as follows:

[0043]

[0044] Each time resources are allocated, only one data packet of the user is processed. The rest of the data packets of the users in the slice are waiting in the buffer. The first data packet of the buffer in the queue of data packets to be sent in slice k at time t is of size δ k (), the transmission time of the data packet As follows:

[0045]

[0046] When a data packet arrives and is not transmitted immediately, it waits in the buffer. The waiting time for the data packet in the buffer is The encoding time of the data packets of the same slice is the same, which is Then the data packet processing time τ k () is as follows:

[0047]

[0048] The computing resources considered in the present invention are the computing resources of the BBU software unit for digital processing and encoding and decoding. The specific relationship expression of the computing resources c(t) of the base station is as follows:

[0049]

[0050] Where ξ is the weight coefficient of computing resources, A is the number of antennas, B is the base station bandwidth (in MHz), θ(t) is the base station rate (in Mbps), computing resources c(t) is in Giga Operations Per Second (GOPS), and ζ is the bias coefficient of computing resources.

[0051] The computing resource c(t) of the base station BBU software unit at time t is expressed as follows:

[0052]

[0053] The optimization goal of the present invention is to obtain high QoS for users while reducing computing resource consumption. Therefore, the system utility function U(t) is the weighted sum of the QoS benefits and computing resource costs of the two slices of the base station, as shown in the following formula:

[0054]

[0055]

[0056] Where, q1(t)=1(t)- 1, (t) is the rate benefit function of the eMBB slice, α1 is the throughput gain coefficient, θ1(t) is the actual observed throughput of the eMBB slice data packet at time t, and θ 1,h (t) is the lower limit of the throughput of eMBB slice data packets at time t; q2(t) = 2(τ 2, (t)-2(t)) is the delay gain of the URLLC slice, α2 is the delay gain coefficient, τ2(t) is the actual processing delay of the URLLC slice data packet at time t, τ 2, (t) is the upper limit of the processing delay of the URLLC slice data packet at time t; β is the calculation cost coefficient.

[0057] Step 3: Maximize the cumulative utility function of the system to obtain the optimization target;

[0058] Specifically, step 3 of the present invention includes:

[0059] Step 31: Constraining the system cumulative utility function using constraint conditions;

[0060] The constraints include: Constraint C1, Constraint C2, Constraint C3, and Constraint C4. Constraint C1 ensures that the allocated bandwidth is an integer number of sub-bandwidths. Constraint C2 ensures that the bandwidth allocated to the two slices does not exceed the total bandwidth of the base station. Constraint C3 ensures that the data packet rate at time t of the URLLC slice is not less than the rate lower limit. Constraint C4 ensures that the processing delay of the data packet at time t of the eMBB slice is not greater than the processing delay upper limit.

[0061] Step 32: Maximizing the system cumulative utility function is taken as the optimization objective;

[0062] Among them, the optimization goal is:

[0063]

[0064] Among them, st represents the constraint, b represents the sub-bandwidth, and the n-equal sub-bandwidth is

[0065] Step 4: Input the slice state into the trained DDQN algorithm network, so that the trained DDQN algorithm network solves the optimization objective and obtains the action with the maximum reward and the bandwidth allocation strategy corresponding to the action;

[0066] The training process of the DDQN algorithm network trained by the present invention is as follows:

[0067] Step 51: Initialize the experience pool. Initialize the capacity of the experience pool D to N, initialize the DDQN algorithm network Q and network parameters ω - =;

[0068] Step 52: for the i-th iteration, obtain the data of the first data packet in the slice buffer buffer to form the state s;

[0069] Step 53: Randomly generate a probability p;

[0070] Step 54: If p≤e, randomly select an action a∈A, otherwise input the state s into the DDQN algorithm network Q to obtain action a;

[0071] Step 55: The system executes action a and obtains reward r based on the system's QoS and formula (12);

[0072] Step 56: Get the data of the next data packet from the slice buffer buffer to form state s -

[0073] Step 57: Transform the quadruple data (s,,, - ) is stored in the experience pool D;

[0074] The experience pool contains n samples, each of which is a set of four-tuple data. Each four-tuple data consists of the state at time t, the action at time t, the reward of the DDQN algorithm network Q at time t, and the state at time t+1.

[0075] Step 58: If the experience pool D is full, randomly sample n groups of samples from the experience pool D to update the network parameters;

[0076] Step 59: Calculate the loss function based on the actual action and selected action of the DDQN algorithm network Q

[0077] Step 60: Calculate the gradient of the loss function with respect to the Q network neural network parameters ω, and update the parameters using the Adam optimizer;

[0078] Step 61: i+1 returns to step 52 until i reaches the number of iterations.

[0079] From the analysis of the optimization objectives of the present invention, it can be seen that the resource allocation problem of the wireless access network has been modeled as a dynamic decision-making problem, and the DDQN algorithm is used to construct a wireless resource allocation solution for wireless access network slices.

[0080] Figure 3 This is a diagram of the wireless resource solution model of the adopted DDQN algorithm. The present invention deploys the DDQN algorithm on the service-oriented RAN control plane server, and the DDQN Agent obtains the status information from the periodic statistics of the RAN user plane status information of the SMF as the state of the algorithm. The DDQN Agent then makes a wireless resource allocation decision based on the input state, that is, the action of the algorithm. Next, the resource allocation decision is sent to the PCF for policy configuration of wireless resource allocation, and then sent to the RCF for parsing and execution of the resource policy, so as to implement the application of the wireless resource allocation solution on the RAN user plane. After the user uses the allocated resources to process and transmit data, it enters the next state and obtains the reward corresponding to the action based on the QoS data analyzed by the QAF. The new state is then input into the DDQN algorithm, and so on. When the algorithm can stably obtain a higher reward, the training iteration process is terminated and the final model is obtained, that is, the service-oriented RAN wireless resource allocation solution model based on the DDQN algorithm.

[0081] In the present invention, the state of slice k at the current time t is defined as follows:

[0082]

[0083] Where k = {1, 2}, representing eMBB and URLLC slices respectively; δ k (t) is the size of the first data packet in the buffer of slice k at time t. The first data packet in the buffer is the data packet currently to be sent; τk (t) is the maximum tolerable delay of the first data packet in the buffer of slice k at time t; k (t) is the number of packets to be sent in the buffer of slice k at time t; θ k (t) is the minimum transmission rate of the first data packet in the buffer of slice k at time t; is the waiting time for the first data packet in the buffer of slice k at time t; h k (t) is the channel gain coefficient of slice k at time t; B is the total bandwidth of the base station; N is the number of sub-bandwidths; p is the total power of the base station, and the state is a 6*+3 dimensional vector;

[0084] The action at time t is defined as follows:

[0085] A(t)={b1(t),2(t)}(11)

[0086] Where b1(t) represents the bandwidth of the eMBB slice at time t, and b2(t) represents the bandwidth of the eMBB slice at time t. They can only be composed of one or more sub-bandwidths. The total bandwidth of the slice at time t cannot exceed the total bandwidth of the base station. Constraints C1 and C2 are satisfied by limiting the action space of the DDQN algorithm.

[0087] The reward R(t) obtained by performing the action A(t);

[0088] The optimization goal of the present invention is to maximize the system's cumulative utility function under constraints, while the training goal of the DDQN algorithm is to maximize the cumulative reward function. However, the DDQN algorithm cannot directly handle optimization goals with constraints. Constraints C1 and C2 can be satisfied by restricting the action space, but constraints C3 and C4 cannot be handled in this way. Therefore, the present invention transforms the optimization goal to satisfy the reward. When constraints C3 and C4 are satisfied, the reward is the utility function. When constraints C3 and C4 are not satisfied, the reward is set to 0, as shown in the following formula:

[0089]

[0090] The detailed steps of the DDQN-based wireless resource allocation algorithm of the present invention are as follows:

[0091]

[0092] Step 5: Allocate bandwidth resources to users in the eMBB slice and URLLC slice according to the bandwidth allocation policy.

[0093] Step 41: Input the current slice state into the trained DDQN algorithm network, randomly select an action, and calculate the reward corresponding to the solution of the optimization goal by executing the action;

[0094] Step 42: Select the action with the maximum reward and the bandwidth allocation strategy corresponding to the action.

[0095] In order to reflect the performance of the resource allocation scheme, the present invention selects the delay single performance optimization algorithm and the static uniform allocation algorithm and the wireless resource allocation method proposed by the present invention for comparison. Figure 4 It can be seen that as the packet arrival rate increases, the bandwidth utilization of the three resource allocation schemes decreases, but the bandwidth utilization of the utility optimization allocation scheme based on DDQN decreases more slowly, thus achieving higher bandwidth utilization. Figure 5 It can be seen that as the packet arrival rate increases, the computing energy consumption of the three methods increases. The computing energy consumption of the utility optimization allocation solution based on DDQN in the present invention increases at a slower rate, thereby achieving lower computing resource consumption. The present invention calculates the ratio of the transmission rate of eMBB service data packets exceeding the minimum rate within 1s and the ratio of the processing delay of URLLC service data packets below the maximum tolerance delay, and uses them as the QoS satisfaction of the slice. The results are as follows: Figure 6 and Figure 7 As shown. Figure 6 It can be seen that the eMBB slice rate satisfaction of the utility optimization allocation scheme based on DDQN is greater than that of the two comparison schemes, and the greater the packet arrival rate, the greater the difference. Figure 7 It can be seen that the latency satisfaction of the URLLC slice using the DDQN-based utility optimization allocation solution is slightly lower than that of the two comparison solutions, and the difference does not increase significantly as the packet arrival rate increases. This is because the proposed method comprehensively considers both rate and latency benefits in its objective function. Compared to the other two methods, the proposed method can better cope with traffic fluctuations, maintaining a high latency satisfaction level for both slices.

[0096] The key points of this invention are: introducing slicing technology into service-oriented RAN, designing two types of RAN slices for eMBB services with high-speed requirements and URLLC services with low-latency requirements, jointly optimizing the multi-dimensional performance indicators of slices and base station operating costs, constructing the system cumulative utility function maximization problem, using the DDQN algorithm to solve it, and designing the algorithm logic based on the DDQN solution in detail.

[0097] Compared with the best existing technologies, the advantages of the present invention are that it can jointly consider the computing resource cost while ensuring the QoS of different services, and has high feasibility. It can achieve higher bandwidth utilization and lower computing energy consumption on the basis of meeting the user QoS.

[0098] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0099] Although the present application is described herein with reference to various embodiments, those skilled in the art will be able to understand and implement other variations of the disclosed embodiments in practicing the claimed application by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0100] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A resource allocation method for a service-oriented wireless access network, characterized in that: include: Step 1: Obtain the performance indicators and slice status of different slices at the current moment; The slices include eMBB slices and URLLC slices, and each slice has multiple users; Step 2: Construct a system cumulative utility function for system data transmission based on the performance indicators of different slices and the computing resource cost of the base station; Step 3: Maximize the cumulative utility function of the system to obtain the optimization target; Step 4: Input the slice state into the trained DDQN algorithm network, so that the trained DDQN algorithm network solves the optimization objective and obtains the action with the maximum reward and the bandwidth allocation strategy corresponding to the action; Step 5: Allocate bandwidth resources to users in the eMBB slice and URLLC slice according to the bandwidth allocation policy; Step 2 includes: Step 21: Calculate the data transmission rate θ of slice k at time t based on the performance index k (t); where k = {1, 2}, representing eMBB and URLLC slices, respectively; Step 22: Based on the first packet size δ k (t), calculate the transmission time of the data packet Step 23: According to the transmission time The time the data packet waits in the buffer And the encoding time of the data packets of the same slice Calculate the packet processing time τ k (t); Step 24: Obtain the computing resources c(t) of the base station. The specific relationship expression of the computing resources c(t) of the base station is as follows: Where ξ is the weight coefficient of computing resources, A is the number of antennas, B is the base station bandwidth in MHz, θ(t) is the base station rate in Mbps, computing resources c(t) is in billion operations per second, ζ is the bias coefficient of computing resources; N is the number of sub-bandwidths; The computing resource c(t) of the base station BBU software unit at time t is expressed as follows: Step 25: According to the data transmission rate θ of slice k at time t k (t), packet processing time τ k (t) and the computing resources c(t) of the base station are used to construct the system cumulative utility function of the system transmission data; the system cumulative utility function U(t) is the weighted sum of the QoS benefits and computing resource costs of the two slices of the base station, as shown in the following formula: Where q1(t)=θ1(t)-θ 1,th (t) is the rate benefit function of the eMBB slice, α1 is the throughput gain coefficient, θ1(t) is the actual observed throughput of the eMBB slice data packet at time t, and θ 1,th (t) is the lower limit of the throughput of eMBB slice data packets at time t; q2(t) = (τ 2,th (t)-τ2(t)) is the delay gain of the URLLC slice, α2 is the delay gain coefficient, τ2(t) is the actual processing delay of the URLLC slice data packet at time t, τ 2,th (t) is the upper limit of the processing delay of the URLLC slice data packet at time t; β is the calculation cost coefficient.

2. The resource allocation method for a service-oriented radio access network according to claim 1, wherein: In step 1, at the current time t, the bandwidths obtained by the eMBB slice and the URLLC slice are b1(t) and b2(t), respectively. The bandwidth resources of the slice are evenly distributed to the M users in the slice, and the user bandwidth is: in, is the bandwidth of user d in slice k at time t.

3. The resource allocation method for a service-oriented radio access network according to claim 2, wherein: The performance indicators include: a variable indicating whether sub-bandwidth i is allocated to slice k at time t in the base station; the base station's transmit power on sub-bandwidth i; the channel gain coefficient of slice k; the size of the first data packet in the queue buffer storing data packets to be sent by slice k at time t is δ k (t); When a data packet arrives and is not transmitted immediately, it waits in the buffer. The time the data packet waits in the buffer Packet encoding time of the same slice 4. The resource allocation method for a service-oriented radio access network according to claim 3, wherein: According to Shannon's formula, the data transmission rate θ of slice k at time t k (t) are as follows: Among them, a i,k (t) is a variable indicating whether sub-bandwidth i is allocated to slice k at time t in the base station, a i,k (t) = {0, 1}; p i (t) is the transmission power of the base station in sub-bandwidth i at time t; h k (t) is the channel gain coefficient of slice k at time t; σ 2 is the power spectral density of additive white Gaussian noise; b represents the sub-bandwidth; The base station's transmission power is constant after system deployment, with a total power of P, which is evenly distributed to each sub-bandwidth, i.e., p i (t) = P / N, substituting into (2), the data transmission rate θ of slice k at time t is k (t) is as follows: Each time resources are allocated, only one data packet of the user is processed. The rest of the data packets of the users in the slice are waiting in the buffer. The first data packet of the buffer in the queue of data packets to be sent in slice k at time t is of size δ k (t), the transmission time of the data packet As follows: When a data packet arrives and is not transmitted immediately, it waits in the buffer. The waiting time for the data packet in the buffer is The encoding time of the data packets of the same slice is the same, which is Then the data packet processing time τ k (t) is as follows:

5. The resource allocation method for a service-oriented radio access network according to claim 4, characterized in that: Step 3 includes: Step 31: Constraining the system cumulative utility function using constraint conditions; The constraints include: Constraint C1, Constraint C2, Constraint C3, and Constraint C4. Constraint C1 ensures that the allocated bandwidth is an integer number of sub-bandwidths. Constraint C2 ensures that the bandwidth allocated to the two slices does not exceed the total bandwidth of the base station. Constraint C3 ensures that the data packet rate at time t of the URLLC slice is not less than the rate lower limit. Constraint C4 ensures that the processing delay of the data packet at time t of the eMBB slice is not greater than the processing delay upper limit. Step 32: Maximizing the system cumulative utility function is taken as the optimization objective; Among them, the optimization goal is: Among them, st represents the constraint, and the bandwidth of n equal molecules is 6. The resource allocation method for a service-oriented radio access network according to claim 5, characterized in that: The training process of the trained DDQN algorithm network is as follows: Step 51: Initialize the experience pool. Initialize the capacity of the experience pool D to N, initialize the DDQN algorithm network Q and network parameters ω - =ω; Step 52: for the i-th iteration, obtain the data of the first data packet in the slice buffer buffer to form the state s; Step 53: Randomly generate a probability p; Step 54: If p≤e, randomly select an action a∈A, otherwise input the state s into the DDQN algorithm network Q to obtain action a; Step 55: The system executes action a and obtains reward r based on the system's QoS; Step 56: Get the data of the next data packet from the slice buffer buffer to form state s - Step 57: Transform the four-tuple data (s, a, r, s - ) is stored in the experience pool D; The experience pool contains n samples, each of which is a set of four-tuple data. Each four-tuple data consists of the state at time t, the action at time t, the reward of the DDQN algorithm network Q at time t, and the state at time t+1. Step 58: If the experience pool D is full, randomly sample n groups of samples from the experience pool D to update the network parameters; Step 59: Calculate the loss function based on the actual action and selected action of the DDQN algorithm network Q Step 60: Calculate the gradient of the loss function with respect to the Q network neural network parameters ω, and update the parameters using the Adam optimizer; Step 61: i+1 returns to step 52 until i reaches the number of iterations.

7. The resource allocation method for a service-oriented radio access network according to claim 6, characterized in that: The state of slice k at the current time t is defined as follows: Among them, τ k (t) is the maximum tolerable delay of the first data packet in the buffer of slice k at time t; k (t) is the number of packets to be sent in the buffer of slice k at time t; θ k (t) is the minimum transmission rate of the first data packet in the buffer of slice k at time t; is the waiting time for the first data packet in the buffer of slice k at time t; The action at time t is defined as follows: A(t)={b1(t),b2(t)}(11) Where b1(t) represents the bandwidth of the eMBB slice at time t, and b2(t) represents the bandwidth of the URLLC slice at time t. Both can only consist of one or more sub-bandwidths, and the total bandwidth of the slices at time t cannot exceed the total bandwidth of the base station. Constraints C1 and C2 are satisfied by limiting the action space of the DDQN algorithm. The reward R(t) obtained by performing the action A(t); Among them, when the constraints C3 and C4 are met, the reward is the utility function, and when the constraints C3 and C4 are not met, the reward is set to 0, as shown in the following formula:

8. The resource allocation method for a service-oriented radio access network according to claim 7, characterized in that: Step 4 includes: Step 41: Input the current slice state into the trained DDQN algorithm network, randomly select an action, and calculate the reward corresponding to the solution of the optimization goal by executing the action; Step 42: Select the action with the maximum reward and the bandwidth allocation strategy corresponding to the action.

Citation Information

Patent Citations

  • Self-adaptive virtual resource allocation method of network slice based on NOMA

    CN107682135A

  • D2D communication network slice allocation method based on deep reinforcement learning

    CN113163451A