A 5G / 5G-A network slice resource allocation method supporting coexistence of high-reliability low-latency services and enhanced mobile broadband services

By using a hierarchical PPO algorithm and transfer learning-assisted methods, physical resource blocks are allocated to URLLC and eMBB nodes, solving the complexity of resource allocation in the coexistence scenario of URLLC and eMBB services and improving resource utilization.

CN120499853BActive Publication Date: 2026-08-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510789994.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2026-08-25
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In industrial 5G/5G-A networks where URLLC and eMBB services coexist, how can we improve resource utilization and solve the complexity and multi-scale allocation problems of resource allocation while ensuring latency and reliability constraints?

Method used

The hierarchical proximal policy optimization (PPO) algorithm, combined with transfer learning and action space truncation mechanism, is used to allocate physical resource blocks to eMBB nodes and URLLC nodes. The effective action space is learned through optimization algorithm to improve resource utilization.

Benefits of technology

While meeting the quality of service requirements of each network slice, the usage of physical resource blocks is effectively minimized, resource utilization is improved, and efficient allocation and utilization of resources are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499853B_ABST
    Figure CN120499853B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of 5G / 5G-A network slice resource allocation methods supporting high reliability low latency service and enhanced mobile bandwidth service coexistence, belong to mobile communication field.It includes: obtaining URLLC and eMBB service coexistence 5G / 5G-A network parameter information;Optimization target and constraint condition of minimum physical resource block usage are constructed;According to system model and optimization target, establish the hierarchical PPO model assisted by transfer learning, and design state space, action space, reward function and action space truncation mechanism;Pre-training and the parameters of the pre-trained resource allocation model are migrated to hierarchical PPO model;Hierarchical PPO model is trained, and network parameters are updated according to loss function until model converges, obtain the physical resource block allocation method for URLLC and eMBB service coexistence.The present application minimizes physical resource block usage while guaranteeing slice service quality, improves resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mobile communications and relates to a 5G / 5G-A network slicing resource allocation method that supports the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services. Background Technology

[0002] 5G mobile communication technology boasts significant advantages such as low latency, high bandwidth, and massive connectivity, primarily supporting three typical application scenarios: Ultra Reliable and Low Latency Communication (URLLC), Enhanced Mobile Broadband (eMBB), and Massive Machine-Type Communication (mMTC). As an evolution of 5G, 5G-A further enhances network performance indicators and meets intelligent requirements. With the deepening integration of 5G / 5G-A in industrial networks, ensuring data transmission latency and reliability in industrial networks where URLLC and eMBB services coexist is a crucial challenge for industrial network development.

[0003] In industrial 5G / 5G-A networks, the coexistence model of URLLC and eMBB services is a commonly used downlink system architecture for radio access networks. Network slices are deployed in industrial 5G / 5G-A systems to construct multiple isolated virtual private networks (VPNs), and the quality of service (QoS) of services in each slice is guaranteed through dynamic allocation of wireless communication resources. Compared with resource allocation problems in traditional industrial networks, resource allocation in the coexistence scenario of URLLC and eMBB services is more challenging. eMBB services allocate resources in each time slot, while to ensure low latency for URLLC services, resource preemption occurs in micro-time slots. This multi-scale allocation makes the problem more complex. Furthermore, strict reliability constraints must be considered, further exacerbating the complexity. Therefore, how to ensure latency and reliability constraints and improve resource utilization in the coexistence scenario of URLLC and eMBB services has become a significant challenge. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a 5G / 5G-A network slice resource allocation method that supports the coexistence of high-reliability, low-latency services and enhanced mobile bandwidth services. In the downlink scenario of the radio access network, based on the hierarchical proximal policy optimization (PPO) algorithm, physical resource blocks are allocated to eMBB nodes and URLLC nodes to minimize the amount of physical resource blocks used while ensuring slice service quality, thereby improving resource utilization. Furthermore, an action space truncation mechanism is added to the hierarchical PPO method to ensure that the PPO method can learn within the effective action space by constraining the range of selectable action spaces. Further, to optimize the learning efficiency of the algorithm, transfer learning is introduced in combination with the hierarchical PPO method to form a transfer learning-assisted hierarchical PPO method.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A method for allocating 5G / 5G-A network slice resources that supports the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services, the method comprising the following steps:

[0007] S1. Obtain network parameter information for the coexistence of URLLC and eMBB services in 5G / 5G-A networks;

[0008] S2. Construct the optimization objective and constraints for minimizing the usage of physical resource blocks;

[0009] S3. Based on the system model and optimization objectives, establish a hierarchical PPO model assisted by transfer learning, and design the state space, action space, reward function, and action space truncation mechanism.

[0010] S4. Perform pre-training and transfer the parameters of the pre-trained resource allocation model to the hierarchical PPO model;

[0011] S5. Train the hierarchical PPO model, update the network parameters according to the loss function until the model converges, and obtain the physical resource block allocation method for the coexistence of URLLC and eMBB services.

[0012] Furthermore, in S1, a downlink scenario for the radio access network where URLLC and eMBB services coexist is established, wherein the transmission reliability of the network slice for URLLC services is... Maximum acceptable latency is Data is transmitted immediately upon arrival; the transmission reliability of network slices for eMBB services is... Maximum acceptable latency is

[0013] Within the coverage area of ​​the base station, URLLC services include N uThere are N nodes, and the eMBB service includes N. e There are 10 nodes; and within each time slot, the node location changes randomly within the base station's range; the channel gain between the base station and the nodes follows the path loss model and Rayleigh distribution, and varies between each time slot.

[0014] The base station allocates a buffer queue to each node to buffer data waiting to be transmitted. The queue update process is represented as follows:

[0015] U(t+1)=max(U(t)+α(t)-rτ-β(t),0)

[0016] Where U(t) and U(t+1) represent the queue backlog at times t and t+1, respectively, α(t) and β(t) represent the received data and the discarded data, respectively, and r and τ are defined as the transmission rate and transmission time.

[0017] Furthermore, in S2, an optimization objective is constructed. Under the long-term average condition of T→∞, the optimization objective is to minimize the usage of physical resource blocks, expressed as:

[0018]

[0019] in, These are the number of physical resource blocks allocated to eMBB node i and URLLC node j, respectively.

[0020] The constraints on the above optimization objectives include:

[0021] (1) The number of allocated physical resource blocks cannot exceed the total number available:

[0022]

[0023] In the formula, K is the total number of physical resource blocks of the base station;

[0024] (2) Each physical resource block can only be allocated to one node within a unit of time:

[0025]

[0026] in, and All are indicator functions, if or This indicates that physical resource block k has been allocated to the corresponding node; otherwise, it indicates that physical resource block k has not been allocated.

[0027] (3) URLLC slicing service quality constraints: the latency of any traffic shall not exceed the maximum acceptable latency of URLLC.

[0028]

[0029] in, This is the latency of URLLC node j;

[0030] (4) eMBB slice service quality constraints, exceeding the maximum acceptable latency of eMBB The proportion of traffic does not exceed

[0031]

[0032] in, It is the latency of eMBB node i. This indicates that the latency exceeds the maximum acceptable latency of eMBB. The proportion of traffic.

[0033] Furthermore, in S3, a hierarchical PPO model is established based on the Actor-Critic reinforcement learning method, including the eMBB resource allocation model PPO. e URLLC resource allocation model PPO u and pre-trained resource allocation model PPO pre eMBB slices allocate resources in each time slot, while URLLC slices preempt resources in each micro-time slot; PPO e Based on the system status, physical resource blocks are allocated to eMBB nodes in each time slot, while PPO u In each micro-timeslot, URLLC nodes are controlled to preempt physical resource blocks using a perforation method based on traffic and physical resource block usage.

[0034] Simultaneously determine PPO pre and PPO e state space PPO u state space Determine PPO pre and PPO e reward function PPO u reward function and PPO pre PPO e and PPO u Action space And determine the truncated action space based on the action space truncation mechanism.

[0035] Furthermore, regarding PPO pre and PPO e The system state space is defined as:

[0036]

[0037] Among them, Q e D e H e α e α u These represent queue backlog, delay, channel gain, eMBB traffic set, and URLLC traffic set, respectively.

[0038] For PPO u The system state space is defined as:

[0039]

[0040] Among them, H u M and M represent the channel gain and the mask of the physical resource block used by eMBB, respectively.

[0041] Furthermore, regarding PPO pre and PPO e The reward function is defined as:

[0042]

[0043] In the formula, K e (t) represents the number of material resource blocks allocated to the eMBB slice, L i (t) represents the amount of data discarded by eMBB node i, μ e These are the corresponding weighting coefficients;

[0044] For PPO u The reward function is defined as:

[0045]

[0046] In the formula, K u (t) represents the number of physical resource blocks allocated to the URLLC slice, L j (t) represents the amount of data discarded by URLLC node j, μ u These are the corresponding weighting coefficients.

[0047] Furthermore, regarding PPO pre PPO e and PPO u The action space is defined as:

[0048]

[0049] Among them, each action a i Both are defined as physical resource block allocation matrices, where rows represent nodes and columns represent physical resource blocks. A matrix element of 1 indicates that a physical resource block is allocated to the node in that row, and otherwise it indicates that it is not allocated.

[0050] Action selection is performed using an action space truncation mechanism, and sub-action spaces are divided according to service quality and traffic conditions:

[0051]

[0052] Wherein, V1 represents traffic that does not meet the quality of service, and V2 represents zero received traffic and zero backlogged traffic; the sub-action spaces are the set of actions with at least one physical resource block, the set of actions with zero allocated physical resource blocks, and the original action space, respectively.

[0053] We obtain the sub-action space corresponding to each node, and take the intersection to obtain the truncated action space:

[0054]

[0055] in, This represents the nth sub-action space, where N is the total number of sub-action spaces.

[0056] Furthermore, in S4, in PPO pre During the model's pre-training phase, resource allocation optimization is performed only for eMBB services, ignoring URLLC services, and PPO is optimized solely based on eMBB services. pre Train until PPO pre convergence;

[0057] Through transfer learning, PPO pre The parameters are migrated to PPO according to the ratio χ. e and PPO u Among them, PPO e The target task is consistent with the pre-training target, and the network structure is the same; the transfer PPO is... pre All network parameters to PPO e PPO u With PPO pre Different network structures, migration of PPO pre Hidden layer parameters to PPO u .

[0058] Furthermore, in S5, in PPO pre PPO e and PPO u In this context, the loss function is defined as:

[0059]

[0060] In the formula, ω1 and ω2 represent weights, ∈ represents the clipping ratio, clip(·) is the clipping function, and V θ(·) represents the current Critic network output value function;

[0061] r t (θ) represents the policy update ratio, which is expressed as:

[0062]

[0063] Where π θ This is the current strategy. This is the strategy before the update, a t and s t These represent the action and state at time t, respectively.

[0064] The dominance function is defined as follows:

[0065]

[0066] Where γ is the discount factor and λ is a parameter that controls the change in bias. It is the timing difference error, defined as:

[0067]

[0068] Where V represents the value function during sampling;

[0069] It is the target reward, defined as:

[0070]

[0071] It is the policy entropy term, and the formula is:

[0072]

[0073] In the formula, V(s) t ) represents a value function, s t State value; π θ (·) represents the probability vector of each action in the discrete action space, and is expressed as:

[0074]

[0075] Where f θ (·) represents the output of the policy;

[0076] The optimal action is obtained by analyzing the policy output, and the action with the highest probability is selected as the optimal action.

[0077] a * =argmax a π θ (a|s t )

[0078] Use gradient ascent to update policy parameters:

[0079]

[0080] in It is the learning rate;

[0081] Under the updated policy parameters, perform several business transmissions to obtain training results, determine whether the model has converged, and if the model has not converged, perform the next training and update the parameters until the model converges.

[0082] The beneficial effects of this invention are as follows:

[0083] 1) In 5G / 5G-A networks, when URLLC and eMBB services coexist, reasonable resource allocation is crucial. This invention proposes an innovative network slice resource allocation method based on a hierarchical PPO algorithm. Specifically, at the time slot scale, physical resource blocks are precisely allocated to eMBB nodes to meet their relatively stable service demands. At the micro-time slot level, URLLC nodes are guided to preemptively acquire physical resource blocks through puncturing, as URLLC services have extremely high latency requirements and need to acquire resources quickly. Through this refined resource allocation method, the usage of physical resource blocks can be effectively minimized while fully meeting the service quality requirements of each network slice, thereby significantly improving the resource utilization of the entire system and achieving efficient resource allocation and utilization.

[0084] 2) To further improve the performance of the hierarchical PPO algorithm in resource allocation, this invention employs transfer learning to assist the hierarchical PPO algorithm. Transfer learning can transfer existing knowledge to the current task, effectively improving the learning efficiency of the algorithm and accelerating the convergence process. Simultaneously, an action space truncation mechanism is introduced to filter the actions generated by the algorithm, eliminating invalid or inefficient actions and reconstructing a new action space. This mechanism ensures that the PPO algorithm focuses on effective actions, avoiding wasting computational resources on invalid actions, thereby helping to better guarantee the quality of service of each network slice.

[0085] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0086] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0087] Figure 1 This is a schematic diagram of the overall process of the 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to an embodiment of the present invention.

[0088] Figure 2 This is a schematic diagram of the PPO algorithm framework according to an embodiment of the present invention;

[0089] Figure 3 This is a flowchart illustrating the 5G / 5G-A network slice management method supporting the coexistence of URLLC and eMBB services according to an embodiment of the present invention. Detailed Implementation

[0090] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0091] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0092] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0093] Please see Figures 1-3 This is a 5G / 5G-A network slicing resource allocation method that supports the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services.

[0094] In this embodiment, for scenarios involving URLLC and eMBB slices, a hierarchical PPO structure is used to decouple the allocation and combination of eMBB nodes and URLLC nodes, reducing complexity. Physical resource blocks are allocated to eMBB nodes in time slots, and URLLC nodes preempt physical resource blocks in micro-time slots, minimizing physical resource block usage while ensuring the quality of service for slices. Simultaneously, an action space truncation mechanism is proposed, dividing the action space into sub-action spaces and forming new action spaces based on the state of each node, ensuring the model explores within the effective action range. To improve learning efficiency, transfer learning is used to transfer the parameters of the pre-trained resource allocation model to the hierarchical PPO model, resulting in a transfer learning-assisted hierarchical PPO model. Figure 1 As shown, the overall process of a 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to the present invention is as follows:

[0095] S1. Obtain network parameter information for the coexistence of URLLC and eMBB services in 5G / 5G-A networks;

[0096] S2. Construct the optimization objective and constraints for minimizing the usage of physical resource blocks;

[0097] S3. Based on the system model and optimization objectives, establish a hierarchical PPO model assisted by transfer learning, and design the state space, action space, reward function, and action space truncation mechanism.

[0098] S4. Perform pre-training and transfer the parameters of the pre-trained resource allocation model to the hierarchical PPO model;

[0099] S5. Train the hierarchical PPO model, update the network parameters according to the loss function until the model converges, and obtain the physical resource block allocation method for the coexistence of URLLC and eMBB services.

[0100] like Figure 2 The flowchart of the PPO algorithm in this embodiment is shown. The PPO model includes an Actor network and a Critic network, which output action distribution and policy value function estimates respectively during interaction with the environment. State, action, reward, log probability of the action, value function estimate, and completion flag are cached in an experience buffer for training. Mini-batches are sampled from the experience buffer and input into the Actor and Critic networks for training, and the policy is updated using a loss function. Specifically, as... Figure 3 As shown, it includes the following steps in detail:

[0101] V1: Algorithm begins;

[0102] V2: Obtain parameter information from the system model, and simultaneously construct the optimization objective and constraints for minimizing the usage of physical resource blocks;

[0103] Specifically, this involves obtaining 5G / 5G-A network parameter information for the coexistence of URLLC and eMBB services. Consider a radio access network downlink scenario where URLLC and eMBB services coexist, where the reliability of the URLLC slice is set to... Maximum acceptable latency is The requirement is for data to be transmitted immediately upon arrival, while the eMBB slice is set to have an acceptable maximum latency of [missing value]. Reliability is defined as The base station contains K physical resource blocks, each corresponding to a bandwidth of BkHz, with an index defined as k∈{1,2,...,K}. The base station's coverage area is assumed to be a radius of 200 meters, where the eMBB service includes N... e There are N nodes, and the URLLC service includes N. u There are several nodes, and within each time slot, the node locations change randomly within the base station's range. Furthermore, the channel gain between the base station and the nodes follows a path loss model and a Rayleigh distribution, varying between each time slot.

[0104] The base station allocates a buffer queue to each node to buffer data waiting to be transmitted. Updating the queue can be represented as:

[0105] U(t+1)=max(U(t)+α(t)-rτ-β(t),0),

[0106] Where U(t) represents queue backlog, α(t) and β(t) represent received data and discarded data, respectively, and r and τ are defined as transmission rate and transmission time.

[0107] The optimization objective is to minimize the physical resource block usage under the long-term average condition as T→∞, expressed as:

[0108]

[0109] in These are the number of physical resource blocks allocated to eMBB node i and URLLC node j, respectively.

[0110] The constraints are as follows:

[0111] The number of allocated physical resource blocks cannot exceed the total number available.

[0112]

[0113] Each physical resource block can only be allocated to one node within a unit of time, either a time slot or a micro-time slot:

[0114]

[0115] in and All are indicator functions, if and This indicates that physical resource block k has been allocated to the corresponding node; otherwise, it indicates that physical resource block k has not been allocated.

[0116] eMBB slice quality of service constraints, indicating that exceeding The proportion of traffic does not exceed

[0117]

[0118] in It is the latency of eMBB node i.

[0119] URLLC slicing service quality constraints mean that no traffic can exceed [a certain limit].

[0120]

[0121] in This is the latency of URLLC node j.

[0122] V3: Building PPOs using a hierarchical PPO algorithm assisted by transfer learning. pre PPO e and PPO u ;

[0123] PPO is a reinforcement learning method based on the Actor-Critic network. The Actor network outputs the action distribution, and the Critic network evaluates the value function of the current state. A hierarchical PPO model with transfer learning assistance is established, which includes PPO... e PPO u and PPO pre These correspond to eMBB resource allocation, URLLC resource allocation, and pre-training, respectively.

[0124] eMBB slices are allocated resources in each time slot, while URLLC slices are preempted for resources in each micro-time slot. PPO e Based on the system status, physical resource blocks are allocated to eMBB nodes in each time slot, while PPO u In each micro-timeslot, the URLLC node is controlled to preempt physical resource blocks using a perforation method based on traffic and physical resource block usage.

[0125] V4: Initializes the system's state space, action space, reward function, and action space truncation mechanism;

[0126] For PPO pre and PPO e The system state is defined as a set:

[0127]

[0128] Q e D e H e α e α u These represent the sets of traffic for queue backlog, delay, channel gain, eMBB, and URLLC, respectively.

[0129] For PPO u The system state is defined as a set:

[0130]

[0131] Where H u M and M represent the channel gain and the mask of the physical resource block used by eMBB, respectively.

[0132] For PPO pre and PPO e The reward function is defined as:

[0133]

[0134] At the same time, for PPO u The reward function is defined as:

[0135]

[0136] Where K e (t) and K u (t) represents the number of physical resource blocks allocated to MBB and URLLC slices, respectively, L i (t) and L j (t) represents the amount of data discarded by eMBB node i and URLLC node j, respectively, μ e and μ u These are all weights.

[0137] For PPO pre PPO e and PPO u The action space can be defined as:

[0138]

[0139] Each action a iBoth are defined as physical resource block allocation matrices, where rows represent nodes and columns represent physical resource blocks. A matrix element of 1 indicates that a physical resource block has been allocated to the node in that row, otherwise it indicates that it has not been allocated.

[0140] To ensure PPO pre PPO e and PPO u In the action selection process, we focus on effective actions and propose an action space truncation mechanism to divide the sub-action space according to service quality and traffic conditions:

[0141]

[0142] Where V1 represents traffic that does not meet the quality of service, and V2 represents zero received traffic and zero backlogged traffic. The sub-action spaces are the set of actions with at least one physical resource block, the set of actions with zero allocated physical resource blocks, and the original action space.

[0143] We obtain the sub-action space corresponding to each node, and take the intersection to obtain the truncated action space:

[0144]

[0145] V5: Pre-trained PPO pre During the pre-training phase, resource allocation optimization is performed only for eMBB services to simplify the learning task and improve the stability of the initial strategy. (PPO) pre During pre-training, only the eMBB resource allocation problem is considered.

[0146] V6: Transfers the parameters of the pre-trained resource allocation model to the hierarchical PPO model; after pre-training, the PPO model is transferred through transfer learning. pre The parameters are migrated to PPO according to the ratio χ. e and PPO u Because of PPO e The target task is consistent with the pre-training objective, and the network structure is the same, so all network parameters are transferred. PPO... u With PPO pre Different network structures require different hidden layer parameters.

[0147] V7: PPO e Perform eMBB node resource allocation; PPO e Allocate resources to eMBB nodes in the time slot and cache state, action, reward, log probability of action, value function estimate and completion flag.

[0148] V8: PPO u Perform URLLQ node resource allocation; PPO uAllocate resources to URLLC nodes in micro-slots and cache state, action, reward, log probability of action, value function estimate and completion flag.

[0149] V9: Policy Update: Sample mini-batch from the buffer for training, and update the policy through backpropagation based on the loss function;

[0150] In PPO pre PPO e and PPO u In this context, the loss function is pruned to limit the magnitude of policy updates, ensuring that the updated policy does not deviate too far from the old policy. It is defined as follows:

[0151]

[0152] in Let r be the expected value, ∈ be the cropping ratio, and r be the expected value. t (θ) is the policy update ratio, defined as:

[0153]

[0154] Where π θ This is the current strategy. This is the strategy before the update, a t and s t These represent the action and state at time t, respectively, and the advantage function is defined as:

[0155]

[0156] Where γ is the discount factor and λ is a parameter that controls the change in bias. It is the timing difference error, defined as:

[0157]

[0158] Where V represents the value function during sampling.

[0159] Furthermore, in the loss function, ω1 and ω2 represent weights, and V θ This represents the function that represents the current output value of the Critic network. It is the target reward, defined as:

[0160]

[0161] and It is the policy entropy term, and the formula is:

[0162]

[0163] Physical resource block allocation is a discrete action space. Softmax is used to obtain the probability of each action, and the output is a probability vector:

[0164] π θ (a t |s t ) = softmax(f θ (s t ))a t ,

[0165] Where f θ (·) represents the output of the strategy.

[0166] The optimal action is obtained by analyzing the policy output, and the action with the highest probability is selected as the optimal action.

[0167] a * =argmax a π θ (a|s t ).

[0168] Use gradient ascent to update policy parameters:

[0169]

[0170] in It is the learning rate.

[0171] Under the updated policy parameters, several service transmissions are performed to obtain training results and determine whether the model has converged. If the model has not converged, the next training iteration is performed, and the parameters are updated until the model converges. Specifically, during the training process, the number of service transmissions under each policy parameter needs to be flexibly adjusted based on the specific performance feedback. If the number of service transmissions under each policy parameter is small, the training is fast, but noise may be introduced; if the number of service transmissions under each policy parameter is large, the training is slow, but it can explore more extensively the characteristics of the system environment.

[0172] V10: Determine if the system has converged. If it has converged, execute V11; otherwise, execute V7.

[0173] V11: Algorithm ends.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services, characterized in that: The method includes the following steps: S1. Obtain network parameter information for the coexistence of URLLC and eMBB services in 5G / 5G-A networks; S2. Construct the optimization objective and constraints for minimizing the usage of physical resource blocks; S3. Based on the system model and optimization objectives, establish a hierarchical PPO model assisted by transfer learning, and design the state space, action space, reward function, and action space truncation mechanism. S4. Perform pre-training and transfer the parameters of the pre-trained resource allocation model to the hierarchical PPO model; S5. Train the hierarchical PPO model, update the network parameters according to the loss function until the model converges, and obtain the physical resource block allocation method for the coexistence of URLLC and eMBB services.

2. The 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 1, characterized in that: In S1, a downlink scenario for the radio access network where URLLC and eMBB services coexist is established. The transmission reliability of the network slice for the URLLC service is... The maximum acceptable latency is Data is transmitted immediately upon arrival; the transmission reliability of network slices for eMBB services is... The maximum acceptable latency is ; Within the coverage area of ​​the base station, URLLC services include Each node, the eMBB service includes There are 10 nodes; and within each time slot, the node location changes randomly within the base station's range; the channel gain between the base station and the nodes follows the path loss model and Rayleigh distribution, and varies between each time slot. The base station allocates a buffer queue to each node to buffer data waiting to be transmitted. The queue update process is represented as follows: in, and Represent Queue backlog at any time , These represent receiving data and discarding data, respectively. and It is defined as transmission rate and transmission time.

3. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 2, characterized in that: In S2, an optimization objective is constructed. Under long-term average conditions, with the optimization objective of minimizing physical resource block usage, the expression is: in, , These are assigned to eMBB nodes. and URLLC nodes The number of physical resource blocks; The constraints on the above optimization objectives include: (1) The number of allocated physical resource blocks cannot exceed the total number available: In the formula, The total number of physical resource blocks for the base station; (2) Each physical resource block can only be allocated to one node within a unit of time: in, and All are indicator functions, if or , representing physical resource blocks It is assigned to the corresponding node; otherwise, it represents a physical resource block. Not assigned; (3) URLLC slicing service quality constraints: the latency of any traffic shall not exceed the maximum acceptable latency of URLLC. : , in, It is a URLLC node The time delay; (4) eMBB slice service quality constraints, exceeding the maximum acceptable latency of eMBB The proportion of traffic does not exceed : , in, It is an eMBB node The time delay, This indicates that the latency exceeds the maximum acceptable latency of eMBB. The proportion of traffic.

4. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 3, characterized in that: In S3, a hierarchical PPO model is built based on the Actor-Critic reinforcement learning method, including an eMBB resource allocation model. URLLC resource allocation model and pre-trained resource allocation models eMBB slices allocate resources in each time slot, while URLLC slices preempt resources in each micro-time slot. Based on the system status, physical resource blocks are allocated to eMBB nodes in each time slot, while In each micro-timeslot, URLLC nodes are controlled to preempt physical resource blocks using perforation based on traffic and physical resource block usage. Simultaneously determine and state space , state space ,Sure and reward function , reward function ,as well as , and Action space And determine the truncated action space according to the action space truncation mechanism. .

5. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 3, characterized in that: for and The system state space is defined as: in, , , , , These represent queue backlog, delay, channel gain, eMBB traffic set, and URLLC traffic set, respectively. for The system state space is defined as: in, , These represent the masks for channel gain and the physical resource blocks used by eMBB, respectively.

6. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 4, characterized in that: for and The reward function is defined as: In the formula, This indicates the number of physical resource blocks allocated to an eMBB slice. Represents eMBB node The amount of data discarded These are the corresponding weighting coefficients; for The reward function is defined as: In the formula, Indicates the number of physical resource blocks allocated to a URLLC slice. Represents URLLC node The amount of data discarded These are the corresponding weighting coefficients.

7. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 4, characterized in that: for , and The action space is defined as: , Each action Both are defined as physical resource block allocation matrices, where rows represent nodes and columns represent physical resource blocks. A matrix element of 1 indicates that a physical resource block is allocated to the node in that row, and otherwise it indicates that it is not allocated. Action selection is performed using an action space truncation mechanism, and sub-action spaces are divided according to service quality and traffic conditions: in, This indicates that the service quality requirement has not been met, but there is still data usage. This indicates that the received traffic and backlog traffic are 0; the sub-action spaces are the set of actions with at least 1 physical resource block, the set of actions with 0 allocated physical resource blocks, and the original action space, respectively. We obtain the sub-action space corresponding to each node, and take the intersection to obtain the truncated action space: in, Indicates the division of the first Physical movement space The number of all sub-action spaces in the partition.

8. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 4, characterized in that: In S4, in During the model's pre-training phase, resource allocation optimization is performed solely for eMBB services, ignoring URLLC services. Conduct training until convergence; Through transfer learning, The parameters, according to the proportion Migrate to and ,in, The target task is consistent with the pre-training target, and the network structure is the same. All network parameters migrated to ; and Different network structures, migration Hidden layer parameters to .

9. A 5G / 5G-A network slicing resource allocation method supporting the coexistence of high-reliability low-latency services and enhanced mobile bandwidth services according to claim 8, characterized in that: In S5, , and In this context, the loss function is defined as: In the formula, and Represents weight, Indicates the cutting ratio. For the clipping function, This function represents the current output value of the Critic network. The policy update ratio is represented as follows: in This is the current strategy. It's the strategy before the update. and They are Actions and states at all times; The dominance function is defined as follows: in It is a discount factor. It is a parameter that controls the change in bias. It is the timing difference error, defined as: in The value function represents the value at the time of sampling; It is the target reward, defined as: It is the policy entropy term, and the formula is: In the formula, Represents a value function. This is the state value; Let represent the probability vector for each action in the discrete action space, expressed as: in It is the output of the strategy; The optimal action is obtained by analyzing the policy output, and the action with the highest probability is selected as the optimal action. Use gradient ascent to update policy parameters: in It is the learning rate; Under the updated policy parameters, perform several business transmissions to obtain training results, determine whether the model has converged, and if the model has not converged, perform the next training and update the parameters until the model converges.

Citation Information

Patent Citations

  • Double-layer satellite wireless access network slicing method based on deep reinforcement learning

    CN117202294A

  • Scheduling method for eMBB and URLLC coexisting data flow in 5G network based on reinforcement learning

    CN118338450A