A multi-time-scale-based ran slice hierarchical wireless resource management method

By adopting a multi-time-scale hierarchical wireless resource management method in RAN, utilizing the superposition mechanism of orthogonal multiple access and non-orthogonal multiple access technologies, combining high-level and low-level control models and deep reinforcement learning, the coordination problem of multiple network slices in RAN is solved, achieving efficient resource allocation for different services and improving network stability.

CN119653392BActive Publication Date: 2025-10-14NINGBO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411815846.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-10-14
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

How to effectively manage and coordinate multiple network slices in the radio access network (RAN), especially in the scenario where ultra-reliable low-latency communication and enhanced mobile broadband services coexist, solve the service conflicts caused by shared resource blocks, and the diversified service quality configuration technology that cannot effectively utilize limited and restricted resources in current 5G communications.

Method used

A multi-time-scale RAN slicing hierarchical radio resource management method is adopted. By dividing the total available bandwidth of the RAN into multiple resource blocks, the puncturing mechanism of orthogonal multiple access and the superposition mechanism of power domain non-orthogonal multiple access technology and space domain non-orthogonal multiple access technology are utilized. Combined with the high-level control model and the optimized low-level control model, the selector-critic framework of deep reinforcement learning is used to allocate radio resources, thus achieving efficient transmission of service model 1 and service model 2 on the same resource block.

Benefits of technology

It improves the adaptability and stability of network slicing in IoT scenarios, realizes intelligent and efficient resource allocation for different business needs, and reduces network complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653392B_ABST
    Figure CN119653392B_ABST
Patent Text Reader

Abstract

The application relates to a multi-time-scale-based RAN slice hierarchical wireless resource management method, which utilizes a puncturing mechanism of orthogonal multiple access and a puncturing mechanism composed of a power-domain non-orthogonal multiple access technology and a space-domain non-orthogonal multiple access technology to allow a service model 1 and a service model 2 to perform transmission on one resource block, so that the users of the service model 1 and the users of the service model 2 can more efficiently utilize wireless resources or share the same wireless resources. On this basis, the application combines a high-level control model with an optimization bottom-level control model to form a multi-time-scale wireless resource management strategy, and the strategy further combines deep reinforcement learning, so that a problem of solving a target function can be converted into a Gaussian Markov decision process by using an option-critic framework on a large time scale, and different service demands are corresponded to sub-targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a multi-time-scale based RAN slice hierarchical wireless resource management method. Background Art

[0002] With the advent of the Internet of Everything (IoE) era, numerous vertical services are emerging, each with diverse quality of service (QoS) requirements. Consequently, network slicing technology has emerged. This technology allows multiple logical networks to be virtualized on the same physical network to meet the specific needs of different users and applications. In this technology, each slice is dedicated to a specific service.

[0003] The main services of network slicing can be divided into three categories: ultra-reliable low-latency communication (uRLLC), enhanced mobile broadband (eMBB), and massive machine-type communication (mMTC). The provision of these services requires different network data rates, reliability, and latency.

[0004] However, effectively managing and coordinating multiple network slices within the radio access network (RAN) presents a significant challenge. Current 5G network slices fail to effectively utilize diverse Quality of Service (QoS) configuration technologies for limited and constrained resources. In scenarios where ultra-reliable low-latency communications and enhanced mobile broadband services coexist, service conflicts may arise during transmission due to shared resource blocks. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to effectively manage and coordinate multiple network slices in a radio access network (RAN). In order to solve the above technical problem, the present invention provides a RAN slice hierarchical wireless resource management method based on multiple time scales.

[0006] The present invention provides a multi-time-scale based RAN slice layered radio resource management method, comprising the following steps:

[0007] S1: In scenarios where two service models coexist, the total available RAN bandwidth is divided into multiple resource blocks, and the duration of a single time slot and the number of mini-time slots contained in each time slot are determined;

[0008] S2: At the beginning of the current time slot, resource blocks are allocated to service model 1 according to the resource requirements of service model 1;

[0009] S3: Activation time of service model 2, judging whether the resource block is available;

[0010] If yes, proceed to the next step;

[0011] If not, proceed to step S5;

[0012] S4: Allocate unused resource blocks to the users requiring service model 2, and then return to step S2 after the next time slot.

[0013] S5: In a mini-time slot after the activation of service model 2, an orthogonal multiple access puncturing mechanism is adopted to allow service model 2 to be transmitted preferentially, or an overlay mechanism composed of a power domain non-orthogonal multiple access technology and a space domain non-orthogonal multiple access technology is adopted to allow service model 1 and service model 2 to be transmitted simultaneously on the same resource block, and then the next step is executed;

[0014] S6: constructing a high-level control model to be parameterized, and optimizing the parameters of the high-level control model to be parameterized using an option-critic strategy of reinforcement learning to obtain a high-level control model;

[0015] The high-level control model is configured to obtain a channel state and a change in user service flow in each time slot based on a channel model and a communication scenario, and utilize the channel state and the change in user service flow to obtain a slice adjustment decision for prioritizing users of service model 1 and users of service model 2;

[0016] S7: Obtain the slice adjustment decision through the high-level control model, construct an optimized low-level control model corresponding to the slice adjustment decision, and solve the optimized low-level control model to obtain an optimal solution;

[0017] S8: Perform wireless resource scheduling based on the optimal solution through the base station to complete the data transmission of the service model 1 and the service model 2 in the current time slot, and return to execute step S2 after the next time slot.

[0018] The disclosed RAN slice hierarchical wireless resource management method based on multiple time scales can realize more efficient utilization of wireless resources or sharing of the same wireless resources for users of the service model 1 and users of the service model 2 by using the puncturing mechanism of orthogonal multiple access and the puncturing mechanism composed of the power domain non-orthogonal multiple access technology and the spatial domain non-orthogonal multiple access technology, and allowing the service model 1 and the service model 2 to transmit on one resource block. On this basis, the application proposes a combination of a high-level control model and an optimized bottom-level control model, forming a multi-time-scale wireless resource management strategy, and the strategy also combines deep reinforcement learning, which can convert the problem of solving the objective function into a Gaussian Markov decision process on a large time scale using the option-critic framework, and correspond different business needs to sub-goals. Through the scheme proposed in the application, the RAN network slice can make more intelligent and efficient wireless resource allocation according to the specific network state, and improve the adaptability and stability of the network slice to the Internet of Things scene.

[0019] In a possible implementation, in the step S6, the process of constructing the high-level control model to be parameter-optimized is as follows:

[0020] constructing a state space for describing the channel state and user traffic variation in each time slot;

[0021] constructing an action space for representing the performance of the slice of the service model 1 and the reliability of the slice of the service model 2;

[0022] constructing a reward function whose instantaneous value reflects the pros and cons of the slice performance and whose expected cumulative value reflects the adaptive adjustment of the slice to different traffic;

[0023] The high-level control model corresponding to the scheme can simultaneously satisfy the performance of the service model 1 and the reliability of the service model 2, and can reduce the complexity of the network to a certain extent.

[0024] In a possible implementation, the state space is expressed as follows:

[0025] ,

[0026] In the formula,

[0027] denotes a set of users of the service model 1;

[0028] denotes a time slot;

[0029] denotes a set of users of the service model 2;

[0030] represents that resource block b is allocated to the user e of the service model 1 at time slot t;

[0031] represents that resource block b is allocated to the user u of the service model 2 at the micro time slot m of time slot t;

[0032] represents whether the user e and the user u adopt the puncturing mechanism;

[0033] represents the superposition of the power domain non-orthogonal multiple access of the user e and the user u;

[0034] represents whether the user e and the user u adopt the superposition of the spatial domain non-orthogonal multiple access;

[0035] represents the channel estimation value of the user u in resource block b;

[0036] represents the channel estimation value of the user e in resource block b;

[0037] The state space corresponding to this scheme can accurately describe the channel state and the user service flow change in each time slot, which is helpful for the high-level control model to obtain the slice adjustment decision for prioritizing the users of the service model 1 and the users of the service model 2.

[0038] In a possible implementation, the action space is expressed as follows:

[0039] ,

[0040] In the formula,

[0041] ; represents that the slice focuses on protecting the transmission rate of the service model 1 in this time period, represents that the slice focuses on the reliability of the user of the service model 2 in this time period; if both are 1, a compromise scheduling mode is adopted;

[0042] The action space corresponding to this scheme is helpful for the high-level control model to select the performance of the slice of the service model 1 and the reliability of the slice of the service model 2, so as to accurately obtain the slice adjustment decision.

[0043] In a possible implementation, the reward function is expressed as follows:

[0044] ,

[0045] ,

[0046] ,

[0047] ,

[0048] ,

[0049] ,

[0050] ,

[0051] ,

[0052] wherein,

[0053] represents the overall packet loss rate of a micro-slot;

[0054] represents the maximum network complexity tolerable in the RAN;

[0055] represents the data transmission rate of the user e of the service model 1 at the time slot t;

[0056] represents the minimum transmission rate of the service model 1;

[0057] The immediate value of the reward function of this scheme reflects the pros and cons of the slice performance; and the expected cumulative value thereof reflects the adaptive adjustment of the slice under different service flows, which helps to optimize the parameters of the high-level control model through reinforcement learning.

[0058] In a possible implementation, the optimized bottom-level control model is expressed as follows:

[0059] ,

[0060] ,

[0061] wherein,

[0062] is a constant, and the selection thereof is based on the priorities of the service model 1 and the service model 2;

[0063] In the optimized model of this scheme, C1, C2 ensure the performance isolation of the same service, C3 ensures that the service model 1 and the service model 2 can only use one multiplexing technology, and C4, C5 ensure the real-time power and transmission rate of the user, respectively. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1A flowchart of a multi-time-scale based RAN slice hierarchical radio resource management method disclosed in an embodiment of the present application;

[0065] Figure 2 This is the scheduling of eMBB and uRLLC services in this embodiment. DETAILED DESCRIPTION

[0066] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.

[0067] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0068] As shown in Figure 1, the embodiment of the present application discloses a multi-time-scale based RAN slice hierarchical radio resource management method, including the following steps:

[0069] S1: In a scenario where two service models coexist, the total available bandwidth of the RAN is divided into multiple resource blocks, and the duration of a single time slot and the number of mini-time slots contained in each time slot are determined.

[0070] Specifically, the RAN in this embodiment has a typical downlink cellular network system with a single base station. The total available bandwidth is divided into B resource blocks. In these two service models, service model 1 is eMBB and service model 2 is uRLLC. In order to ensure the slicing performance of eMBB, the user's transmission status is strictly guaranteed in this embodiment. That is, maximizing

[0071] ,

[0072] Where,

[0073] represents the rate of eMBB user e at time slot t;

[0074] Indicates the maximum rate defined for user e.

[0075] At the same time, to ensure the reliability and low latency of uRLLC slices, this embodiment defines that a uRLLC user u sends a fixed-size data packet, which must be transmitted within a mini-time slot during base station transmission, otherwise the data packet will be discarded. The packet loss of the uRLLC user is as follows:

[0076] ,

[0077] Where,

[0078] represents the transmission rate of the uRLLC user at mini-timeslot m in time slot t;

[0079] It represents the fixed packet size of uRLLC. The formula ensures that the uRLLC packet is transmitted within one micro-slot.

[0080] Figure 1 This field indicates the scheduling of eMBB and uRLLC services. A frame is 1ms and contains 14 OFDM symbols (15kHz). A frame is defined as one time slot, which is further divided into multiple mini-slots. eMBB services are scheduled in each time slot, while uRLLC services are scheduled in each mini-slot.

[0081] S2: At the initial moment of the current time slot, resource blocks are allocated to service model 1 according to the resource requirements of service model 1.

[0082] S3: Activation time of service model 2, judging whether the resource block is available;

[0083] If yes, proceed to the next step;

[0084] If not, execute step S5.

[0085] S4: Allocate unused resource blocks to users requiring service model 2, and return to step S2 after the next time slot.

[0086] S5: In a mini-time slot after service model 2 is activated, an orthogonal multiple access puncturing mechanism is used to allow service model 2 to be transmitted first, or an overlay mechanism consisting of power domain non-orthogonal multiple access technology and space domain non-orthogonal multiple access technology is used to allow service model 1 and service model 2 to be transmitted simultaneously on the same resource block, and then the next step is executed.

[0087] When considering performance, we must also pay attention to network complexity. When the channel gain gap between overlapping users is too small, the receiver requires more power to decode the overlapping users. Alternatively, when the spatial correlation is lower, the network complexity is lower. The network complexity is shown in the following formula, which can be expressed as the slice complexity introduced by the puncturing and superposition mechanisms respectively.

[0088] ,

[0089] ,

[0090] ,

[0091] Where,

[0092] Indicates whether users e and u adopt the hole punching mechanism;

[0093] represents the superposition of power domain non-orthogonal multiple access of user e and user u;

[0094] Indicates whether user e and user u use the superposition of spatial domain non-orthogonal multiple access;

[0095] represents the channel estimation value of user u in resource block b;

[0096] is the channel estimation value of user e in resource block b.

[0097] S6: Construct a high-level control model for parameter optimization and optimize the parameters of the high-level control model using a reinforcement learning option-critic strategy to obtain a high-level control model. The high-level control model is configured to obtain the channel state and user traffic flow changes within each time slot based on the channel model and communication scenario. The channel state and user traffic flow changes are used to determine the slice adjustment decisions used to prioritize users of service model 1 and service model 2.

[0098] The high-level control model in this embodiment selects two slice modes when making high-level decisions: radio resource allocation favors eMBB slices or radio resource allocation favors uRLLC slices. For example, when transmitting important high-definition video streams, the controller favors eMBB services to ensure video stability and continuity. However, during autonomous driving, to ensure reliable control within a short timeframe, uRLLC services must be prioritized.

[0099] Specifically, in step S6, the process of constructing the high-level control model to be optimized is as follows:

[0100] First, a state space is constructed to describe the channel state and user traffic flow changes in each time slot.

[0101] The state space is a comprehensive description of the channel state and traffic flow, and is the key to ensuring the effectiveness of the algorithm. The state space established in this embodiment is expressed as:

[0102] ,

[0103] Where,

[0104] represents the set of eMBB users, which is the buffering status of eMBB users at time slot t;

[0105] represents a time slot;

[0106] The uRLLC user set is the active set of uRLLC devices at time slot t and mini-time slot m.

[0107] Indicates that resource block b is allocated to eMBB user e at time slot t;

[0108] It indicates that resource block b is allocated to user u of uRLLC at mini-slot m of time slot t.

[0109] Next, an action space is constructed to represent the performance of selecting the slice of business model 1 and the reliability of the slice of business model 2.

[0110] The action space represents the choice of eMBB slice performance and uRLLC slice reliability. Given that the high-level control model is to decide the scheduling mode of radio resources, this embodiment sets the action vector of the high-level control model as the priority scheduling of the slice for eMBB and uRLLC at the current moment. The corresponding action space is expressed as follows:

[0111] ,

[0112] Where,

[0113] , Indicates that the slice in this time period focuses on protecting the transmission rate of eMBB. Indicates the reliability of the slice for uRLLC users during this time period. If both are 1, a compromise scheduling method is adopted. Different action selections will affect the user perception function of the service in the underlying model.

[0114] Finally, a reward function is constructed in which the instantaneous value reflects the performance of the slice, and the expected cumulative value reflects the adaptive adjustment of the slice under different business traffic.

[0115] The instantaneous value of the reward function reflects the performance of the slice, while the expected cumulative value of the reward function reflects the slice's adaptive adjustment to different service traffic conditions. Based on this, this embodiment designs a combination of a long-term reward function and a short-term reward function to form the reward function.

[0116] (1) Short-term reward function

[0117] The short-term reward function mainly refers to the reliability of the uRLLC service. Specifically, once the uRLLC service is activated, it needs to be immediately allocated resource blocks within the micro-time slot. The reliability of the service can be measured by the packet loss rate. If the uRLLC service fails to complete the data transmission within the specified micro-time slot, this situation will be regarded as packet loss. Therefore, the reliability of the activated uRLLC service instance can be analyzed by calculating the total packet loss rate of the uRLLC service at a specific micro-time. Specifically, at any micro-time slot, the packet loss rate of uRLLC cannot exceed the maximum tolerable packet loss probability in the RAN slice, that is,

[0118] ,

[0119] The overall packet loss rate of the mini-timeslot can be calculated using the following formula based on the total number of activated uRLLC devices:

[0120] .

[0121] In addition, the multiplexing of eMBB and uRLLC must be considered in each mini-timeslot. The introduction of the superposition / puncturing mechanism will increase the complexity of the network. When the channel gain difference between superimposed users is too small, the receiver requires more power to decode the overlapping users. Or, when the spatial correlation is lower, the network complexity is reduced. Ensuring low network complexity in each mini-timeslot can easily ensure the difficulty of radio resource management of the slice. Therefore, to ensure the feasibility of slicing, this embodiment also designs a reward function for network complexity:

[0122] ,

[0123] Indicates the maximum tolerable network complexity in the RAN.

[0124] (2) Long-term reward function

[0125] To ensure that slices have relatively fair rewards for different services, this embodiment further designs a reward function for the eMBB transmission rate. Since its training time scale is the sum of m mini-slots, a larger reward feedback is required. The eMBB transmission rate is affected by the uRLLC at each mini-slot:

[0126] ,

[0127] Where,

[0128] represents the signal-to-interference-and-noise ratio of user e at mini-slot m of time slot t in resource block b.

[0129] Since the eMBB service must guarantee its transmission performance, this embodiment designs a minimum rate limit for each eMBB. When the eMBB slice meets the rate requirement, an additional reward is obtained at time slot t. The reward function at each time slot is defined as:

[0130] ,

[0131] Note that even users belonging to the same category have inherent service requirements, such as the minimum transmission rate of eMBB. , may also be different. Therefore, the high-level control model reward function is:

[0132] ,

[0133] is the uRLLC business reward in the short-term reward function;

[0134] Rewards for the network;

[0135] For different user priority scheduling strategies, the convergence and reliability of training can be ensured by setting different reward weights.

[0136] In step S6, when optimizing the parameters of the high-level control model using the option-critic strategy of reinforcement learning, the option-critic (OC) structure is used to ensure that the high-level control model makes the optimal slice adjustment decision in different environments, ensuring flexibility on the time scale, that is, allocating wireless resources in each micro-time slot and determining the service priority of the next time period on a flexible and variable long time scale.

[0137] The option-critic architecture optimizes the internal policy of an option and the termination network based on the policy gradient principle to maximize the discounted expected return. In the option-critic architecture, the policy gradient within the option and the policy gradient of the termination condition are of primary interest. Since this embodiment models the internal options as a convex optimization problem rather than a Markov decision process, only the policy gradient of the termination condition is considered here.

[0138] The gradient of the termination network of the high-level control model can be expressed as:

[0139] ,

[0140] Where;

[0141] Parameters representing the termination network;

[0142] Indicates that the thinking cost parameter is introduced to prevent frequent switching of options;

[0143] represents the advantage function of the termination network;

[0144] represents the option value function, which is updated using the following formula:

[0145] ; (1)

[0146] The network updates are as follows:

[0147] . (2)

[0148] The training process based on the option-critic structure is as follows: the channel state and service traffic at the current moment are used as input, the Q network outputs the Q value of each slice scheduling behavior, and the wireless resource scheduling criterion with the largest Q value is selected; the Q network is updated according to formula (1), and then the termination network is updated according to formula (2), and the output is the termination probability of each slice strategy. If the slice strategy at the current moment is terminated at the next moment, the slice strategy with the largest Q value is re-selected using the Q network at the next moment.

[0149] Specifically, this embodiment uses the parameters A deep neural network is used to represent the long-time optimization policy πe, and a real-time resource allocation objective function is used to represent the short-time optimization policy πu. Because the lower layer is a convex optimization problem, the slicing strategy on the short time scale is guaranteed to be stable and convergent. On this basis, the upper layer model can be trained using a deep neural network to achieve convergence of πe.

[0150] S7: Obtain a slice adjustment decision through a high-level control model, construct an optimized low-level control model corresponding to the slice adjustment decision, and solve the optimized low-level control model to obtain an optimal solution.

[0151] To further analyze the selection of radio resource scheduling methods in slice adjustment decisions derived from the high-level control model, this embodiment specifically considers the traffic characteristics and service patterns of different QoS categories in eMBB and uRLLC services. To this end, different utility models are developed in this embodiment to quantitatively measure the UE's service satisfaction level.

[0152] (1) QoS categories that guarantee transmission performance

[0153] For enhanced services (e.g., IP voice transmission and adaptive video streaming), due to the user's inherent and preferred service requirements, the user's performance gain mainly depends on the quality of resource allocation and is therefore susceptible to fluctuations. Therefore, the strategy tends to focus on the service quality of eMBB.

[0154] (2) QoS categories focusing on reliability and low latency

[0155] For real-time, mission-critical, and demanding applications (e.g., autonomous driving, vehicle platooning), consistent and guaranteed service delivery levels are required.

[0156] During the parameter optimization process of the high-level control model, the priority of the slice service has been obtained through the option-critic structure, allowing the base station to reasonably and differentially adjust its resource allocation strategy to adapt to the real-time network load, dynamic resource conditions and various user needs. Different utility functions are proposed for the action selection of the high-level control model:

[0157] ,

[0158] Service model 1 represents the eMBB service, and service model 2 represents the uRLLC service. This embodiment further proposes an objective function p0 for allocating wireless resources in mini-time slots:

[0159] ,

[0160] Among them, C1 and C2 ensure the performance isolation of the same service, C3 ensures that eMBB and uRLLC can only use one multiplexing technology, and C4 and C5 respectively guarantee the user's real-time power and transmission rate. The choice is mainly affected by the priorities of eMBB and uRLLC.

[0161] The above optimization model is a deterministic convex optimization problem, and there must be a point where the constraints are not equal, thus satisfying the Slater condition. Therefore, we use the dual decomposition and KKT optimality conditions to solve the original convex problem. Now let's process the objective function:

[0162] ,

[0163] like is the optimal solution, then The following constraints must be satisfied and there must be a multiplier Satisfies the following formula:

[0164] ,

[0165] ,

[0166] ,

[0167] ,

[0168] ,

[0169] Because the original problem is a convex problem, the points that satisfy the KKT condition are also the original and dual optimal solutions. This method can be used to allocate power to users in real time at micro-time slots, and the power allocation meets the stability characteristics.

[0170] S8: The base station performs wireless resource scheduling based on the optimal solution to complete the data transmission of service model 1 and service model 2 in the current time slot, and returns to execute step S2 after the next time slot.

[0171] The multi-timescale hierarchical radio resource management method for RAN slices disclosed in this embodiment utilizes an orthogonal multiple access puncturing mechanism and a puncturing mechanism composed of power-domain non-orthogonal multiple access technology and space-domain non-orthogonal multiple access technology to allow service model 1 and service model 2 to be transmitted on the same resource block. This allows for more efficient use of radio resources or sharing of the same radio resources for users of service model 1 and service model 2. On this basis, this embodiment proposes a combination of a high-level control model and an optimized low-level control model to form a multi-timescale radio resource management strategy. This strategy also incorporates deep reinforcement learning, utilizing an option-critic framework to transform the objective function solution into a Gaussian Markov decision process on a large timescale, mapping different service requirements to sub-goals. Through the proposed solution, RAN network slices can make more intelligent and efficient radio resource allocation based on specific network conditions, improving the adaptability and stability of network slices for IoT scenarios.

[0172] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0173] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0174] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A multi-time-scale based RAN slice hierarchical radio resource management method, characterized in that: The steps include: S1: In scenarios where two service models coexist, the total available RAN bandwidth is divided into multiple resource blocks, and the duration of a single time slot and the number of mini-time slots contained in each time slot are determined; S2: At the beginning of the current time slot, resource blocks are allocated to service model 1 according to the resource requirements of service model 1; S3: Activation time of service model 2, judging whether the resource block is available; If yes, proceed to the next step; If not, proceed to step S5; S4: Allocate unused resource blocks to the users requiring service model 2, and then return to step S2 after the next time slot. S5: In a mini-time slot after the activation of service model 2, an orthogonal multiple access puncturing mechanism is adopted to allow service model 2 to be transmitted preferentially, or an overlay mechanism composed of a power domain non-orthogonal multiple access technology and a space domain non-orthogonal multiple access technology is adopted to allow service model 1 and service model 2 to be transmitted simultaneously on the same resource block, and then the next step is executed; S6: constructing a high-level control model to be parameterized, and optimizing the parameters of the high-level control model to be parameterized using an option-critic strategy of reinforcement learning to obtain a high-level control model; The high-level control model is configured to obtain a channel state and a change in user service flow in each time slot based on a channel model and a communication scenario, and utilize the channel state and the change in user service flow to obtain a slice adjustment decision for prioritizing users of service model 1 and users of service model 2; S7: Obtain the slice adjustment decision through the high-level control model, construct an optimized low-level control model corresponding to the slice adjustment decision, and solve the optimized low-level control model to obtain an optimal solution; S8: performing wireless resource scheduling based on the optimal solution by the base station to complete the data transmission of the service model 1 and the service model 2 in the current time slot, and returning to execute step S2 after the next time slot; In step S6, the process of constructing the high-level control model to be optimized is as follows: Construct a state space for describing the channel state and user traffic flow changes in each time slot; Constructing an action space for representing the performance of selecting the slice of the business model 1 and the reliability of the slice of the business model 2; The instantaneous value is constructed to reflect the performance of the slice, and the expected cumulative value is constructed to reflect the reward function of the slice's adaptive adjustment under different business traffic.

2. The multi-time-scale based RAN slice layered radio resource management method according to claim 1, characterized in that The state space is expressed as follows: , Where, A set of users representing the business model 1; represents a time slot; A user set representing the business model 2; Indicates that resource block b is allocated to user e of service model 1 at time slot t; Indicates that resource block b is allocated to user u of service model 2 at mini-time slot m of time slot t; Indicates whether users e and u adopt the hole punching mechanism; represents the superposition of power domain non-orthogonal multiple access of user e and user u; Indicates whether user e and user u use the superposition of spatial domain non-orthogonal multiple access; represents the channel estimation value of user u in resource block b; represents the channel estimation value of user e in resource block b.

3. The multi-time-scale based RAN slice hierarchical radio resource management method according to claim 2, characterized in that: The action space is expressed as follows: , Where, ; Indicates that the slice in this time period focuses on protecting the transmission rate of service model 1. It indicates that the slice in this time period focuses on the reliability of users of the service model 2; if both are 1 at the same time, a compromise scheduling method is adopted.

4. The multi-time-scale based RAN slice layered radio resource management method according to claim 3, characterized in that The reward function is calculated as follows: , , , , , , , , Where, Indicates the overall packet loss rate of a mini-slot; Indicates the maximum tolerable network complexity in RAN; represents the data transmission rate of user e in the service model 1 at time slot t; Indicates the minimum transmission rate of the service model 1.

5. The multi-time-scale based RAN slice hierarchical radio resource management method according to claim 4, characterized in that: The optimal underlying control model is expressed as follows: , , Where, is a constant, which is selected based on the priority of the business model 1 and the business model 2.

Citation Information

Patent Citations

  • D2D communication network slice allocation method based on deep reinforcement learning

    CN113163451A

  • Network slice optimization processing method and system

    CN113992524A