A multi-service joint downlink resource allocation method
By using neural networks to predict user connection status and the Dueling Deep Q-Learning algorithm to optimize resource allocation, the resource allocation problem in the coexistence scenario of eMBB and URLLC services is solved, achieving a balance between high reliability and high transmission rate, and avoiding resource waste.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-05-15
- Publication Date
- 2026-04-21
AI Technical Summary
In scenarios where eMBB and URLLC services coexist, existing technologies suffer from unbalanced resource allocation, waste, and degraded transmission performance. They also lack effective scheduling methods and cannot simultaneously meet the requirements of high reliability and high transmission rate.
By constructing a neural network model to predict the connection status between users and base stations, and combining it with the Dueling Deep Q-Learning algorithm, resource allocation strategies are optimized to ensure the reliability of URLLC services and the transmission rate of eMBB services. NOMA and puncturing transmission models are used for resource allocation.
This approach maximizes the transmission rate of eMBB services while meeting the reliability requirements of URLLC services, avoiding resource waste and improving resource utilization.
Smart Images

Figure CN116489774B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communications and relates to the resource allocation problem in the scenario where eMBB users and URLLC users coexist in mobile communications. Specifically, it relates to a method for multi-service joint downlink resource allocation. Background Technology
[0002] Currently, with the development of 5G and the gradual confirmation of expected 6G scenarios, applications such as autonomous driving and intelligent vehicle-to-everything (V2X) have received stronger theoretical support. The coexistence of eMBB (enhanced Mobile Broadband) users and URLLC (Ultra-Reliable Low-Latency Communication) users will become a typical scenario for future mobile V2X. URLLC services, represented by autonomous driving, require high communication reliability and ultra-low latency to ensure vehicle safety. Meanwhile, eMBB services, represented by intelligent V2X users, require high-speed mobile networks to support VR videos, games, and other applications. Therefore, the allocation of downlink resources for multiple services presents new challenges to traditional wireless resource scheduling methods.
[0003] To address the coexistence problem of eMBB and URLLC services, the industry has proposed several solutions: In [1], it was proposed to reserve specific time-frequency resources for URLLC services to meet the needs of bursty URLLC services, while other resources are used for eMBB service transmission. This resource reservation method can well meet the latency and reliability requirements of bursty URLLC service transmission, but at other times, the reserved time-frequency resources are not used for service transmission, which causes an imbalance and waste in resource allocation. In [2], a resource preemption method for URLLC services was proposed, that is, under normal circumstances, the communication resources required by eMBB users are pre-allocated, and then in each mini-time slot, the communication resources are preferentially allocated to bursty URLLC users, forming a situation where high-priority URLLC users preempt the communication resources of low-priority eMBB users. This method can ensure that time-frequency resources are not wasted and meet the latency and reliability requirements of URLLC services, but eMBB services will be allocated discontinuous time-frequency resources, causing a decrease in transmission performance. The authors of [3] proved that the NOMA scheme can be used in the resource allocation of URLLC services. In [4], the authors used the NOMA scheme to allocate resources to both eMBB and URLLC users simultaneously, thereby maximizing resource utilization and meeting the high-speed requirements of eMBB services. However, this would reduce the reliability of URLLC services. Article [5] combines the resource preemption and NOMA schemes in the above literature and adopts a learning-based method to intelligently allocate strategies according to different services and channel information. To a certain extent, it achieves a trade-off between the latency reliability of URLLC services and the high speed of eMBB services. However, it lacks the estimation of the expected completion of services and the modification of the allocation strategy based on this estimation.
[0004] In summary, some existing technologies lack effective utilization of frequency domain resources, resulting in resource waste; some focus only on a specific service while ignoring the transmission requirements of another type of service; some achieve a trade-off between the transmission requirements of the two services, but lack consideration for the prediction of service completion, thus reducing the utilization rate of some resources; and there is a lack of an effective scheduling method that fully considers the transmission requirements of the two services and predicts the service completion status for resource scheduling.
[0005] References
[0006] [1] M.Setayesh, S.Bahrami and VWSWong, "Joint PRB and PowerAllocation for Slicing eMBB and URLLC Services in 5G C-RAN," GLOBECOM 2020-2020IEEE Global Communications Conference, 2020, pp.1-6, doi:10.1109 / GLOBECOM42002.2020.9322568.
[0007] [2]M.Alsenwi, NHTran, M.Bennis, SRPandey, AKBairagi and CSHong, "Intelligent Resource Slicing for eMBB and URLLC Coexistence in 5G and Beyond: A Deep Reinforcement Learning Based Approach," in IEEE Transactions onWireless Communications, vol.20, no.7, pp.4585-4600, July2021, doi:10.1109 / TWC.2021.3060514.
[0008] [3]G.Liu, Z.Wang, J.Hu, Z.Ding and P.Fan, "Cooperative NOMA Broadcasting / Multicasting for Low-Latency and High-Reliability 5G Cellular V2XCommunications," in IEEE Internet of Things Journal, vol.6, no.5, pp.7828-7838, Oct.2019, doi:10.1109 / JIOT.2019.2908415.
[0009] [4] A. Manzoor, SMA Kazmi, SRPandey and CSHong, "Contract-BasedScheduling of URLLC Packets in Incumbent EMBB Traffic," in IEEE Access, vol.8, pp.167516-167526, 2020, doi:10.1109 / ACCESS.2020.3023128.
[0010] [5] Y.Prathyusha and T.-L.Sheu, "Coordinated Resource Allocations foreMBB and URLLC in 5GCommunication Networks," in IEEE Transactions on VehicularTechnology, vol.71, no.8, pp.8717-8728, Aug.2022, doi:10.1109 / TVT.2022.3176018. Summary of the Invention
[0011] To address the issue of multi-service resource allocation in mobile communication scenarios, this invention proposes a multi-service joint downlink resource allocation method. This method fully considers the coexistence of multiple services in actual communication scenarios and efficiently utilizes frequency domain resources to meet the significance of multi-service transmission indicators.
[0012] The specific steps of the multi-service joint downlink resource allocation method are as follows:
[0013] Step 1: Build a communication scenario that includes base stations and users. Divide the users connected to each base station into two types: eMBB service and URLLC service. Record the distance from all users to the base station in each time slot.
[0014] The base station set is BS = {BS1, BS2, ..., BS} i ,...BS n}, for base station BS i The connected users include: U = {1,2,...,u} representing the set of users who require URLLC services, and E = {1,2,...,e} representing the set of users who require eMBB services;
[0015] Let the m-th time slot be T. m The corresponding user base station distance is:
[0016] in, This indicates that user e, who requires eMBB services, is in time slot T. m Corresponding base station BS i distance, This indicates that user u, who requires URLLC services, is in time slot T. m Corresponding base station BS i The distance.
[0017] Step 2: Quantitatively define the transmission models for eMBB and URLLC users respectively, and calculate the maximum transmission rate of users under different transmission models;
[0018] Transmission models include the puncturing transmission model and the NOMA transmission model;
[0019] For eMBB users e i 1) Under the NOMA transmission model, the base station BS i Connected user e i The SINR is:
[0020]
[0021] in, Indicates base station BS i The power currently allocated to eMBB services, where N0 represents Gaussian white noise. User e i Current channel gain includes large-scale fading, small-scale fading, and shadowing fading. The SIC coefficient, Indicates base station BS i The power currently allocated to the URLLC service.
[0022] 2) Under the punched-hole transmission model, user e i No power is allocated, and SINR is not calculated.
[0023] Finally, according to user e iThe SINR, combined with the Shannon channel capacity, yields the user e i Maximum transmission rate:
[0024]
[0025] B represents the allocated channel bandwidth, and (·) indicates that SINR applies to both shared modes.
[0026] For URLLC user u i 1) Under the NOMA transmission model, the base station BS i Connected user u i The SINR is:
[0027]
[0028] Representing user u i The current channel gain;
[0029] 2) In the punched-hole transmission model, user u i The SINR is:
[0030]
[0031] Finally, according to user u i SINR, combined with finite block length coding theory, for user u i Calculate the maximum transmission rate:
[0032]
[0033] in, Represents the subcarrier bandwidth, T represents the mini-slot length, and ε c Indicates the probability of decoding error. This indicates the number of resource blocks allocated to this URLLC service at the same time. Indicates channel dispersion; Q G () represents the right-tail function of the Gaussian normal distribution;
[0034] Step 3: Aggregate the distance from the current time slot user to the base station and the distance from the previous time slot user to the base station, input the aggregated distance into the constructed neural network model, and output the connection prediction matrix.
[0035] The m-th time slot T m The corresponding user base station distance D m The latency threshold matrix for user services is as follows:
[0036] Construct a neural network model with multiple hidden layers, taking the distance matrix D at the current time m as input. mThe distance matrix D at the previous time step m-1 and user service latency threshold The network is trained using binary cross-entropy as the loss function, and the output binary scalar matrix is used as the connection prediction matrix F. pred ,as follows:
[0037]
[0038] This represents the predicted connectivity of a specific user in the current m-th time slot. This indicates that, based on the prediction, the user's requested service can be completed within the service time threshold. This indicates that the transmission could not be completed within the delay threshold or that the user left the cell during transmission. All users within the m-th time slot... The connection prediction matrix F is formed. pred .
[0039] Step 4: Establish optimization objectives based on the high reliability requirements of URLLC services and the high transmission rate requirements of eMBB services, and combine them with the connection prediction matrix F. pred A comprehensive decision-making process was adopted to maximize eMBB transmission rate and eMBB reliability while ensuring URLLC reliability.
[0040] The optimization issues are as follows:
[0041]
[0042] Among them, R eMBB Indicates the transmission rate of eMBB services; Pr E Indicates the probability of eMBB service transmission failure; Pr U Indicates the probability of URLLC service transmission failure;
[0043] Constraint C1 represents the sum of the reward weights for eMBB rate, eMBB reliability, and URLLC reliability, which equals 1; α i These are the weighting coefficients.
[0044] Constraint C2 represents the total power of the base station's transmitted signals within the cell being lower than the power limit P. max ;P i For base station BS i Power allocated to users within the community;
[0045] Constraint C3 means that the decoding priority of URLLC services is higher than that of eMBB services in the same time slot; They represent base stations (BS) respectively. i Power allocated to URLLC and eMBB users;
[0046] Constraint C4 means filtering out users predicted to be disconnected in the connection prediction matrix in the overall objective function, and the impact of the power allocated to these users on the objective function is ignored.
[0047] Constraint C5 indicates that URLLC services must strictly meet reliability QoS requirements.
[0048] Constraint C6 indicates that for any type of service transmission to be successful, its transmission delay must be less than the delay threshold. (·) refers to various services including URLLC and eMBB;
[0049] Constraint C7 indicates that for any type of service transmission to be successful, its signal-to-interference-plus-noise ratio (SIR) must be greater than the threshold condition.
[0050] Step 5: Train the reinforcement learning resource scheduling model based on the Dueling Deep Q-Learning algorithm, solve the optimization objective function, and thus give the optimal selection strategy for URLLC and eMBB services to preempt or share resources, and obtain the results of multi-service joint downlink resource allocation.
[0051] The reinforcement learning resource scheduling model consists of three layers: the first layer network provides the reserved resource ratio for subsequent URLLC services based on the current maximum threshold delay of eMBB services, channel gain, and connection prediction matrix; the second layer network provides the allocation action for each resource block based on the resource reservation rate and connection prediction matrix provided by the upper layer network, combined with the channel gain; and the third layer network provides resource scheduling decisions based on the number of arriving URLLC services and the power allocation of hourly slots.
[0052] The advantages of this invention are:
[0053] 1) A multi-service joint downlink resource allocation method that senses the user's location information within the cell in real time and predicts the future connection status between the user and the base station based on the user's temporal distance matrix and service threshold delay. Through a special design of the reward function, the reinforcement learning algorithm's strategy tends to allocate resources to services that are predicted to remain in the cell within the threshold delay and complete transmission within the delay threshold, thus optimizing the resource scheduling strategy and avoiding resource waste.
[0054] 2) A multi-service joint downlink resource allocation method first allocates resources for eMBB services, and then intersperses URLLC short message transmission in the time slots of data transmission. The latency-sensitive URLLC services are transmitted through puncturing preemptive scheduling and NOMA. This method maximizes the protection of the transmission rate of eMBB services while ensuring the latency and reliability QoS requirements of URLLC services, and realizes joint resource allocation for the two services. Attached Figure Description
[0055] Figure 1 This is a flowchart of a multi-service joint downlink resource allocation method according to the present invention;
[0056] Figure 2 This is a schematic diagram of a communication scenario including a base station and a user, as constructed by the present invention.
[0057] Figure 3 This is a schematic diagram of the resource blocks used in the frame structure model of this invention. Detailed Implementation
[0058] The present invention will be further described below with reference to the accompanying drawings and examples.
[0059] Based on distance-aware prediction of user base station connection prediction, this invention presents a multi-service joint downlink resource allocation method. This method allocates downlink resources for users of different services within a cell. It generates corresponding connection opportunities based on the distribution of base stations and users, and predicts the transmission success rate of arriving services based on the perceived user distance-time series matrix, generating a connection prediction matrix. Based on the prediction matrix, user distribution, and channel gain, it performs joint resource scheduling for arriving eMBB and URLLC services, thereby maximizing the eMBB transmission rate while meeting the QoS requirements of URLLC reliability. By incorporating the connection prediction matrix, this method considers the transmission success rate within the base station coverage area in addition to the decision information on transmission resources and channels, improving the resource utilization of the given decisions. It is an efficient multi-service joint resource optimization method.
[0060] The multi-service joint downlink resource allocation method generates connection opportunities between the base station and the user based on the effective coverage of the base station. For communication between the base station and the user, it is divided into eMBB service and URLLC service. At the beginning of the first time slot, all time slots are allocated to eMBB service for transmission. If URLLC service suddenly arrives during the transmission of eMBB service, the transmission of URLLC service is scheduled first, and resources are preempted or shared in the subsequent small time slots.
[0061] Regarding how to make decisions on URLLC service resource scheduling strategies, this method, based on the transmission quality of eMBB service and the reliability requirements of URLLC service, and considering user distribution and channel resources in the cell, presents a joint optimization objective function based on eMBB rate, transmission success rate and URLLC service reliability; the goal of the service multiplexing strategy is to maximize this objective function.
[0062] For service reuse strategies, base stations within a cell can also sense the distance between the user and the base station in real time. Combining this with time-series information, they can predict whether the user will leave the base station's connection range. This is a typical spatiotemporal prediction problem, requiring consideration of the characteristics and patterns of spatiotemporal data. This problem can be solved using recurrent neural networks (RNNs), a type of neural network model suitable for modeling time-series data. For spatiotemporal prediction problems, RNNs can perform predictions by modeling time-series data.
[0063] Since URLLC services may arrive at any time slot, and the resource reuse strategies and resource allocation of adjacent services can influence each other, the optimal decision for each time slot is difficult to solve through mathematical derivation. This method proposes a reinforcement learning approach as a solution to this complex optimization problem. By combining the user access prediction matrix, it provides the optimal strategy for URLLC and eMBB services to preempt or share resources. This approach maximizes the objective function while satisfying relevant constraints, thus yielding the optimal decision.
[0064] Based on reinforcement learning, a three-layer network model based on Dueling Deep Q-Learning is proposed. The first layer network gives the reserved resource ratio for subsequent URLLC services based on the maximum threshold delay of the current eMBB service, channel gain, and connection prediction matrix. The second layer network gives the allocation action for each resource block based on the resource reservation rate given by the upper layer network, the connection prediction matrix, and the channel gain. The third layer network gives the resource scheduling decision based on the number of arriving URLLC services and the power allocation of the hourly slot.
[0065] like Figure 1 As shown, the specific steps are as follows:
[0066] Step 1: Build a communication scenario that includes base stations and users. Divide the users connected to each base station into two types: eMBB service and URLLC service. Record the distance from all users to the base station in each time slot.
[0067] like Figure 2 As shown, the base station set is BS = {BS1, BS2, ..., BS} i ,...BS n}, for base station BS i The connected users include the following two types: U = {1,2,...,u} represents the set of users who require URLLC services, and E = {1,2,...,e} represents the set of users who require eMBB services;
[0068] Base station attributes include: base station coordinates, antenna height, maximum transmit power, etc.
[0069] User attributes include: user coordinates, user service type, and corresponding service QoS requirements;
[0070] Based on the generated data point information, the distance from all users to the base station is recorded in each time slot: let the m-th time slot be T. m The corresponding user base station distance is:
[0071] in, This indicates that user e, who requires eMBB services, is in time slot T. m Corresponding base station BS i distance, This indicates that user u, who requires URLLC services, is in time slot T. m Corresponding base station BS i The distance.
[0072] The simulation environment setup for this embodiment is as follows:
[0073]
[0074]
[0075] Step 2: Quantitatively define the transmission models for eMBB and URLLC users respectively, and calculate the maximum transmission rate of users under different transmission models;
[0076] Depending on the shared strategy, the transmission models include the puncturing transmission model and the NOMA transmission model;
[0077] For eMBB users e i 1) Under the NOMA transmission model, considering the communication interference from other users within the cell, the environmental noise can be obtained at the base station BS. i In the corresponding community, user e i The SINR is:
[0078]
[0079] in, Indicates base station BS i The power currently allocated to eMBB services, where N0 represents Gaussian white noise. User e i Current channel gain includes large-scale fading, small-scale fading, and shadowing fading. The SIC coefficient, Indicates base station BS i The power currently allocated to the URLLC service.
[0080] 2) Under the punched-hole transmission model, user e iNo power is allocated; all power is allocated to URLLC services for transmission, and SINR is not calculated.
[0081] Finally, according to user e i The SINR, combined with the Shannon channel capacity, yields the user e i Maximum transmission rate:
[0082]
[0083] B represents the allocated channel bandwidth, and (·) indicates that SINR applies to both shared modes.
[0084] For URLLC user u i 1) Under the NOMA transmission model, the base station BS i Connected user u i The SINR is:
[0085]
[0086] Representing user u i The current channel gain;
[0087] 2) In the punched-hole transmission model, user u i The SINR is:
[0088]
[0089] Finally, considering the short packet transmission characteristics of URLLC services, the Shannon channel capacity formula is no longer applicable to this scenario. (Based on user u) i SINR, combined with finite block length coding theory, for user u i Calculate the maximum transmission rate:
[0090]
[0091] in, Represents the subcarrier bandwidth, T represents the mini-slot length, and ε c Indicates the probability of decoding error. This indicates the number of resource blocks allocated to this URLLC service at the same time. Q G () denotes the right-tail function of the Gaussian normal distribution, denoted as: Channel dispersion, used to measure the randomness of a channel relative to a deterministic channel of the same capacity, can be expressed as the following formula:
[0092] Step 3: Aggregate the distance from the current time slot user to the base station and the distance from the previous time slot user to the base station, input the aggregated distance into the constructed neural network model, and output the connection prediction matrix.
[0093] The m-th time slot T m The corresponding user base station distance D m The latency threshold matrix for user services is as follows:
[0094] Construct a neural network model with multiple hidden layers, taking the distance matrix D at the current time m as input. m The distance matrix D at the previous time step m-1 and user service latency threshold The label data indicates whether the corresponding user has left the base station's range. Binary cross-entropy (BCE) is used as the loss function for network training, and the output binary scalar matrix serves as the connection prediction matrix F. pred ,as follows:
[0095]
[0096] This represents the predicted connectivity of a specific user in the current m-th time slot. This indicates that, based on the prediction, the user's requested service can be completed within the service time threshold. This indicates that the transmission could not be completed within the delay threshold or that the user left the cell during transmission. All users within time m... The connection prediction matrix F is formed. pred .
[0097] Step 4: Establish optimization objectives based on the high reliability requirements of URLLC services and the high transmission rate requirements of eMBB services, and combine them with the connection prediction matrix F. pred A comprehensive decision-making process was adopted to maximize eMBB transmission rate and eMBB reliability while ensuring URLLC reliability.
[0098] The optimization issues are as follows:
[0099]
[0100] Among them, R eMBB Indicates the transmission rate of eMBB services; Pr E Indicates the probability of eMBB service transmission failure; Pr U Indicates the probability of URLLC service transmission failure;
[0101] Constraint C1 represents the sum of the reward weights for eMBB rate, eMBB reliability, and URLLC reliability, which equals 1; α i These are the weighting coefficients.
[0102] Constraint C2 represents the total power of the base station's transmitted signals within the cell being lower than the power limit P. max ;P i For base station BS i Power allocated to users within the community;
[0103] Constraint C3 means that the decoding priority of URLLC services is higher than that of eMBB services in the same time slot; They represent base stations (BS) respectively. i Power allocated to URLLC and eMBB users;
[0104] Constraint C4 means filtering out users predicted to be disconnected in the connection prediction matrix in the overall objective function, and the impact of the power allocated to these users on the objective function is ignored.
[0105] Constraint C5 indicates that URLLC services must strictly meet reliability QoS requirements.
[0106] Constraint C6 indicates that for all types of service transmissions, including URLLC and eMBB, to be successful, the transmission delay must be less than the delay threshold. (·) refers to various services including URLLC and eMBB;
[0107] Constraint C7 indicates that for any type of service transmission to be successful, its signal-to-interference-plus-noise ratio (SIR) must be greater than the threshold condition.
[0108] To maximize the transmission rate R, power needs to be allocated to the RBs of different users. According to Shannon's theory, the transmission rate of eMBB users is:
[0109]
[0110] Solving for:
[0111]
[0112] The URLLC user transmission rate can be obtained as follows:
[0113]
[0114] To maximize the above equation, the power distribution should be:
[0115]
[0116] For different eMBB and URLLC shared strategies The calculations are as follows:
[0117]
[0118]
[0119] Step 5: Train the reinforcement learning resource scheduling model based on the Dueling Deep Q-Learning algorithm, solve the optimization objective function, and thus give the optimal selection strategy for URLLC and eMBB services to preempt or share resources, and obtain the results of multi-service joint downlink resource allocation.
[0120] The reinforcement learning resource scheduling model consists of three layers: the first layer network provides the reserved resource ratio for subsequent URLLC services based on the current maximum threshold delay of eMBB services, channel gain, and connection prediction matrix; the second layer network provides the allocation action for each resource block based on the resource reservation rate and connection prediction matrix provided by the upper layer network, combined with the channel gain; and the third layer network provides resource scheduling decisions based on the number of arriving URLLC services and the power allocation of hourly slots.
[0121] like Figure 3 As shown, this is a possible decision-making method for joint NOMA and puncturing scheduling. The frame structure in the model is used to determine the basic time and frequency attributes of the RB. In the time domain, in this embodiment, a time frame consists of 8 mini-frames and the length of a time frame is 1 millisecond. In the frequency domain, an RB bandwidth consists of 12 subcarrier intervals. Given the bandwidth range and RB bandwidth, the total number of RBs in each mini-frame can be determined accordingly.
[0122] The specific process is as follows:
[0123] Input agent BS i The input consists of a resource reservation network, an RB allocation network, and a power allocation decision network. The learning rate is {σ1,σ2,σ3}, the greed factor is {ε1,ε2,ε3}, the experience replay pool is D, the discount factor is γ, and environmental parameters are used.
[0124] Initialize the training round EP = 0 and the current time t = 0; randomly initialize the network parameters; output action 1 based on the current state as the resource reservation rate decision; use action 1 and channel gain as the input state of the second layer network, and output action 2 as the user RB power allocation decision; use action 2 and real-time service arrival as the input of the third layer network, and output action 3 as the joint power allocation decision for eMBB and URLLC; calculate the reward r(s). t ,a t ) and the next state s t+1 and the training set (s t ,a t ,r(s t ,a t ),s t+1 Add it to the experience replay pool D.
[0125] Calculate the loss function and update the network parameters of each layer using backpropagation.
[0126] Termination condition: Is the current EP number less than the set total number of rounds? If it is less, the EP number is increased by 1; otherwise, the loop ends and the trained reinforcement learning model is obtained.
[0127] This method proposes traffic models for URLLC and eMBB users, considering communication interference from adjacent base stations. It predicts whether a user will exceed the connection range of a base station by combining the user-to-base station distance matrix with time-series information. To generate the optimal scheduling strategy, this paper uses the Dueling Deep Q-learning method as an example, enabling the agent to adopt different resource allocation strategies based on the channel state and sequence position information of different users. This maximizes the experience rate for eMBB users while ensuring the reliability of ULRRC transmission. It has the advantage of fully considering the coexistence problem of multiple services in actual communication scenarios and the significance of efficiently utilizing frequency domain resources to meet the transmission indicators of multiple services.
[0128] It should be noted that the Dueling Deep Q-learning method used in this invention is only intended to help readers understand the resource allocation decision-making process of this method. Any modifications and substitutions made to this invention due to different scheduling decision-making methods should be included within the scope of protection of this invention.
Claims
1. A method for multi-service joint downlink resource allocation, characterized in that, The specific steps are as follows: Step 1: Build a communication scenario that includes base stations and users. Divide the users connected to each base station into two types: eMBB service and URLLC service. Record the distance from all users to the base station in each time slot. The base station set is BS = {BS1, BS2, ..., BS} i ,...BS n }, for base station BS i The connected users include: U = {1,2,...,u} representing the set of users who require URLLC services, and E = {1,2,...,e} representing the set of users who require eMBB services; Let the m-th time slot be T. m The corresponding user base station distance is: in, This indicates that user e, who requires eMBB services, is in time slot T. m Corresponding base station BS i distance, This indicates that user u, who requires URLLC services, is in time slot T. m Corresponding base station BS i The distance; Step 2: Quantitatively define the transmission models for eMBB and URLLC users respectively, and calculate the maximum transmission rate of users under different transmission models; Transmission models include the puncturing transmission model and the NOMA transmission model; Step 3: Aggregate the distance from the current time slot user to the base station and the distance from the previous time slot user to the base station, input the aggregated distance into the constructed neural network model, and output the connection prediction matrix. The m-th time slot T m The corresponding user base station distance D m The latency threshold matrix for user services is as follows: Construct a neural network model with multiple hidden layers, taking the distance matrix D at the current time m as input. m The distance matrix D at the previous time step m-1 and user service latency threshold The network is trained using binary cross-entropy as the loss function, and the output binary scalar matrix is used as the connection prediction matrix F. pred ,as follows: This represents the predicted connectivity of a specific user in the current m-th time slot. This indicates that, based on the prediction, the user's requested service can be completed within the service time threshold. This indicates that the transmission could not be completed within the delay threshold or that the user left the cell during transmission; all users within the m-th time slot The connection prediction matrix F is formed. pred ; Step 4: Establish optimization objectives based on the high reliability requirements of URLLC services and the high transmission rate requirements of eMBB services, and combine them with the connection prediction matrix F. pred A comprehensive decision-making process was adopted to maximize eMBB transmission rate and eMBB reliability while ensuring URLLC reliability. The optimization issues are as follows: C2:ΣP i ≤P max Among them, R eMBB Indicates the transmission rate of eMBB services; Pr E Indicates the probability of eMBB service transmission failure; Pr U Indicates the probability of URLLC service transmission failure; Constraint C1 represents the sum of the reward weights for eMBB rate, eMBB reliability, and URLLC reliability, which equals 1; α i These are the weighting coefficients; Constraint C2 represents the total power of the base station's transmitted signals within the cell being lower than the power limit P. max ;P i For base station BS i Power allocated to users within the community; Constraint C3 means that the decoding priority of URLLC services is higher than that of eMBB services in the same time slot; They represent base stations (BS) respectively. i Power allocated to URLLC and eMBB users; Constraint C4 means filtering out users predicted to be disconnected in the connection prediction matrix in the overall objective function, and the impact of the power allocated to these users on the objective function is ignored; Constraint C5 indicates that URLLC services must strictly meet reliability QoS requirements; Constraint C6 indicates that for any type of service transmission to be successful, its transmission delay must be less than the delay threshold. (·) refers to various services including URLLC and eMBB; Constraint C7 indicates that for any type of service transmission to be successful, its signal-to-interference-plus-noise ratio (SIR) must be greater than the threshold condition. Step 5: Train the reinforcement learning resource scheduling model based on the Dueling Deep Q-Learning algorithm, solve the optimization objective function, and thus give the optimal selection strategy for URLLC and eMBB services to preempt or share resources, and obtain the results of multi-service joint downlink resource allocation. The reinforcement learning resource scheduling model consists of three layers: the first layer network provides the reserved resource ratio for subsequent URLLC services based on the current maximum threshold delay of eMBB services, channel gain, and connection prediction matrix; the second layer network provides the allocation action for each resource block based on the resource reservation rate and connection prediction matrix provided by the upper layer network, combined with the channel gain; and the third layer network provides resource scheduling decisions based on the number of arriving URLLC services and the power allocation of hourly slots.
2. The multi-service joint downlink resource allocation method as described in claim 1, characterized in that, In step two, for eMBB user e i : 1) Under the NOMA transmission model, the base station (BS) i Connected user e i The SINR is: in, Indicates base station BS i The power currently allocated to eMBB services, where N0 represents Gaussian white noise. User e i Current channel gain includes large-scale fading, small-scale fading, and shadowing fading. For SIC coefficient, Indicates base station BS i The power currently allocated to URLLC services; 2) Under the punched-hole transmission model, user e i No power is allocated, and SINR is not calculated. Finally, according to user e i The SINR, combined with the Shannon channel capacity, yields the user e i Maximum transmission rate: B represents the allocated channel bandwidth, and (·) indicates that SINR applies to both shared modes; For URLLC user u i : 1) Under the NOMA transmission model, the base station (BS) i Connected user u i The SINR is: On behalf of user u i The current channel gain; 2) Under the punched-hole transmission model, user u i The SINR is: Finally, according to user u i SINR, combined with finite block length coding theory, for user u i Calculate the maximum transmission rate: in, Represents the subcarrier bandwidth, T represents the mini-slot length, and ε c Indicates the probability of decoding error. This indicates the number of resource blocks allocated to this URLLC service at the same time; Indicates channel dispersion; Q G () represents the right-tail function of the Gaussian normal distribution.
Citation Information
Patent Citations
Reserved URLLC hybrid multiple access transmission optimization method and system during eMBB coexistence
CN113099460A
EMBB URLLC multiplexing resource allocation method based on queuing preemption mechanism
CN115734359A